WaveNet

E39544

WaveNet is a deep generative neural network architecture for raw audio that produces highly natural-sounding speech and other audio signals.

AI illustration

How this image was made

AI-generated illustration of WaveNet

This AI-generated illustration was produced by black-forest-labs/FLUX.2-dev (1024x1024) from a prompt written by openai/gpt-oss-120b from the entity's label + description.

Prompt

Generate an image of WaveNet (WaveNet is a deep generative neural network architecture for raw audio that produces highly natural-sounding speech and other audio signals.)

All labels observed (4)

Label Occurrences
WaveNet canonical 14
WaveNet: A Generative Model for Raw Audio 3
WaveNet prior 1

How this entity was disambiguated

Statements (51)

Predicate Object
instanceOf autoregressive model ⓘ
deep generative model ⓘ
neural network architecture ⓘ
text-to-speech model ⓘ
application music generation ⓘ
neural vocoder for parametric TTS ⓘ
voice conversion ⓘ
arxivId 1609.03499 ⓘ
basedOn causal convolutional neural networks ⓘ
designedFor audio signal modeling ⓘ
raw audio generation ⓘ
speech synthesis ⓘ
developedBy DeepMind ⓘ
Google DeepMind ⓘ
linked to: DeepMind
hasProperty high computational cost at inference ⓘ
highly natural-sounding speech output ⓘ
parallelization across time is limited by autoregressive structure ⓘ
improvedUpon concatenative text-to-speech systems ⓘ
parametric HMM-based TTS systems ⓘ
inputType discretized audio waveform samples ⓘ
inspired PixelCNN ⓘ
introducedBy Aaron van den Oord ⓘ
Alex Graves ⓘ
Heiga Zen ⓘ
Karen Simonyan ⓘ
Nal Kalchbrenner ⓘ
Oriol Vinyals ⓘ
Sander Dieleman ⓘ
introducedIn 2016 ⓘ
introducedInPaper WaveNet: A Generative Model for Raw Audio ⓘ
linked to: WaveNet
language Python ⓘ
ledTo Parallel WaveNet ⓘ
WaveGlow ⓘ
WaveRNN ⓘ
neural vocoder architectures ⓘ
outputType probability distribution over next audio sample ⓘ
publishedAt arXiv ⓘ
relatedTo PixelRNN ⓘ
autoregressive image models ⓘ
supports general raw waveform modeling ⓘ
music audio generation ⓘ
speaker-conditioned speech synthesis ⓘ
text-conditioned speech synthesis ⓘ
trainingObjective cross-entropy loss over quantized samples ⓘ
maximum likelihood estimation ⓘ
usedIn Google Assistant text-to-speech ⓘ
linked to: Google Assistant

Google Cloud Text-to-Speech ⓘ
uses autoregressive sample-by-sample prediction ⓘ
conditional generative modeling ⓘ
dilated causal convolutions ⓘ
softmax output over quantized audio samples ⓘ

How these facts were elicited

Referenced by (19)

Full triples — surface form annotated when it differs from this entity's canonical label.

DeepMind → developed → WaveNet ⓘ
WaveNet → introducedInPaper → WaveNet: A Generative Model for Raw Audio ⓘ
linked to: WaveNet
Heiga Zen → notableWork → WaveNet: A Generative Model for Raw Audio ⓘ
linked to: WaveNet
Nal Kalchbrenner → notableWork → WaveNet ⓘ
WaveRNN → comparedTo → WaveNet ⓘ
WaveRNN → improvesOn → WaveNet ⓘ
WaveRNN → moreEfficientThan → WaveNet ⓘ
WaveGlow → basedOn → WaveNet ⓘ
WaveGlow → comparedWith → WaveNet ⓘ
Karen Simonyan → knownFor → WaveNet ⓘ
Google Cloud Text-to-Speech → supports → WaveNet voices ⓘ
linked to: WaveNet
Parallel WaveNet → basedOn → WaveNet ⓘ
Aaron van den Oord → knownFor → WaveNet ⓘ
Aaron van den Oord → developed → WaveNet ⓘ
Aaron van den Oord → notableWork → WaveNet: A Generative Model for Raw Audio ⓘ
linked to: WaveNet
ClariNet → relatedTo → WaveNet ⓘ
Tacotron → canBeUsedWith → WaveNet ⓘ
VQ-VAE → usedWith → WaveNet prior ⓘ
linked to: WaveNet