HuBERT

E435884

HuBERT is a self-supervised speech representation learning model that learns powerful audio features from unlabeled speech for tasks like automatic speech recognition and audio classification.

All labels observed (5)

How this entity was disambiguated

Statements (51)

Predicate Object
instanceOf deep learning model ⓘ
self-supervised speech representation learning model ⓘ
speech foundation model ⓘ
basedOn Transformer architecture ⓘ
designedFor audio classification ⓘ
automatic speech recognition ⓘ
downstream speech tasks ⓘ
self-supervised learning from speech ⓘ
speech representation learning ⓘ
spoken language understanding ⓘ
developedBy Facebook AI Research ⓘ
Meta AI ⓘ
evaluationBenchmark Libri-light ⓘ
LibriSpeech ⓘ
TIMIT ⓘ
hasAuthor Benjamin Bolte ⓘ
James Glass ⓘ
Kyunghyun Cho ⓘ
Wei-Ning Hsu ⓘ
Yao-Hung Hubert Tsai ⓘ
hasVariant HuBERT Base ⓘ
linked to: HuBERT

HuBERT Large ⓘ
linked to: HuBERT

HuBERT X-Large ⓘ
linked to: HuBERT
improvesOver wav2vec 2.0 on several speech benchmarks ⓘ
inputType acoustic features such as log-mel filterbanks ⓘ
raw audio waveforms ⓘ
introducedIn 2021 ⓘ
introducedInPaper HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units ⓘ
linked to: HuBERT
language primarily English in original experiments ⓘ
learningSignalSource discrete units obtained by clustering MFCC or filterbank features ⓘ
maskingStrategy masking of contiguous time spans in the input sequence ⓘ
openSourceImplementation Hugging Face Transformers ⓘ
fairseq ⓘ
linked to: Fairseq
outputType contextualized speech representations ⓘ
frame-level audio embeddings ⓘ
pretrainingStage masked region prediction of cluster assignments ⓘ
offline clustering of acoustic features ⓘ
publishedAt Interspeech 2021 ⓘ
linked to: INTERSPEECH
relatedTo Data2Vec ⓘ
WavLM ⓘ
wav2vec 2.0 ⓘ
linked to: Wav2Vec2
supportsTask audio event classification ⓘ
automatic speech recognition fine-tuning ⓘ
emotion recognition from speech ⓘ
keyword spotting ⓘ
phoneme recognition ⓘ
speaker recognition ⓘ
trainingDataType unlabeled speech audio ⓘ
trainingParadigm self-supervised learning ⓘ
usesObjective cluster-based prediction task ⓘ
masked prediction of latent speech units ⓘ

How these facts were elicited

Referenced by (6)

Full triples — surface form annotated when it differs from this entity's canonical label.

Wav2Vec2 → inspired → HuBERT ⓘ
HuBERT → introducedInPaper → HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units ⓘ
linked to: HuBERT
HuBERT → hasVariant → HuBERT Base ⓘ
linked to: HuBERT
HuBERT → hasVariant → HuBERT Large ⓘ
linked to: HuBERT
HuBERT → hasVariant → HuBERT X-Large ⓘ
linked to: HuBERT