DeBERTa

E435870

DeBERTa is a transformer-based language model developed by Microsoft that improves upon BERT and RoBERTa using disentangled attention and enhanced mask decoder mechanisms for superior natural language understanding.

All labels observed (10)

How this entity was disambiguated

Statements (50)

Predicate Object
instanceOf pretrained language model ⓘ
transformer-based language model ⓘ
architectureComponent embedding layer ⓘ
feed-forward network ⓘ
layer normalization ⓘ
multi-head self-attention ⓘ
availableOn Hugging Face Transformers ⓘ
basedOn Transformer architecture ⓘ
benchmark GLUE ⓘ
SQuAD ⓘ
linked to: SQuAD 2.0

SuperGLUE ⓘ
developer Microsoft ⓘ
hasVersion DeBERTa-XL ⓘ
linked to: DeBERTa

DeBERTa-base ⓘ
linked to: DeBERTa

DeBERTa-large ⓘ
linked to: DeBERTa

DeBERTa-v1 ⓘ
linked to: DeBERTa

DeBERTa-v2 ⓘ
linked to: DeBERTa

DeBERTa-v3 ⓘ
linked to: DeBERTa

DeBERTa-xlarge ⓘ
linked to: DeBERTa

DeBERTa-xxlarge ⓘ
linked to: DeBERTa
implementation PyTorch ⓘ
improves context representation ⓘ
word representation ⓘ
improvesUpon BERT ⓘ
RoBERTa ⓘ
introducedBy Microsoft Research ⓘ
language English ⓘ
license MIT License (for official Microsoft implementation) ⓘ
linked to: MIT License
optimization Adam optimizer (typical training setup) ⓘ
outperforms BERT on GLUE ⓘ
RoBERTa on GLUE (for larger variants) ⓘ
paperTitle DeBERTa: Decoding-enhanced BERT with Disentangled Attention ⓘ
linked to: DeBERTa
publicationType research paper ⓘ
releasedYear 2020 ⓘ
supportsTask natural language inference ⓘ
question answering ⓘ
sentiment analysis ⓘ
sequence labeling ⓘ
text classification ⓘ
token classification ⓘ
task natural language understanding ⓘ
trainingObjective masked language modeling ⓘ
replaced token detection (for some versions) ⓘ
uses absolute position embeddings (for some variants) ⓘ
relative position embeddings ⓘ
usesMechanism disentangled attention ⓘ
enhanced mask decoder ⓘ
usesPretrainingData BookCorpus (for some variants) ⓘ
linked to: BookCorpus

Wikipedia ⓘ
large-scale web text ⓘ

How these facts were elicited

Referenced by (11)

Full triples — surface form annotated when it differs from this entity's canonical label.

DeBERTa → hasVersion → DeBERTa-v1 ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-v2 ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-v3 ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-XL ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-large ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-base ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-xlarge ⓘ
linked to: DeBERTa
DeBERTa → hasVersion → DeBERTa-xxlarge ⓘ
linked to: DeBERTa
DeBERTa → paperTitle → DeBERTa: Decoding-enhanced BERT with Disentangled Attention ⓘ
linked to: DeBERTa