Attention Is All You Need

E457850

"Attention Is All You Need" is the landmark 2017 research paper that introduced the Transformer architecture and revolutionized modern natural language processing and sequence modeling.

All labels observed (2)

Label Occurrences
Attention Is All You Need canonical 9
"Attention Is All You Need" 8

How this entity was disambiguated

Statements (53)

Predicate Object
instanceOf computer science paper ⓘ
research paper ⓘ
scientific paper ⓘ
affiliatedInstitution Google Brain ⓘ
Google Research ⓘ
applicationDomain machine translation ⓘ
architectureType encoder-decoder ⓘ
benchmarkDataset WMT 2014 English-to-French translation ⓘ
WMT 2014 English-to-German translation ⓘ
citationStatus highly cited paper ⓘ
enabled parallel training of sequence models ⓘ
field deep learning ⓘ
machine learning ⓘ
natural language processing ⓘ
sequence modeling ⓘ
hasAuthor Aidan N. Gomez ⓘ
Ashish Vaswani ⓘ
Illia Polosukhin ⓘ
Jakob Uszkoreit ⓘ
Llion Jones ⓘ
Niki Parmar ⓘ
Noam Shazeer ⓘ
Łukasz Kaiser ⓘ
linked to: Lukasz Kaiser
impact became foundational for large language models ⓘ
revolutionized modern natural language processing ⓘ
inspiredModel BERT ⓘ
GPT series ⓘ
T5 ⓘ
introducedConcept Transformer architecture ⓘ
multi-head attention ⓘ
positional encoding ⓘ
scaled dot-product attention ⓘ
self-attention mechanism ⓘ
optimizationMethod Adam optimizer ⓘ
outperformed previous state-of-the-art machine translation models ⓘ
proposedModel Transformer ⓘ
publicationYear 2017 ⓘ
publishedIn Advances in Neural Information Processing Systems 30 ⓘ
linked to: NeurIPS
publishedInConference NeurIPS 2017 ⓘ
linked to: NeurIPS
publisher Neural Information Processing Systems Foundation ⓘ
linked to: NeurIPS
reduced sequential computation in sequence models ⓘ
replacedArchitecture GRU networks ⓘ
LSTM networks ⓘ
recurrent neural networks ⓘ
title Attention Is All You Need ⓘ
usesComponent dropout regularization ⓘ
layer normalization ⓘ
multi-head self-attention layers ⓘ
position-wise feed-forward networks ⓘ
residual connections ⓘ
stacked decoder layers ⓘ
stacked encoder layers ⓘ
usesTechnique label smoothing ⓘ

How these facts were elicited

Referenced by (17)

Full triples — surface form annotated when it differs from this entity's canonical label.

Transformer → introducedInPaper → Attention Is All You Need ⓘ
Łukasz Kaiser → coAuthorOf → Attention Is All You Need ⓘ
subject linked to: Lukasz Kaiser
Attention Is All You Need → title → Attention Is All You Need ⓘ
Ashish Vaswani → notableWork → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Ashish Vaswani → coAuthorOf → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Noam Shazeer → coAuthorOf → Attention Is All You Need ⓘ
Niki Parmar → notableWork → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Niki Parmar → coAuthorOf → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Jakob Uszkoreit → coAuthorOf → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Jakob Uszkoreit → notableWork → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Llion Jones → coAuthorOf → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Llion Jones → notableWork → "Attention Is All You Need" ⓘ
linked to: Attention Is All You Need
Aidan N. Gomez → notableWork → Attention Is All You Need ⓘ
Aidan N. Gomez → coAuthorOf → Attention Is All You Need ⓘ
Illia Polosukhin → coAuthorOf → Attention Is All You Need ⓘ
Illia Polosukhin → hasPublication → Attention Is All You Need ⓘ
DETR → inspiredBy → Attention Is All You Need ⓘ