Transformer

E102296

Transformer is a neural network architecture based on self-attention mechanisms that has become the foundation for modern large language models and many state-of-the-art systems in natural language processing.

All labels observed (3)

Label Occurrences
Transformer canonical 3
BERT 2
Transformer architecture 1

How this entity was disambiguated

Statements (50)

Predicate Object
instanceOf deep learning model
neural network architecture
appliedIn computer vision
language modeling
machine translation
multimodal learning
question answering
speech recognition
text summarization
architectureType encoder-decoder
basedOn self-attention mechanism
coreIdea compute attention over all positions in a sequence
enables long-range dependency modeling
foundationFor BERT
GPT
T5
Vision Transformer
linked to: ViT

many large language models
hasVariant Transformer decoder-only
Transformer encoder-only
encoder-decoder Transformer
linked to: EncoderDecoderModel
implementedIn JAX
PyTorch
TensorFlow
inputRepresentation positional embeddings
token embeddings
inspired subsequent attention-based architectures
introducedBy Aidan N. Gomez
Ashish Vaswani
Illia Polosukhin
Jakob Uszkoreit
Llion Jones
Niki Parmar
Noam Shazeer
Łukasz Kaiser
linked to: Lukasz Kaiser
introducedInPaper Attention Is All You Need
introducedInYear 2017
keyOperation scaled dot-product attention
limitation quadratic complexity in sequence length due to self-attention
notableProperty high parallelizability on GPUs and TPUs
primaryComponent multi-head self-attention
position-wise feed-forward network
publishedAtConference NeurIPS 2017
linked to: NeurIPS
reducedRelianceOn convolutional neural networks in sequence modeling
replaced recurrent neural networks in many NLP tasks
supports parallel sequence processing
trainingObjective maximum likelihood estimation for sequence modeling
uses layer normalization
positional encoding
residual connections

How these facts were elicited

Referenced by (6)

Full triples — surface form annotated when it differs from this entity's canonical label.

GPT-3 architecture Transformer
GPT-4 architecture Transformer
Google Search usesAlgorithm BERT
linked to: Transformer
Hugging Face Transformers supportsModelType BERT
linked to: Transformer
Łukasz Kaiser knownFor Transformer architecture
subject linked to: Lukasz Kaiser
linked to: Transformer
Gnarls Barkley song Transformer