Training Compute-Optimal Large Language Models
E1346775
UNEXPLORED
"Training Compute-Optimal Large Language Models" is a research paper that analyzes how to most efficiently allocate computational resources when scaling large language models to maximize performance.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Training Compute-Optimal Large Language Models canonical | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18877030 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Training Compute-Optimal Large Language Models Context triple: [Tom Henighan, coAuthorOf, Training Compute-Optimal Large Language Models]
-
A.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity" is a research paper that introduces a sparsely activated mixture-of-experts transformer architecture enabling efficient training and inference of language models with up to a trillion parameters.
-
B.
Megatron-LM
Megatron-LM is a large-scale language model training framework developed by NVIDIA, designed to efficiently train massive transformer models through model, tensor, and pipeline parallelism.
-
C.
Language Models are Few-Shot Learners
"Language Models are Few-Shot Learners" is a landmark research paper that demonstrated large-scale transformer-based language models can perform diverse tasks from just a few examples without task-specific training.
-
D.
Language Models are Unsupervised Multitask Learners
"Language Models are Unsupervised Multitask Learners" is a 2019 OpenAI research paper that demonstrated how large-scale unsupervised language models like GPT-2 can perform a wide range of tasks without task-specific training.
-
E.
OPT: Open Pre-trained Transformer Language Models
OPT: Open Pre-trained Transformer Language Models is a family of openly released large-scale transformer-based language models developed by Meta AI to provide transparent, reproducible alternatives to proprietary models like GPT-3.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Training Compute-Optimal Large Language Models Target entity description: "Training Compute-Optimal Large Language Models" is a research paper that analyzes how to most efficiently allocate computational resources when scaling large language models to maximize performance.
-
A.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity" is a research paper that introduces a sparsely activated mixture-of-experts transformer architecture enabling efficient training and inference of language models with up to a trillion parameters.
-
B.
Megatron-LM
Megatron-LM is a large-scale language model training framework developed by NVIDIA, designed to efficiently train massive transformer models through model, tensor, and pipeline parallelism.
-
C.
Language Models are Few-Shot Learners
"Language Models are Few-Shot Learners" is a landmark research paper that demonstrated large-scale transformer-based language models can perform diverse tasks from just a few examples without task-specific training.
-
D.
Language Models are Unsupervised Multitask Learners
"Language Models are Unsupervised Multitask Learners" is a 2019 OpenAI research paper that demonstrated how large-scale unsupervised language models like GPT-2 can perform a wide range of tasks without task-specific training.
-
E.
OPT: Open Pre-trained Transformer Language Models
OPT: Open Pre-trained Transformer Language Models is a family of openly released large-scale transformer-based language models developed by Meta AI to provide transparent, reproducible alternatives to proprietary models like GPT-3.
- F. None of above. chosen
Referenced by (1)
Full triples — surface form annotated when it differs from this entity's canonical label.