Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
52,669-word document condensed to 103 words. Meta AI · Aug 20, 2026
TL;DR
“Llama Team, AI @ Meta1 1A detailed contributor list can be found in the appendix of this paper. Modern artificial intelligence (AI) systems are powered by foundation models.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| GRE | verbal, scaled_score | 167.00 |
| GRE | verbal, scaled_score | 167.00 |
| GRE | verbal, scaled_score | 166.00 |
| GRE | verbal, scaled_score | 166.00 |
| GRE | quantitative, scaled_score | 166.00 |
| GRE | quantitative, scaled_score | 164.00 |
| GRE | verbal, scaled_score | 162.00 |
| GRE | quantitative, scaled_score | 162.00 |
Showing top 8 of 164. See full list below.
Capability claim
- “we present a new set of foundation models for language, calledLlama 3.”
Safety findings
- “not released or because the API does not provide access to log-probabilities.”
Deployment scope
- “available at the time for reward modeling, while only using the latest batches from various capabilities for DPO training.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 578dc31aa9a4 · version dated Aug 20, 2026.
Extracted Evaluations(164 results)
Sort by:⚠ 2 conflicting reports3/164 rows fully reproducible (2%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| agent | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| coding | mentioned | — | pretrainedmissing: shot countmissing: methodmissing: language | self-reported | |
| coding | mentioned | — pass at 1 | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| coding | mentioned | — pass at 1 | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
/ plus | coding | mentioned | — pass at 1 | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
/ evalplus_base | coding | mentioned | — pass at 1 | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
| coding | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| coding | mentioned | — | pretrainedmissing: shot countmissing: methodmissing: language | self-reported | |
| instruction_following | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
/ multilingual⚠ 16 others disagree | knowledge | scored | 85.5% accuracy | 5-shotAveragemissing: methodmissing: training state | self-reported |
/ multilingual | knowledge | scored | 83.2% accuracy | 5-shotAverageinstruction-tunedmissing: method | self-reported |
/ multilingual | knowledge | scored | 80.2% accuracy | 5-shotAveragemissing: methodmissing: training state | self-reported |
/ multilingual | knowledge | scored | 78.2% accuracy | 5-shotAverageinstruction-tunedmissing: method | self-reported |
/ multilingual | knowledge | scored | 64.3% accuracy | 5-shotAveragemissing: methodmissing: training state | self-reported |
/ multilingual | knowledge | scored | 58.8% accuracy | 5-shotAveragemissing: methodmissing: training state | self-reported |
/ multilingual | knowledge | scored | 58.6% accuracy | 5-shotAverageinstruction-tunedmissing: method | self-reported |
/ multilingual | knowledge | scored | 46.8% accuracy | 5-shotAveragemissing: methodmissing: training state | self-reported |
| knowledge | mentioned | — | pretrainedmissing: shot countmissing: methodmissing: language | self-reported | |
/ pro | knowledge | mentioned | — | pretrainedmissing: shot countmissing: methodmissing: language | self-reported |
| knowledge | mentioned | — | 5-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
/ pro | knowledge | mentioned | — | 5-shotcotinstruction-tunedmissing: language | self-reported |
| knowledge | mentioned | — | 5-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| knowledge | mentioned | — | 5-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| knowledge | mentioned | — | 5-shotinstruction-tunedmissing: methodmissing: language | self-reported | |
| knowledge | mentioned | — | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | 5-shotmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | mentioned | — | 5-shotcotinstruction-tunedmissing: language | self-reported |
/ pro | knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ en_mc | long_context | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
/ en_qa | long_context | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| math | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| multilingual | scored | 91.6% accuracy | 0-shotcotAverageinstruction-tuned | self-reported | |
| multilingual | scored | 91.6% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 90.5% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 86.9% accuracy | 0-shotcotAverageinstruction-tuned | self-reported | |
⚠ 2 others disagree | multilingual | scored | 85.9% accuracy | 0-shotcotAveragemissing: training state | self-reported |
| multilingual | scored | 71.1% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 68.9% accuracy | 0-shotcotAverageinstruction-tuned | self-reported | |
| multilingual | scored | 53.2% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 51.4% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
| multilingual | scored | 29.9% accuracy | 0-shotcotAveragemissing: training state | self-reported | |
/ test | multimodal | scored | 95.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 93.1 accuracy | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 92.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 92.6 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
/ test | multimodal | scored | 92.2 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
/ test | multimodal | scored | 90.8 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 88.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 87.2 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 85.8 accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
/ test | multimodal | scored | 85.7 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ test | multimodal | scored | 84.4 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
/ test | multimodal | scored | 83.2 accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
/ test | multimodal | scored | 78.7 accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
/ test | multimodal | scored | 78.4 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 167.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 167.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 166.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 166.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 166.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 164.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 162.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 162.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 161.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 158.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 155.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 154.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ quantitative | other | scored | 152.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GRE/ verbal | other | scored | 149.0 scaled score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 100.0 retrieval accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
Jailbreaks | other | scored | 99.9 true positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Injections | other | scored | 99.5 true positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Out-of-Distribution Jailbreaks | other | scored | 97.5 true positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AI2 Diagram/ test | other | scored | 94.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AI2 Diagram/ test | other | scored | 94.4 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AI2 Diagram/ test | other | scored | 94.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AI2 Diagram/ test | other | scored | 94.1 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
AP | other | scored | 93.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AI2 Diagram/ test | other | scored | 93.0 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
AP | other | scored | 93.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AP | other | scored | 92.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Multilingual Jailbreaks | other | scored | 91.5 true positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AP | other | scored | 87.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
TextVQA/ val | other | scored | 84.8 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
AI2 Diagram/ test | other | scored | 84.4 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
TextVQA/ val | other | scored | 83.4 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
AP | other | scored | 81.3 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
VQAv2/ test-dev | other | scored | 80.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
VQAv2/ test-dev | other | scored | 80.2 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
VQAv2/ test-dev | other | scored | 79.1 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
TextVQA/ val | other | scored | 78.7 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
TextVQA/ val | other | scored | 78.2 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
AI2 Diagram/ test | other | scored | 78.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
VQAv2/ test-dev | other | scored | 78.0 accuracy | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
TextVQA/ val | other | scored | 78.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
VQAv2/ test-dev | other | scored | 77.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AP | other | scored | 74.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Indirect Injections | other | scored | 71.4 true positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
AP | other | scored | 70.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 precision | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 precision | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 f1 | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 f1 | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.9 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 recall | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.9 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.9 recall | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.9 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.9 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.9 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.9 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.8 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.8 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.8 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.8 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.8 recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.8 precision | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.2 false positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ tool_use | other | scored | 0.2 false positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.0 false positive rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ english | other | scored | 0.0 false positive rate | ENmissing: shot countmissing: methodmissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.0 false positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Llama Guard/ multilingual | other | scored | 0.0 false positive rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Safety Violation Rate/ indiscriminate_weapons | other | scored | 0.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ privacy | other | scored | -40.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ suicide_self_harm | other | scored | -62.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ violent_crimes | other | scored | -67.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ specialized_advice | other | scored | -70.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ sex_related_crimes | other | scored | -75.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ non_violent_crimes | other | scored | -80.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ intellectual_property | other | scored | -88.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
Safety Violation Rate/ sexual_content | other | scored | -100.0 violation rate reduction | ENinstruction-tunedmissing: shot countmissing: method | self-reported |
| other | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
API-Bank | other | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported |
API-Bench | other | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported |
/ multi_needle | other | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported |
ZeroSCROLLS | other | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported |
ZeroSCROLLS | other | mentioned | — | 0-shotinstruction-tunedmissing: methodmissing: language | self-reported |
Conic10k | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
VQAv2 | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ICL Consistency Test | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoVoST/ 2 | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Super-NaturalInstructions | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| reasoning | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| reasoning | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| reasoning | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| reasoning | mentioned | — | instruction-tunedmissing: shot countmissing: methodmissing: language | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ val | vision | scored | 69.1% accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ val | vision | scored | 68.3% accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ val | vision | scored | 64.5% accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
/ val | vision | scored | 62.2% accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ val | vision | scored | 60.6% accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
/ val | vision | scored | 56.4% accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported |
/ val | vision | scored | 49.6% accuracy | cotinstruction-tunedmissing: shot countmissing: language | self-reported |
| vision | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |