Model Cards / Meta AI

Llama 3.2 Model Card

model card4,158 words·18 min read·Mar 31, 2026·Source
Summary

Llama 3.2 Model Card

A 555-word brief of a 4,158-word document. Published by Meta AI. Version dated Mar 31, 2026.
01

What this is

Llama 3.2 is a collection of multilingual large language models (1B and 3B parameters) released by Meta on October 24, 2024. The collection includes pretrained and instruction-tuned text-only variants, plus quantized versions (SpinQuant and QLoRA) designed for on-device deployment. It extends Llama 3.1 to smaller-scale use cases, targeting multilingual dialogue, agentic retrieval, and summarization tasks.

02

Capabilities

The instruction-tuned 3B model scores 63.4 on MMLU (5-shot), 77.7 on GSM8K (CoT), and 77.4 on IFEval; the 1B model scores 49.3, 44.4, and 59.5 on the same benchmarks respectively. Both models support multilingual text input and output across 8 officially supported languages (English, German, French, Italian, Portuguese, Hindi, Spanish, Thai), with a 128k token context window reduced to 8k in quantized variants. Models were pretrained on up to 9 trillion tokens with a knowledge cutoff of December 2023.

03

Evaluation methodology

Meta used an internal evaluations library to run standard automatic benchmarks covering general reasoning, math, instruction following, tool use, long context, and multilingual categories. Dedicated adversarial evaluation datasets were built for safety evaluation and applied to systems composed of Llama models paired with Purple Llama safeguards. The card does not describe contamination controls for the benchmark suite.

04

Safety testing

Red teaming was conducted by experts in cybersecurity, adversarial machine learning, responsible AI, and integrity, including multilingual content specialists with market-specific background. For CBRN risks, uplift testing was performed on Llama 3.1 70B and 405B, and Meta states it "broadly believe[s] that the testing conducted for the 405B model also applies to Llama 3.2 models." Child safety assessments used expert-led, objective-based red teaming across multiple attack vectors and languages. Cyber attack uplift testing evaluated both skill augmentation and fully autonomous offensive agent scenarios, again extrapolated from Llama 3.1 405B results.

05

Mitigations

Instruction-tuned models undergo multiple rounds of SFT, rejection sampling, and DPO using safety-oriented training data that combines human-generated and synthetic examples; LLM-based classifiers select high-quality prompts and responses. Refusal training emphasizes tone guidelines and covers both borderline and adversarial prompts. System-level safeguards—Llama Guard, Prompt Guard, and Code Shield—are released as open-source tools and included by default in Meta's reference implementations. For constrained or mobile environments, Meta recommends Llama Guard 3-1B or its mobile-optimized variant.

06

Deployment and access

Llama 3.2 is governed by the Llama 3.2 Community License, a custom commercial agreement. Models are available for commercial and research use across supported languages, with quantized variants specifically targeting on-device and mobile deployment via the ExecuTorch framework. Use in violation of applicable laws, the Acceptable Use Policy, or in languages beyond those explicitly supported is out of scope.

07

Limitations

Meta states that "testing conducted to date has not covered, nor could it cover, all scenarios," and the model "may in some instances produce inaccurate, biased or other objectionable responses." The 1B and 3B models "will have a different alignment profile and safety/helpfulness tradeoff than more complex, larger systems." The model is static, trained on an offline dataset with a December 2023 knowledge cutoff, and future outputs cannot be predicted in advance.

08

What's new

Relative to Llama 3.1, the 1B and 3B models introduce knowledge distillation during pretraining, using logits from Llama 3.1 8B and 70B as token-level targets, followed by pruning recovery. New quantized variants—SpinQuant and QLoRA—are released for on-device inference, achieving 2.4–2.6x decode speed improvements and 45–60% memory reductions over BF16 baselines on an Android OnePlus 12 device. The quantization scheme is designed around the ExecuTorch framework with ARM CPU backends.

Generated by Claude sonnet from the cleaned source on Apr 23, 2026. Passages in double quotes are verbatim from the source; other text is neutral paraphrase. For citation, use the original: original document · source SHA 1ab343363789.

Extracted Evaluations(83 results)

Sort by:0/83 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
agentscored
38.5
accuracy
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
agentscored
30.1
accuracy
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
knowledgescored
62.5%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
62.3%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
62.1%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
61.6%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
60.6%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
55.1%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
54.6%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
54.5%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
53.8%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
53.6%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
53.6%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
53.4%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
53.3%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
53.3%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
53.3%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
53.3%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
52.2%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
52.1%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
51.9%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
51.7%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
51.3%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
51.2%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
50.9%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
50.9%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
50.3%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
50.0%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
49.9%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
44.5%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
44.0%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
43.3%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
42.2%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
42.1%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
42.0%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
41.8%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
41.5%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
41.3%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
40.8%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
40.6%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
40.5%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
40.4%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
40.2%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
39.8%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
39.8%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
39.8%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
39.6%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
39.2%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
39.2%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
38.9%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
38.1%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
37.5%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
36.0%
accuracy
5-shotSpanishinstruction-tunedmissing: method
self-reported
knowledgescored
34.9%
accuracy
5-shotGermaninstruction-tunedmissing: method
self-reported
knowledgescored
34.9%
accuracy
5-shotItalianinstruction-tunedmissing: method
self-reported
knowledgescored
34.9%
accuracy
5-shotPortugueseinstruction-tunedmissing: method
self-reported
knowledgescored
34.9%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
34.8%
accuracy
5-shotFrenchinstruction-tunedmissing: method
self-reported
knowledgescored
34.7%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
34.0%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
33.5%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
32.4%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
32.1%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
knowledgescored
31.2%
accuracy
5-shotThaiinstruction-tunedmissing: method
self-reported
knowledgescored
30.0%
accuracy
5-shotHindiinstruction-tunedmissing: method
self-reported
/ en_mc
long_contextscored
72.2
accuracy
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ en_mc
long_contextscored
63.3
accuracy
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ en_mc
long_contextscored
38.0
accuracy
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ en_qa
long_contextscored
27.3
f1
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ en_qa
long_contextscored
20.3
f1
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ en_qa
long_contextscored
19.8
f1
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
multilingualscored
68.9%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
58.2%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
56.8%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
54.3%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
48.9%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
24.5%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
24.4%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
18.2%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
multilingualscored
13.7%
exact match
0-shotcotinstruction-tunedmissing: language
self-reported
/ multi_needle
otherscored
98.8
recall
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ multi_needle
otherscored
84.7
recall
0-shotinstruction-tunedmissing: methodmissing: language
self-reported
/ multi_needle
otherscored
75.0
recall
0-shotinstruction-tunedmissing: methodmissing: language
self-reported