“In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February”
| Benchmark | Variant | Score |
|---|---|---|
| HellaSwag | accuracy | 93.3% |
| BIG-Bench | hard, accuracy | 89.2% |
| MGSM | 8-shot, high_resource, accuracy | 89.1% |
| MGSM | 8-shot, Average, accuracy | 87.5% |
| MGSM | 8-shot, low_resource, accuracy | 86.3% |
| MGSM | 8-shot, high_resource, accuracy | 85.3% |
| MGSM | 8-shot, Average, accuracy | 82.5% |
| MGSM | 8-shot, mid_resource, accuracy | 82.4% |
Showing top 8 of 141. See full list below.
- “we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio.”
- “not released in the Gemini 1.0 models.”
- “mitigations include: Safety filters with established thresholds to set responsible default behaviors.”
- “available to the model.”
- “future work. These models also show improvements in jailbreak robustness and do not respond to “garbage” token attacks, but they do respond to handcrafted prompt injection attacks – potentially due to their increased ability to follow the kind of instructions in the prompt injection.”
- “future work. 9.5. Assurance Evaluations Assurance evaluations are our ‘arms-length’ internal evaluations for responsibility governance decision- making (Weidinger et al., 2024).”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 600781b60d37 · version dated Aug 20, 2026.
Extracted Evaluations(141 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | scored | 74.4% pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| knowledge | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
⚠ 3 others disagree | math | scored | 67.7% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| math | scored | 58.5% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ intermediate_algebra_levels_4_5 | math | scored | 20.6% solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ intermediate_algebra_levels_4_5⚠ 3 others disagree | math | scored | 18.6% solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ intermediate_algebra_levels_4_5 | math | scored | 12.5% solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ high_resource | multilingual | scored | 89.1% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
| multilingual | scored | 87.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
/ low_resource | multilingual | scored | 86.3% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ high_resource | multilingual | scored | 85.3% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
| multilingual | scored | 82.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
/ mid_resource | multilingual | scored | 82.4% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ high_resource | multilingual | scored | 81.6% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ low_resource | multilingual | scored | 79.4% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
| multilingual | scored | 79.0% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
/ mid_resource | multilingual | scored | 78.8% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ low_resource | multilingual | scored | 76.4% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ mid_resource | multilingual | scored | 73.2% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ high_resource | multilingual | scored | 65.7% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
| multilingual | scored | 63.5% accuracy | 8-shotAveragemissing: methodmissing: training state | self-reported | |
/ low_resource | multilingual | scored | 62.5% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
/ mid_resource | multilingual | scored | 53.6% accuracy | 8-shotmissing: methodmissing: languagemissing: training state | self-reported |
| multimodal | scored | 63.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multimodal | scored | 52.1 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multimodal | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multimodal | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multimodal | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
BIG-Bench/ hard | other | scored | 89.2 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ mid_resource | other | scored | 75.8 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ en_to_xx | other | scored | 75.4 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23 | other | scored | 75.3 bleurt | 1-shotAveragemissing: methodmissing: training state | self-reported |
WMT23/ xx_to_en | other | scored | 75.1 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ en_to_xx | other | scored | 74.8 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ high_resource | other | scored | 74.8 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ mid_resource | other | scored | 74.7 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23 | other | scored | 74.4 bleurt | 1-shotAveragemissing: methodmissing: training state | self-reported |
WMT23/ mid_resource | other | scored | 74.3 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ xx_to_en | other | scored | 74.2 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ high_resource | other | scored | 74.2 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23 | other | scored | 74.1 bleurt | 1-shotAveragemissing: methodmissing: training state | self-reported |
WMT23/ en_to_xx | other | scored | 74.0 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ xx_to_en | other | scored | 73.9 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ high_resource | other | scored | 73.9 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 72.0 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ xx_to_en | other | scored | 72.0 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 71.8 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ mid_resource | other | scored | 71.8 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23 | other | scored | 71.7 bleurt | 1-shotAveragemissing: methodmissing: training state | self-reported |
WMT23/ high_resource | other | scored | 71.7 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
WMT23/ en_to_xx | other | scored | 71.5 bleurt | 1-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 71.3 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 70.8 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 70.3 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 65.0 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 63.7 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 63.5 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 63.4 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 62.7 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 58.3 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 57.0 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 56.9 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 54.1 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 52.2 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 51.6 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 50.0 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 49.3 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 48.6 bleurt | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 48.0 bleurt | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 47.3 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 46.5 bleurt | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 43.5 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 39.0 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MTOB/ eng_to_kgv | other | scored | 36.9 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
HiddenMath | other | scored | 36.0 problems solved | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 34.6 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 34.3 bleurt | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 33.3 bleurt | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 32.8 chrf | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 25.0 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MTOB/ eng_to_kgv | other | scored | 21.4 chrf | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
HiddenMath | other | scored | 20.0 problems solved | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 19.0 raw score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MTOB/ eng_to_kgv | other | scored | 17.8 chrf | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 17.8 chrf | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 16.0 chrf | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
HiddenMath | other | scored | 12.0 problems solved | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
HiddenMath | other | scored | 11.0 problems solved | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 5.6 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 5.5 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 5.5 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 5.4 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 5.0 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 4.4 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 4.2 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 4.1 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 3.2 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 2.9 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 2.8 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 2.0 human eval score | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 0.4 human eval score | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 0.3 human eval score | 5-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ kgv_to_eng | other | scored | 0.2 human eval score | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
MTOB/ eng_to_kgv | other | scored | 0.1 human eval score | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
InfographicVQA | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
V* Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
EgoSchema | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
EgoSchema | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FLEURS | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
YouTube | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CoVoST/ 2 | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
1H-VideoQA | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
1H-VideoQA | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Video Needle-in-a-Haystack | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ multi_needle | other | mentioned | — recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BetterChartQA | other | mentioned | — | 0-shotmissing: methodmissing: languagemissing: training state | self-reported |
InfographicVQA | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
HarmBench | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
TDC | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Dolomites | other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ multi_needle | other | mentioned | — recall | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | scored | 93.3% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 46.2% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 41.5% accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 39.5% accuracy | 0-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 35.7% accuracy | 4-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | scored | 27.9% accuracy | 4-shotmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| reasoning | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| vision | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |