Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
2,243-word document condensed to 64 words. Google DeepMind · Jul 27, 2026
TL;DR
“the Google DeepMind site for a comprehensive list of model cards. This model card includes more essential information about the Gemini 3 family of models than previous model”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| key skills benchmark | v1_hard, solve_rate | 91.7% |
| Misalignment | situational_awareness, solve_rate | 27.3% |
| Misalignment | stealth, solve_rate | 25.0% |
| Tone | — | 7.9% |
| Unjustified-refusals | — | 3.7% |
| Image to Text Safety | — | 3.1% |
| Multilingual Safety | Average | 20.0% |
| key skills benchmark | v2, solve_rate | 0.0% |
Showing top 8 of 10. See full list below.
Deployment scope
- “available via Notebook LM.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 108e93b43e9d · version dated Jul 27, 2026.
Extracted Evaluations(10 results)
Sort by:0/10 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
key skills benchmark/ v1_hard | other | scored | 91.7 solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Misalignment/ situational_awareness | other | scored | 27.3 solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Misalignment/ stealth | other | scored | 25.0 solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 7.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 3.7 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 3.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 0.2 | Averagemissing: shot countmissing: methodmissing: training state | self-reported | |
key skills benchmark/ v2 | other | scored | 0.0 solve rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | -10.4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |