“Large language models (LLMs) are being deployed in many domains of our lives ranging from browsing, to voice assistants, to coding assistance tools, and have potential for vast societal impacts.[1, 2, 3, 4, 5, 6, 7] This system card analyzes GPT-4, the latest LLM in the GPT family of models.[ 8, 9, 10] First, we highlight safety challenges presented by the model’s limitations (e.g., producing conv”
| Benchmark | Variant | Score |
|---|---|---|
| TruthfulQA | post-mitigation, accuracy | 60.0% |
| TruthfulQA | pre-mitigation, accuracy | 30.0% |
Showing top 2 of 4. See full list below.
- “we trained a range of classifiers on new risk vectors and have incorporated these into our monitoring workflow, enabling us to better enforce our API usage policies.”
- “We believe this has reduced the risk surface, though has not completely eliminated it. Today’s deployment represents a balance between minimizing risk from deployment, enabling positive use cases, and learning from deployment.”
- “available to proliferators, especially in comparison to traditional search tools.”
- “Further research is needed to fully characterize these risks.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 91d85b8a8a8e · version dated Aug 20, 2026.
Extracted Evaluations(4 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| safety | scored | 60.0% accuracy | post-mitigationmissing: shot countmissing: languagemissing: training state | self-reported | |
| safety | scored | 30.0% accuracy | pre-mitigationmissing: shot countmissing: languagemissing: training state | self-reported | |
| safety | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| safety | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |