“GPT-Rosalind-5.5 is a new model in the GPT-Rosalind series, our frontier reasoning model built to support research across biology, drug discovery, and translational medicine. We are deploying it in research preview to trusted organizations, providing access limited to qualified scientists, research institutes and government partners who are working on beneficial uses and who have a strong security”
| Benchmark | Variant | Score |
|---|---|---|
| biorisk_knowledge | cons_at_32 | 81.7% |
| biorisk_knowledge | cons_at_32 | 81.1% |
| biorisk_knowledge | cons_at_32 | 78.3% |
| LifeSciBench | pass_at_1 | 63.4% |
| LabworkBench | pass_at_1 | 63.2% |
| LifeSciBench | pass_at_1 | 58.8% |
| LabworkBench | pass_at_1 | 55.8% |
| virology_troubleshooting | multi_select, pass_at_1 | 55.3% |
Showing top 8 of 27. See full list below.
- “not deploying automated monitors for real-time blocking of potentially unsafe generations.”
- “available only to approved customers.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 8a65255cd054 · version dated Aug 3, 2026.
Extracted Evaluations(27 results)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
biorisk_knowledge | other | scored | 81.7 cons at 32 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
biorisk_knowledge | other | scored | 81.1 cons at 32 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
biorisk_knowledge | other | scored | 78.3 cons at 32 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
LifeSciBench | other | scored | 63.4 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
LabworkBench | other | scored | 63.2 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
LifeSciBench | other | scored | 58.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
LabworkBench | other | scored | 55.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
virology_troubleshooting/ multi_select | other | scored | 55.3 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
virology_troubleshooting/ multi_select | other | scored | 53.7 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 53.3 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
virology_troubleshooting/ multi_select | other | scored | 51.2 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 49.6 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 44.1 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | scored | 37.3 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ open_ended | other | scored | 36.7 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ open_ended | other | scored | 36.4 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
medicinal_chemistry | other | scored | 27.5 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
medicinal_chemistry | other | scored | 25.1 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Genebench | other | scored | 21.6 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Genebench | other | scored | 20.4 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
dna_sequence_design_tf_binding | other | scored | 16.5 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
dna_sequence_design_tf_binding | other | scored | 13.8 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
dna_sequence_design_tf_binding | other | scored | 13.6 pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 3.1 pass at 4 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 0.4 pass at 4 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 0.0 pass at 4 | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |