Will the highest Humanity's Last Exam (HLE) accuracy score in Epoch AI's benchmark data reach at least 58.5 % for a model released on or before 2026-12-06?
deadline 2026-12-06
confidence medium
tier 4 evidence
36%PolySignal estimate
Resolution criterion
Resolves YES if, in a snapshot of Epoch AI's benchmark data (https://epoch.ai/data/benchmark_data.zip, file "hle_external.csv", column "Accuracy") taken on or after 2026-12-20, at least one model with a release date on or before 2026-12-06 has a value >= 58.5 %. Reference at creation: 54.8 % (gpt-6-astra_unknown, 2026-09-03). Voided if Epoch AI revises the methodology so that values are no longer comparable, or retires the file. Data: Epoch AI, CC BY 4.0.
Verified starting state
Computed base rate: 36% of past 60-day windows in Epoch AI's benchmark data saw the running record rise by at least the required amount (threshold 0.585, current record 0.5479999999999999). Data: Epoch AI, CC BY 4.0, epoch.ai/benchmarks. Overlapping windows: small effective sample.
What pushes it up
- Rapid frontier model reasoning advances have historically produced record jumps of this magnitude within a 60-day window [tier 1].
What pushes it down
- Humanity's Last Exam is designed specifically to resist rapid saturation, making a +3.7 percentage point increase within two months difficult [tier 1].
- No verified announcements or test results indicate an imminent model release exceeding 58.5% before December 2026 [tier 4].
What would move this number most
Whether a major frontier AI lab releases a new flagship reasoning model or update between October 8 and December 6, 2026, that Epoch AI evaluates on HLE.
How this number is built
| Deadline | 2026-12-06 |
| Base rate | 36% |
| Best evidence | tier 4 · 1 article(s) used |
| Independent runs | 35 · 36 (spread 1 pts) |
| Confidence | medium |
| Evidence cap | not triggered |
| Last revised | 2026-10-08 19:50 |
| Model | gemini-3.7-flash |
Sources consulted