AI & tech · Français

Will the highest Humanity's Last Exam (HLE) accuracy score in Epoch AI's benchmark data reach at least 58.5 % for a model released on or before 2026-12-06?

deadline 2026-12-06 confidence medium tier 4 evidence
36%PolySignal estimate

Resolution criterion

Resolves YES if, in a snapshot of Epoch AI's benchmark data (https://epoch.ai/data/benchmark_data.zip, file "hle_external.csv", column "Accuracy") taken on or after 2026-12-20, at least one model with a release date on or before 2026-12-06 has a value >= 58.5 %. Reference at creation: 54.8 % (gpt-6-astra_unknown, 2026-09-03). Voided if Epoch AI revises the methodology so that values are no longer comparable, or retires the file. Data: Epoch AI, CC BY 4.0.

Verified starting state

Computed base rate: 36% of past 60-day windows in Epoch AI's benchmark data saw the running record rise by at least the required amount (threshold 0.585, current record 0.5479999999999999). Data: Epoch AI, CC BY 4.0, epoch.ai/benchmarks. Overlapping windows: small effective sample.

What pushes it up

What pushes it down

What would move this number most

Whether a major frontier AI lab releases a new flagship reasoning model or update between October 8 and December 6, 2026, that Epoch AI evaluates on HLE.

How this number is built

Deadline2026-12-06
Base rate36%
Best evidencetier 4 · 1 article(s) used
Independent runs35 · 36 (spread 1 pts)
Confidencemedium
Evidence capnot triggered
Last revised2026-10-08 19:50
Modelgemini-3.7-flash

Sources consulted

Modern AI Benchmarks: What Practitioners Actually Need to Knowtier 4 · Medium · 2026-05-29