Humanity's Last Exam (benchmark)

From Systems Analysis Wiki
Jump to navigation Jump to search

Humanity's Last Exam (HLE) is a benchmark designed to evaluate the capabilities of advanced artificial intelligence (AI) systems on closed-ended academic tasks requiring knowledge and reasoning at the level of top human experts. The benchmark was developed in 2024–2025 by the non-profit organization Center for AI Safety (CAIS) in collaboration with the data company Scale AI[1]. A peer-reviewed version was later published in the journal Nature in January 2026 under the title "A benchmark of expert-level academic questions to assess AI capabilities"[2].

HLE is conceived as a "final academic exam" for AI models—an exceptionally difficult, closed-ended test intended to measure how close modern models are to an expert level and where the gaps in their abilities remain[1]. The benchmark consists of a publicly released set of 2,500 questions covering over one hundred different disciplines, supplemented by a private holdout set of questions used to detect overfitting[1].

Background

By the mid-2020s, leading language models such as GPT-4 and Claude had reached such high performance on popular test suites (for example, MMLU) that many benchmarks ceased to be a reliable measure of progress. Where these tests had once been a challenging frontier, frontier models now exceeded 90% accuracy on them, making it difficult to objectively assess further improvements[3]. Stanford HAI's AI Index 2025 report later cited Humanity's Last Exam as one of the "more challenging benchmarks" developed specifically because the popular tests had reached saturation[4].

In this context, Dan Hendrycks—director of CAIS and the creator of the earlier MMLU benchmark—proposed the idea of a "last exam": a set of maximally difficult questions that could distinguish AI capabilities from the level of a true expert. According to Hendrycks, the direct impetus came from a conversation with entrepreneur Elon Musk, who considered existing tests too easy, remarking that the MMLU questions were "undergrad level" and that he wanted problems "a world-class expert could do"[5]. The project was initially given a working title along the lines of "The Last Stand," which was changed before the public launch to "Humanity's Last Exam" as being less apocalyptic while better matching the closed-form, exam-style format of the questions[5].

Creation of the benchmark

To implement the idea, CAIS partnered with Scale AI, whose chief executive Alexandr Wang and research director Summer Yue helped lead the effort alongside Hendrycks. On September 16, 2024, a global call for the most difficult questions for the future exam was announced. Scientists and specialists worldwide were invited to submit problems capable of stumping the most advanced AI models, and a prize fund of $500,000 was established to attract contributions[3]. The reward structure paid $5,000 for each of the top 50 questions and $500 for the next 500; authors of accepted questions were also offered co-authorship on the resulting paper[6]. The submission deadline, initially set for November 1, 2024, was later extended to November 15, 2024.

The selection of problems proceeded in several stages. First, submissions were filtered using frontier AI models: if the models answered a question correctly (or, for multiple-choice items, did better than random guessing), the question was discarded as not difficult enough. Questions that the models failed were then passed to human experts, who reviewed them in two rounds to verify correctness and ensure a single, unambiguous, verifiable answer that could not be obtained through a simple internet search. From an initial pool of tens of thousands of submissions, nearly 1,000 experts from over 500 institutions across 50 countries—mostly professors, researchers, and graduate-degree holders—contributed to the final dataset[7].

The benchmark was first announced in late January 2025, with the arXiv preprint submitted on January 24, 2025[1]. After release, a public "community feedback bug bounty" program ran until March 21, 2025 to identify and remove erroneous or searchable questions; on April 3, 2025 the dataset was finalized at 2,500 publicly released questions, with flagged items removed and replaced[7]. A portion of the questions is withheld in a private holdout set for control testing and to prevent models from overfitting to a fixed set[6]. The peer-reviewed version of the work appeared in Nature on January 28, 2026[2].

Structure and content of the benchmark

The HLE question set covers a wide range of academic disciplines. The questions are distributed by subject area as follows[2]:

  • Mathematics: ~41%
  • Biology and Medicine: ~11%
  • Computer Science and AI: ~10%
  • Physics: ~9%
  • Humanities and Social Sciences: ~9%
  • Chemistry: ~7%
  • Engineering: ~4%
  • Other fields: ~9%

Approximately 14% of all tasks are multimodal, requiring the analysis of an accompanying image (a diagram, figure, or inscription) in addition to text[2]. About 76% of the questions are exact-match (short-answer) items, where the model must independently produce a precise answer (a number, term, or string); the remaining 24% are multiple-choice questions with five or more options[2]. The subject matter is deliberately specialized—examples range from translating ancient Palmyrene inscriptions to identifying microanatomical structures in birds or analyzing features of Biblical Hebrew pronunciation[2].

All tasks in HLE share common properties:

  • Extremely high difficulty: each problem requires knowledge and skill comparable to that of a qualified specialist in the field[8].
  • Verifiable answer: each question has a specific, provably correct answer suitable for automated grading.
  • Resistance to search: the tasks are designed so that the answer cannot be found with a simple search query; success requires deep understanding and reasoning[1].

For automated grading, the benchmark's maintainers use a separate model (o3-mini) as an extractor and judge to compare a model's response against the ground-truth answer, and they note that results can vary slightly with different judge models and prompts[7].

Model performance results

Humanity's Last Exam immediately confirmed its reputation as an extremely challenging test: at release, no contemporary AI model came close to expert-level performance, and the gap has narrowed only gradually as models have improved. Because scores depend heavily on the date and on whether external tools are allowed, the figures below should be read as a time series rather than a fixed ranking.

  • Initial release (early 2025, no tools). Leading models scored in the single digits. GPT-4o reached roughly 3%, while the strongest reasoning models of the moment—OpenAI's o1 and DeepSeek-R1—answered only about 9% of questions correctly[7].
  • Tool-augmented agents (February 2025). OpenAI's experimental Deep Research agent, built on the o3 model and allowed to perform automatic web searches, correctly solved 26.6% of the tasks—roughly three times the best tool-free score at the time, though still far from a passing grade. Its access to search makes direct comparison with tool-free models uneven[9].
  • Through 2025 (no tools). Newer reasoning models pushed tool-free accuracy into the low 20s. Google's Gemini 2.5 Pro was reported at about 18.8% at its March 2025 launch and later in the low-20s, while xAI's Grok 4 reached roughly 25% by mid-year[7].
  • Current state (2026). The frontier has continued to climb steeply. By early 2026, top models on the official leaderboard reached the mid-30s; by mid-2026, public leaderboards placed the leading systems in the mid-40s and above. Reported text-only figures included the mid-40s for Google's Gemini 3.x Pro line and the low-to-mid 40s for the strongest GPT-5 variants, while on the Artificial Analysis board Anthropic's Claude Fable 5 and Claude Opus 4.8 held the top positions (about 53% and 46% respectively), among the first public-leaderboard results above the 50% mark[7][10].

A notable secondary finding concerns calibration. At initial publication, models combined low accuracy (under 10%) with very high stated confidence (calibration error above 80%)—strong evidence that the systems were confabulating rather than recognizing the limits of their own knowledge[7].

Criticism and limitations

Despite its influence, HLE has drawn substantive criticism.

  • Trivia versus intelligence. Some contributors and observers have questioned whether answering highly specialized, graduate-level questions is a meaningful measure of general intelligence. Kevin Zhou, a theoretical-physics researcher at UC Berkeley who contributed questions, noted the large gap between exam performance and genuine research capability[6]. (The peer-reviewed Nature version used the more descriptive title "A benchmark of expert-level academic questions to assess AI capabilities.")
  • Answer accuracy. In July 2025, the AI research organization FutureHouse published an investigation reporting that roughly 29% (95% confidence interval ±3.7 percentage points) of the 321 text-only biology/health and chemistry questions it audited had answers that conflicted with the peer-reviewed literature. The audit deliberately covered only those subsets rather than the benchmark as a whole, and used FutureHouse's PaperQA2 research agent to cross-check each answer's rationale against published evidence[11]. In a follow-up, the HLE team conducted its own targeted re-review of a biology, chemistry, and health subset and reported an expert-disagreement rate of about 18%, arguing that part of the gap reflects genuine disagreement among experts on very hard questions rather than outright errors; FutureHouse separately released a curated "HLE Bio/Chem Gold" subset of validated items[12].
  • Systematic revision. By 2026 these reliability concerns had prompted systematic re-verification of the dataset. The 2026 "HLE-Verified" effort re-audited the public set through expert review and model-based cross-checks, certifying 1,811 of the 2,500 questions (either verified as correct or repaired under preserved evaluation intent) and releasing the remaining 689 as a documented "uncertain" set; the authors report that verification and repair measurably shift downstream model scores, indicating that a portion of measured performance had reflected annotation artifacts rather than genuine capability differences[13].
  • Compensation and process. Some contributors reported unclear payment structures and shifting timelines during development, with several PhD-level experts expressing frustration over ambiguous expectations around the $500,000 prize pool[6].
  • Not a test of general intelligence. By design, HLE measures structured, closed-ended academic knowledge and reasoning. It does not evaluate creativity, initiative, open-ended research ability, or the skill of posing new scientific questions, so even a perfect score would not by itself indicate artificial general intelligence (AGI)[7].

Significance and outlook

The emergence of HLE was a significant event in the AI community, as the benchmark filled a pressing need for a new, more challenging measure of progress.

  • A common baseline. HLE offers researchers and policymakers a standardized—if methodology-sensitive—reference point for assessing AI capabilities, allowing them to track improvement over time and gauge how close machines are to the level of human experts.
  • A tool to inform policy. A shared, standardized reference test supports more substantive discussion of AI development trajectories, potential risks, and possible governance measures.
  • The final frontier of academic testing. The name "Last Exam" reflects the idea that this set of problems could be the final closed-book exam needed to evaluate AI. Passing HLE would mean that, in terms of formal knowledge and rigorously verifiable reasoning, a machine had reached the level of the best human experts[7].

The authors predicted that, given the pace of progress, models might exceed 50% accuracy on HLE by the end of 2025[7]. In practice this threshold was not reached on that timeline—at the close of 2025 the strongest models stood roughly in the mid-20s to high-30s on the official leaderboard—and public leaderboards did not show frontier systems clearly crossing the 50% mark until 2026, roughly half a year later than projected. Crossing it means that machines have come very close to expert level on a narrow but important measure of academic knowledge, while still leaving open-ended scientific and creative work as a separate, unmeasured frontier.

Model results by source

HLE scores are not absolute: they depend on the evaluation date, the judge model and prompt used for grading, whether external tools (e.g. web search) are allowed, and whether the multimodal questions are included or only the text-only subset. Different leaderboards therefore report different numbers for the same model, and figures are best compared within a single source rather than across sources. The two leaderboards below—the benchmark's own (Scale AI / SEAL) and an independent evaluator (Artificial Analysis)—illustrate this: Gemini 3.1 Pro Preview is reported at 47.3% on the Scale text-only board and 44.7% by Artificial Analysis, and the two boards do not even agree on the current leader. Both tables are point-in-time snapshots that change frequently; each is dated below, and for current standings the live leaderboards should be consulted.

Scale AI / SEAL — official leaderboard (text-only)

The benchmark's own leaderboard, run with the Center for AI Safety. Models are evaluated on the text-only subset (about 86% of the dataset) at temperature 0; grading uses o3-mini as an automatic extractor and judge. Rank is the statistical upper bound (a model is ranked above another only when its lower 95% confidence bound exceeds the other's upper bound), so models can share a rank. "Calib. err." is the RMS calibration error—how far a model's stated confidence is from its actual accuracy; high values indicate overconfidence. The standings below are a mid-2026 snapshot and change frequently[14].

Rank (UB) Model Accuracy, % (±95% CI) Calib. err.
1 Gemini 3.1 Pro Preview (thinking high) 47.3 ±2.1 50
1 GPT-5.4 Pro 45.3 ±2.1 37
3 Muse Spark 40.9 ±2.1 51
3 Gemini 3 Pro Preview 37.7 ±2.0 57
4 GPT-5.4 (xhigh thinking) 36.5 ±2.0 42
4 Claude Opus 4.6 Thinking Max 36.2 ±2.0 46
5 GPT-5 Pro 33.3 ±2.0 49
8 GPT-5.2 28.5 ±1.9 45
8 Claude Opus 4.5 Thinking 26.3 ±1.9 55
8 GPT-5 26.3 ±1.9 50
9 GPT-5.1 Thinking 24.7 ±1.8 54
11 Gemini 2.5 Pro Preview (06-05) 22.1 ±1.8 72
For reference — initial-era models (late 2024 / early 2025):
34 o1 (December 2024) 7.8 ±1.1 84
48 Claude 3.5 Sonnet (October 2024) 4.3 ±0.9 83
59 GPT-4o (November 2024) 2.3 ±0.6 88

Artificial Analysis — independent leaderboard

An independent evaluator that runs its own harness on the text-only subset of HLE (2,158 questions), excluding the multimodal questions for cross-model comparability. Its figures are not directly comparable to the Scale board above because of differences in methodology, judge model, and question set. The top ten below are read directly from the Artificial Analysis leaderboard as of 1 July 2026; standings change frequently[10].

# Model Accuracy, %
1 Claude Fable 5 (with fallback) 53.3
2 Claude Opus 4.8 (max) 45.7
3 Gemini 3.1 Pro Preview 44.7
4 GPT-5.5 (xhigh) 44.3
5 GPT-5.5 (high) 43.0
6 Gemini 3.5 Flash 41.0
7 GLM-5.2 (max) 40.1
8 Muse Spark 39.9
9 Claude Sonnet 5 (max) 39.6
10 Qwen3.7 Max 38.1

Literature

  • Liang, P. et al. (2022). Holistic Evaluation of Language Models (HELM). arXiv:2211.09110.
  • Chang, Y. et al. (2023). A Survey on Evaluation of Large Language Models. arXiv:2307.03109.
  • Ni, S. et al. (2025). A Survey on Large Language Model Benchmarks. arXiv:2508.15361.
  • Biderman, S. et al. (2024). The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned. arXiv:2405.14782.
  • Kiela, D. et al. (2021). Dynabench: Rethinking Benchmarking in NLP. arXiv:2104.14337.
  • Ma, Z. et al. (2021). Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking. arXiv:2106.06052.
  • Goel, K. et al. (2021). Robustness Gym: Unifying the NLP Evaluation Landscape. arXiv:2101.04840.
  • Xu, C. et al. (2024). Benchmark Data Contamination of Large Language Models: A Survey. arXiv:2406.04244.
  • Liu, S. et al. (2025). A Comprehensive Survey on Safety Evaluation of LLMs. arXiv:2506.11094.
  • Chiang, W.-L. et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132.
  • Boubdir, M. et al. (2023). Elo Uncovered: Robustness and Best Practices in Language Model Evaluation. arXiv:2311.17295.
  • Huang, L. et al. (2023). A Survey on Hallucination in Large Language Models. arXiv:2311.05232.

References

  1. 1.0 1.1 1.2 1.3 1.4 Phan, L., Gatti, A., Han, Z. et al. "Humanity's Last Exam". arXiv:2501.14249, 2025. [1]
  2. 2.0 2.1 2.2 2.3 2.4 2.5 "A benchmark of expert-level academic questions to assess AI capabilities". Nature, vol. 649, pp. 1139–1146, 2026. DOI: 10.1038/s41586-025-09962-4. [2]
  3. 3.0 3.1 Dastin, J. & Paul, K. "AI experts ready 'Humanity's Last Exam' to stump powerful tech". Reuters, 2024. [3]
  4. Maslej, N. et al. The AI Index 2025 Annual Report. Stanford Institute for Human-Centered AI, April 2025, pp. 141–142.
  5. 5.0 5.1 Roose, K. "When A.I. Passes This Test, Look Out". The New York Times, 23 January 2025.
  6. 6.0 6.1 6.2 6.3 "Humanity's Last Exam". In Wikipedia. [4]
  7. 7.00 7.01 7.02 7.03 7.04 7.05 7.06 7.07 7.08 7.09 "Humanity's Last Exam". Center for AI Safety. [5]
  8. "Could you pass 'Humanity's Last Exam'? Probably not, but neither can AI". TechRadar. [6]
  9. "OpenAI's deep research can complete 26% of 'Humanity's Last Exam': What is it and what does it mean?". Hindustan Times. [7]
  10. 10.0 10.1 "Humanity's Last Exam Benchmark Leaderboard". Artificial Analysis. [8]
  11. FutureHouse. "About 30% of Humanity's Last Exam chemistry/biology answers are likely wrong", 2025. [9]
  12. The HLE organizing team revised its preprint in response to the FutureHouse audit and, in a September 2025 re-review of a biology/chemistry/health subset, reported an expert-disagreement rate of roughly 18%. See arXiv:2501.14249 (revised) and the FutureHouse update at [10].
  13. Zhai, W., Wang, Z., Wang, J. et al. "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam". arXiv:2602.13964, 2026. [11]
  14. "Humanity's Last Exam (Text Only)". Scale AI / SEAL Leaderboards. [12]