HLE

    ๐Ÿ† Leaderboard

    As of September 2, 2026, Claude Fable 5.1 is #1 for HLE at 65%. Ranked by Humanity's Last Exam: 2,500 expert questions designed so search engines fail. Humanity's Last Exam leaderboard with HLE scores and API prices. Compare the hardest public reasoning exam across GPT, Claude, Gemini, and open models.

    Updated September 2, 2026347 models36 providers

    HLE vs price

    117 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    Claude Mythos 5.1
    Anthropic ยท Proprietary
    65%โ€”โ€”$60.00
    Claude Fable 5.1
    Anthropic ยท Proprietary
    65%โ€”โ€”$60.00
    Claude Opus 5
    Anthropic ยท Proprietary
    64.7%โ€”92.9%$30.00
    Claude Mythos Preview
    Anthropic ยท Proprietary
    64.7%94.6%โ€”$60.00
    Claude Fable 5
    Anthropic ยท Proprietary
    64.5%โ€”85.9%$60.00
    GLM-5.3OSS
    Z AI ยท Open Source
    62.5%โ€”โ€”$5.80
    Muse Spark 1.1
    Meta ยท Proprietary
    62.1%โ€”91.2%$5.50
    DeepSeek-V4-Pro-0813OSS
    DeepSeek ยท Open Source
    60%โ€”92.4%$5.28
    Claude Opus 4.8
    Anthropic ยท Proprietary
    57.9%93.6%85.3%$30.00
    Claude Sonnet 5
    Anthropic ยท Proprietary
    57.4%โ€”80.3%$12.00
    GPT-5.5 Pro
    OpenAI ยท Proprietary
    57.2%โ€”โ€”$540.00
    Kimi K3OSS
    Moonshot AI ยท Open Source
    56%93.5%93.5%$18.00
    Seed 2.1 Pro
    ByteDance ยท Proprietary
    55.7%โ€”โ€”$5.34
    GLM-5.3-FlashOSS
    Z AI ยท Open Source
    55.3%โ€”โ€”$0.65
    DeepSeek-V4-Flash-Vision-Exp
    DeepSeek ยท Proprietary
    55.1%โ€”โ€”$0.88
    GLM-5.2OSS
    Z AI ยท Open Source
    54.7%91.2%71.2%$5.80
    Claude Opus 4.7
    Anthropic ยท Proprietary
    54.7%94.2%86.4%$30.00
    Seed 2.1 Turbo
    ByteDance ยท Proprietary
    54.6%โ€”โ€”$3.00
    Claude Opus 4.6
    Anthropic ยท Proprietary
    53.1%91.3%88.4%$30.00
    GLM-5.1OSS
    Z AI ยท Open Source
    52.3%86.2%89.9%$5.80
    GPT-5.5
    OpenAI ยท Proprietary
    52.2%93.6%77.3%$35.00
    Gemini 3.1 Pro
    Google ยท Proprietary
    51.4%94.3%94.4%$17.50
    Kimi K2-Thinking-0905OSS
    Moonshot AI ยท Open Source
    51%84.5%โ€”$2.47
    Kimi K2.5OSS
    Moonshot AI ยท Open Source
    50.2%87.6%โ€”$3.68
    Qwen3.5-27BOSS
    Qwen ยท Open Source
    48.5%85.5%โ€”$2.70
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek ยท Open Source
    48.2%90.1%89.7%$5.22
    Qwen3.5-122B-A10BOSS
    Qwen ยท Open Source
    47.5%86.6%โ€”$3.60
    Qwen3.5-35B-A3BOSS
    Qwen ยท Open Source
    47.4%84.2%โ€”$2.25
    Claude Sonnet 4.6
    Anthropic ยท Proprietary
    46.8%89.9%78.8%$18.00
    Gemini 3 Pro
    Google ยท Proprietary
    45.8%91.9%91.9%$14.00
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek ยท Open Source
    45.1%88.1%โ€”$0.42
    Qwen3.8 MaxOSS
    Qwen ยท Open Source
    43.6%92.6%92.7%$8.00
    Gemini 3 Flash
    Google ยท Proprietary
    43.5%90.4%89.4%$3.50
    GLM-4.7OSS
    Z AI ยท Open Source
    42.8%85.7%83.3%$2.80
    Qwen3.7 Max
    Qwen ยท Proprietary
    41.4%92.4%90.9%$5.00
    DeepSeek-V3.2OSS
    DeepSeek ยท Open Source ยท via OpenRouter
    40.8%82.4%โ€”$0.57
    DeepSeek-V4-Flash-0423OSS
    DeepSeek ยท Open Source
    40.3%87.4%โ€”$0.30
    Gemini 3.5 Flash
    Google ยท Proprietary
    40.2%โ€”88.9%$10.50
    Grok-4
    xAI ยท Proprietary
    40%87.5%87.5%$18.00
    GPT-5.4
    OpenAI ยท Proprietary
    39.8%92.8%89.4%$17.50
    ERNIE 5.0
    Baidu ยท Proprietary
    39%85%โ€”$7.00
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA ยท Open Source ยท via OpenRouter
    37.4%87%86.1%$2.70
    GPT-5.2 Pro
    OpenAI ยท Proprietary
    36.6%93.2%โ€”$189.00
    Kimi K2.6OSS
    Moonshot AI ยท Open Source
    36.4%90.5%90.8%$4.93
    Qwen3.8-Flash-NextOSS
    Qwen ยท Open Source
    35.9%91.7%โ€”$0.62
    Qwen3.8 Flash
    Qwen ยท Proprietary
    35.9%91.7%โ€”$0.62
    Qwen3.7-Plus
    Qwen ยท Proprietary
    34.7%90.3%87.9%$1.60
    GPT-5.2
    OpenAI ยท Proprietary
    34.5%92.4%73.2%$15.75
    MiMo-V2.5-ProOSS
    Xiaomi ยท Open Source
    34%66.7%82.6%$1.30
    Inkling-SmallOSS
    Thinking Machines ยท Open Source ยท via Thinking Machines Lab
    31.6%89.5%88.5%$1.50
    Showing 1โ€“50 of 347 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    347 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads HLE right now?

    As of September 2, 2026, Claude Fable 5.1 by Anthropic is #1 for HLE at 65%. Ranked by Humanity's Last Exam: 2,500 expert questions designed so search engines fail. This board also tracks HLE, GPQA, GPQA Diamond, Humanity's Last Exam (With Tools). Next on the same board: Claude Mythos 5.1 and Claude Opus 5. Related leaders: GPT-5.6 Sol on GPQA at 94.6%; Gemini 3.7 Flash on GPQA Diamond at 94.8%. This hle leaderboard ranks models by HLE. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: Humanity's Last Exam (https://agi.safe.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for hle. Ranked by Humanity's Last Exam: 2,500 expert questions designed so search engines fail. Input and output are dollars per million tokens.
    RankModelHLEInput /MOutput /M
    1Claude Fable 5.165%$10.00$50.00
    2Claude Mythos 5.165%$10.00$50.00
    3Claude Opus 564.7%$5.00$25.00
    4Claude Mythos Preview64.7%$10.00$50.00
    5Claude Fable 564.5%$10.00$50.00
    6GLM-5.362.5%$1.40$4.40
    7Muse Spark 1.162.1%$1.25$4.25
    8DeepSeek-V4-Pro-081360%$1.32$3.96

    HLE FAQ

    Who ranks #1 on the HLE leaderboard?

    As of September 2, 2026, Claude Fable 5.1 by Anthropic ranks #1 on HLE at 65%. API pricing is $10.00/M input and $50.00/M output.

    What are the top models on HLE?

    The current HLE ranking as of September 2, 2026 is 1. Claude Fable 5.1 at 65%; 2. Claude Mythos 5.1 at 65%; 3. Claude Opus 5 at 64.7%.

    Which hle model is the cheapest?

    Nemotron 3.5 Lightning (30B A3B) is the cheapest scored model on this hle leaderboard at $0.05/M input and $0.20/M output ($0.25 blended). Claude Fable 5.1 still leads HLE at 65%.

    Should I always pick the #1 HLE model?

    Not automatically. Claude Fable 5.1 leads HLE, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh HLE against input/output price, context window, and related evals.

    How often is the HLE leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled September 2, 2026. Treat it as a current index, not a one-off blog post.

    What is Humanity's Last Exam?

    Humanity's Last Exam, or HLE, is a broad expert-level test from the Center for AI Safety. It is meant to stay hard after models saturate MMLU and GPQA.

    What is a good HLE score?

    Frontier models still sit well below human experts on the full exam. Compare models on this page rather than treating any single percentage as a pass mark.

    Does HLE use tools?

    Labs sometimes report HLE with tools and HLE without tools. We keep both columns when the score exists, plus GPQA as a second reasoning check.