top of page

Anthropic's Mid-Tier Model Just Beat OpenAI's Flagship (and Everyone Else) at Expert Judgment

  • Jul 18
  • 3 min read

Fable 5 was supposed to be the story.

But when we released our latest leaderboard scores, the real headline turned out to be that Anthropic's mid-tier model beat OpenAI's best one.

Sonnet 5 scored 74.3% on expert alignment. Sol, OpenAI's flagship GPT-5.6 model, scored 72.9%. It's a clean win, and one not many expected.

What is Expert Alignment?


An expert alignment benchmark measures how closely an AI model’s answers match what a qualified professional would actually say on expert-authored questions. For buyers, developers, and professionals using AI in high-stakes knowledge work, from legal and health tasks to other regulated workflows, that is often a more useful evaluation than public leaderboard scores. We built our benchmark from expert-authored questions the models have never seen, translating complex industry standards into measurable grading rubrics rather than testing for generic benchmark performance. We score on correctness, completeness, prioritization, and judgment. That evaluation combines human annotators with automated scoring to produce a more nuanced review of whether models are correct, complete, and practically useful.



Here are the latest additions to our leaderboard:

 

  • Fable 5 (Anthropic): 80.0%

     

  • Sonnet 5 (Anthropic): 74.3%


  • Sol (OpenAI): 72.9%


  • Terra (OpenAI): 71.8%


  • Luna (OpenAI): 69.2%

 

Fable 5 leads by a wide margin: 5.7 points clear of Sonnet 5, and more than 7 points clear of Sol. In a field this tight, that’s a gap worth taking seriously. The rest of this article compares Anthropic and OpenAI across this benchmark, explains why expert judgment matters more than headline-friendly public tests, and shows how to use expert alignment scores when choosing models for work where accuracy and reliability matter.

Two Companies, Two Very Different Pitches


OpenAI is selling a ladder. Sol at the top for frontier work, Terra in the middle for "everyday work," Luna at the bottom for speed and cost. It's a portfolio built around performance per dollar and agentic flexibility, and outlets like Axios and Barron's have largely repeated that framing back to readers.


Anthropic is selling something narrower: judgment. Fable 5 for the hardest, highest-stakes knowledge work. Sonnet 5 for scaling that same judgment across more volume at a lower price.


Our numbers back Anthropic's pitch almost exactly. That's the uncomfortable part for OpenAI. Its three models cluster tightly together (69.2% to 72.9%), which suggests reasoning mode and price tier aren't buying much expert alignment. Meanwhile Anthropic's two models land well above and set the pace.



Why This Matters More Than It Looks


Public benchmarks are getting easier to game, so evaluation strategies need to adapt as models scale and mature. They leak into training data, they get optimized for headlines, and they rarely test whether a model behaves like a professional under real conditions. We built this benchmark to close that gap: private questions, expert grading, no leaderboard-chasing. Benchmark design often moves through five stages from proof of concept to continuous evolution, with simple metrics giving way to more custom tasks and continuous evaluation as a key way to keep benchmarks aligned with evolving best practices.


Which means the practical question for anyone buying model access is "which one gets my domain right when it matters," not "which one is fastest or cheapest."


Your Buying Guide


Fable 5: When the answer is the product. Legal intake, including benchmark task work where models are evaluated on interpreting statutes and analyzing precedents, health guidance, regulated workflows, anything where "mostly right" isn't good enough.


Sonnet 5: When you need that same judgment at scale, without Fable's price tag. It beat every OpenAI model tested here.


Sol: Strong if you care about agentic workflows, coding, and cost-efficiency, but don't assume expert-grade output without testing it in your own vertical first.


Terra: Close to Sol on our score, so worth a look for scaled production if your domain testing holds up.


Luna: Fast and cheap, good for drafting and routing, not built for high-stakes answers without a human checking behind it.



The Bottom Line


OpenAI built a portfolio around speed, cost, and flexibility. Anthropic built one around judgment. Our benchmark says both companies are telling the truth about what they built. Expert alignment is a separate axis from cheap and fast, and buyers should treat it that way. Right now, on this test, Anthropic's models own that category outright.


 
 
 

Comments


Start using our API solution

bottom of page