ZeroShotMind

Series

LLM Evaluation & Benchmarks

A frontier-level survey of how language models are measured — instruction-following, chat quality, creative writing, formatting, and reward models. Per benchmark: leaderboard data, saturation analysis, and what the numbers actually mean.

Defining Frontier
llm-evaluationbenchmarksleaderboardssaturation
  1. 1

    Instruction-Following Benchmarks: IFEval Is Saturated, IFBench Exposes What's Left

    IFEval's top-10 spread is 2.9pp — six of the top 12 spots go to Qwen3.5 variants. IFBench then shows a 29pp gap between reasoning models and standard instruction-tuned ones on constraints none of them trained against.

    2026-07-17

  2. 2

    Chat & Arena Benchmarks: Human Votes Scale, MT-Bench Doesn't

    MT-Bench's frontier models cluster above 9.0 — the scale is out of room. AlpacaEval 2.0's length-controlled win rate is now saturated above 95% for top models. LMArena Elo keeps separating models as long as votes keep coming in — and they do.

    2026-07-17

  3. 3

    Creative Writing Benchmarks: When the Rubric Runs Out of Signal

    EQ-Bench CW v3 rubric scores are already saturated at the top — a 0.35-point spread across 10 models. Elo still discriminates. Here's what that gap reveals about how we evaluate creative writing.

    2026-07-17

  4. 4

    Formatting & Output Validation Benchmarks: IFEval Is Saturated, IFBench Isn't

    IFEval's top-5 span fewer than 2 percentage points — frontier models have converged on its constraint set. IFBench exposes a 29pp gap between Grok and Claude on out-of-distribution constraints. And SOB shows that JSON schema compliance is not the same as correct field values.

    2026-07-17

  5. 5

    Reward Model Benchmarks: RewardBench Is Saturated, RewardBench 2 Isn't

    RewardBench v1's top-6 spread is 5.7 points — small specialist models now dominate it. RewardBench 2 drops scores by 20 points and actually correlates with downstream RLHF. RM-Bench finds that style bias can push SOTA models below random performance.

    2026-07-17