RewardBench v1's top-6 spread is 5.7 points — small specialist models now dominate it. RewardBench 2 drops scores by 20 points and actually correlates with downstream RLHF. RM-Bench finds that style bias can push SOTA models below random performance.
Chatbot Arena ranks LLMs with ELO, InstructGPT trains a reward model on pairwise preferences, and chess has rated players for seventy years. All three rest on the same one-line probabilistic model — Bradley–Terry — which turns out to be logistic regression over comparisons.
Chatbot Arena ranks LLMs with ELO, InstructGPT trains a reward model on pairwise preferences, and chess has rated players for seventy years. All three rest on the same one-line probabilistic model — Bradley–Terry — which turns out to be logistic regression over comparisons.