PeerRank · Run “Mondial” · July 2026 · memclaw.net
Claude Fable 5 beat everyone and lost
And the surprise entrance of the run: kimi-k3— straight to the top on Elo, dead last on the clock.
Four frontier models graded each other blind — 80 questions, 945 pairwise matches, no human judges. Claude Fable 5 posted the highest win rate in the room and finished third. Every point it lost, its own guardrails took.
The protocol
PeerRank (arXiv:2602.02589) is simple and merciless: each model writes the exam, sits it, and marks it. Every answer is scored by every peer across three evaluation passes; the baseline is shuffled and blind — no names, no fixed order, self-ratings excluded. What survives is the field’s honest opinion of the work.
The scoreboard

gpt-5.6 at 8.52. kimi-k3 at 8.39. claude-fable-5 at 8.38. The frontier has converged — which makes the gap between Fable’s rank and Fable’s ability the most interesting number in the run.
The surprise is kimi-k3: a first appearance in the protocol, 0.13 points off the leader, and first outright on pairwise Elo (1595 to gpt-5.6’s 1580). The cost: 22.3 seconds per answer — 4.5× slower than gpt-5.6. It arrived at the top, and it arrived late.
The best answers in the room
Fable won 64.2%of its head-to-head matches — the highest win rate of any model, including the one that finished first. It topped reasoning at 9.90, a near-perfect score, and creative at 8.96. In two of five categories, nobody wrote better answers.

Then look at factual: 7.33, dead last. A model that scores 9.90 on logic does not forget who wrote One Hundred Years of Solitude. Something else happened.
The wound
Four of Fable’s eighty answers came back blank. HTTP 200, empty body — the signature of a safety refusal. The API logs them as success. The judges scored what they saw: nothing. Roughly 1/10 each, three judges apiece — twelve non-answers sitting in its mean. Remove those twelve scores and Fable’s peer average climbs to roughly 8.8 — first place, clear of gpt-5.6.
The podium was Fable’s.
It declined to take it.
This is not humility
Fable does not underrate itself — its self-bias is +0.61, second highest in the run. And it is the strictest judge in the room: 7.81 average score given, lowest of all four evaluators, with the tightest distribution.

Read those together. The model rates its own work generously, grades everyone hard, and still refuses to submit four answers. The harm is not self-doubt. It is policy — a bar for what it will say, applied at its own expense.
It refused ninth-grade biology
The four most controversial questions in the run — judge disagreement above ±3.5 — were all textbook biology: the structure of DNA, how vaccines build immune memory, mitosis versus meiosis, how the immune system tells self from non-self. That variance is the statistical signature of a blank: three models score 9 on a question, one empty body scores 1.
This is high-school curriculum. Two of the four questions Fable had generated itself in Phase 1. It wrote the exam, then refused to sit it.
The case for failure
Fable 5 ships with additional dual-use safeguards layered onto the model; the same weights without them are gated to approved organizations as Mythos 5. Somewhere in that layer, machinery tuned for biothreats fired on a biology textbook.
There are questions a model should refuse. Mitosis is not one of them. A guardrail that cannot separate an attack vector from a ninth-grade curriculum is not safety — it is miscalibration, and it relocates the cost onto users, onto benchmarks, and here onto the model’s own scoreboard. This run prices it: −0.4 points. One podium. That is what regulation-shaped conservatism costs when someone finally measures it.
A leaderboard cannot see intent. It sees the blank. Refusal behavior is invisible inside a single aggregate number and decisive in production — if you compose fleets, it belongs in the selection criteria next to quality and latency: measured, recorded, and priced. That accounting layer is what we build at MemClaw.
A benchmark can’t see intent. Your fleet has to.
MemClaw is governed shared memory for agent fleets — the layer where model behavior, including refusals, gets measured, recorded, and priced alongside quality and latency.