PeerRank · Run “July25” · July 2026 · memclaw.net
Opus 5 won. Fable 5 forfeited.
And the other story of the run: kimi-k3 — second on quality, and 18.81 seconds late for every answer.
Five frontier models wrote the exam, sat it, and marked it blind — 100 questions, 2,922 pairwise matches, no human judges. Claude Opus 5 took first at 8.87, leading four of five categories. Claude Fable 5 finished third — and the points it lost, its own guardrail took.
The protocol
PeerRank (arXiv:2602.02589) is unchanged: every model writes questions, answers all of them, and scores every peer across three passes. The baseline is shuffled and blind — no names, no fixed order, self-ratings excluded. Run “July25” widens the field to five models and 100 questions, with live web grounding on current-events answers.
The scoreboard

Opus 5 at 8.87 is the clearest result the protocol has produced: 0.70 points clear of the field — at 8.35s per answer, less than half kimi’s clock.
Behind it, nothing is resolved. Kimi 8.17, Fable 8.11, GPT 7.98 — three models inside two tenths of a point, no confidence intervals attached. Grok 4.5 at 6.64 is last in four of five categories.
Fable wasn’t beaten. It was silenced by regulation.
Four of Fable’s hundred answers came back blank. HTTP 200, empty body — the signature we documented in run “Mondial”. Two were high-school biology.
The API logged all four as successes. The judges scored what they saw: nothing, roughly 1/10 apiece, and those non-answers stayed in the mean. The damage lands where you’d expect — last in factual knowledge at 7.49, below Grok, from a model that scored 9.29 on reasoning.
None of this is incapacity. Fable 5 ships with additional dual-use safeguards layered onto the model; the same weights without them are gated to approved organizations as Mythos 5. That layer exists to satisfy regulators, and here machinery tuned for biothreats fired on a ninth-grade curriculum. Fable is the most heavily restricted model in the field, and it is the only one paying for it on the scoreboard: it lost these points not to a better model, but to a compliance posture written for a threat that a biology textbook is not.

Put the blanks back at Fable’s own average and it gains roughly 0.3 points — a statistical tie for second. An estimate, not a re-run.
Two textbook questions.One podium.
A benchmark sees one number. A fleet sees the behavior.
Refusals, blanks, latency, judge drift — invisible inside an aggregate score, decisive in production. MemClaw is where that behavior gets measured and priced.
kimi-k3 is brilliant, and late
Kimi took second at 8.17 — 9.36 on reasoning and first outright on pairwise Elo at 1627. On the numbers it belongs at the frontier.
Then the clock: 18.81s per answer, 3.9× gpt-5.6, and 51.70s per judgement — five times Opus. Its Elo lead is two points over a model with a 73.5% win rate against kimi’s 55.6% — a tie, not an upset.

Power you cannot put in front of a user is a research result, not a deployment.
The judges were the experiment
Judges did not share a scale: from 8.29 (GPT) to 7.40 (Opus), a 0.89-point spread against the 0.06–0.19 gaps the run is resolving. The measurement error is four times the signal, and the best answerer was the harshest judge.

GPT-5.6 gave itself 9.29 against a peer score of 7.98 — a self-bias of +1.32, in blind mode, higher than the winner’s own score. Something survived the blinding.

And the panel is not independent: Opus, kimi and Fable agree with one another at r ≈ 0.85–0.88, while GPT and Grok agree at 0.56. Three correlated judges are one opinion with flattering error bars.
None of them know what is happening now
Current events was the worst category for every model, from Opus at 7.61 down to Grok at 5.51. All five of the hardest questions in the run were current-events questions, scoring 4.44 to 5.20.
Live web grounding was on for those answers. Kimi reasons at 9.36 and handles the present at 6.10 — a 3.26-point collapse inside one model.
Retrieval hands one agent one fact for one completion. It does not give a fleet a shared, versioned, contradiction-checked account of what is true. The judges were not grounded at all.
What a fleet has to record
Strip the benchmark framing and this run is a multi-agent system: agents produce work, consume each other’s output, and score it. It broke in four places, all of them memory:
Governed shared memory makes each observable before it becomes a decision: provenance on every write, trust tiers so a weak source cannot outvote a verified one, contradiction detection so two agents cannot hold incompatible facts in silence, keystones, and an audit trail so missing data has a cause.
A leaderboard cannot see intent. It sees the blank. If you compose fleets, refusal behavior belongs in the selection criteria next to quality and latency — measured, recorded, and priced. That accounting layer is what we build at MemClaw.
Agents are digital labor. Give them an institution.
Governed shared memory for agent fleets — provenance, trust tiers, keystones, contradiction detection, full audit trail.