You're reading one of the top pains on OpenAI Platform & APIs — a free sample of the catalog. Create a free account to browse all Shopify & Atlassian pains and the basic Chrome catalog.
LLM developers can't trust LLM-as-judge scores due to self-disagreement and drift
.21
Demand — how loud & how many (score)
Willing to payImplicit
Momentum (7d)▼ 68.8%
Cluster12 posts · 11 people
MarketplaceOpenAI Platform & APIs
Updated1d ago
The story
The pain
We rely on LLM judges to evaluate our agents, but the judges disagree with themselves—scoring the same case 5/5, then 2/5, then 5/5 within minutes. A single bad run is not evidence, and the live agreement rate stays pretty even when the judge and traffic drift together. We catch it downstream, not from the score. We need a frozen, hand-labelled canary set, a pinned judge version, and enough sample size to distinguish noise from real degradation.
Evidence
What people said · 12 complaints, 11 people
An LLM judge disagrees with itself. I’ve had the same case, same everything, score 5/5, then 2/5, then 5/5 within fifteen minutes. So a single bad run is not evidence: before I look at execution I run the case a few times against the approved version and see how much it moves on its own.
You don't read drift off the live agreement rate. Live traffic and the judge can move together, so the number stays pretty. Such-Process5697's story is the usual one: you catch it downstream.
What bit me was writing my own label set and then tuning the judge against it. After that the agreement rate stays high and means nothing, you optimised toward your own ruler. Freeze a small hand labelled set before you touch the judge and never re-label it when the model 'improves', that's the only part that stays honest.
Export the pain — problem, evidence and market signals — as Markdown, PDF or JSON. The Markdown export ships with a built-in briefing you can hand straight to a developer or an AI agent.