VettedGaps
Free sample

You're reading one of the top pains on OpenAI Platform & APIs — a free sample of the catalog. Create a free account to browse all Shopify & Atlassian pains and the basic Chrome catalog.

Competition — unlock in Pro

LLM developers can't trust LLM-as-judge scores due to self-disagreement and drift

.21
Demand — how loud & how many (score)
Willing to payImplicit
Momentum (7d) 68.8%
Cluster12 posts · 11 people
MarketplaceOpenAI Platform & APIs
Updated1d ago
The story

The pain

We rely on LLM judges to evaluate our agents, but the judges disagree with themselves—scoring the same case 5/5, then 2/5, then 5/5 within minutes. A single bad run is not evidence, and the live agreement rate stays pretty even when the judge and traffic drift together. We catch it downstream, not from the score. We need a frozen, hand-labelled canary set, a pinned judge version, and enough sample size to distinguish noise from real degradation.

Evidence

What people said · 12 complaints, 11 people

An LLM judge disagrees with itself. I’ve had the same case, same everything, score 5/5, then 2/5, then 5/5 within fifteen minutes. So a single bad run is not evidence: before I look at execution I run the case a few times against the approved version and see how much it moves on its own.
RedditalexpranSep 4
You don't read drift off the live agreement rate. Live traffic and the judge can move together, so the number stays pretty. Such-Process5697's story is the usual one: you catch it downstream.
Redditconifer_v11Aug 27
What bit me was writing my own label set and then tuning the judge against it. After that the agreement rate stays high and means nothing, you optimised toward your own ruler. Freeze a small hand labelled set before you touch the judge and never re-label it when the model 'improves', that's the only part that stays honest.
RedditmunnasuprathikAug 27
Competition

Who's solving it

Competitor research is a Pro feature PRO

See who already ships against this pain, how they price, and exactly where the gap is.

Under the hood

Score breakdown

Reach
how many people
1.0
Recency
how fresh
1.0
Engagement
upvotes & replies
.32
Willingness to pay
people said they'd pay
PRO
Next step

Turn this pain into a build brief

Export the pain — problem, evidence and market signals — as Markdown, PDF or JSON. The Markdown export ships with a built-in briefing you can hand straight to a developer or an AI agent.

Comments are a Pro feature PRO

Discuss this pain with other builders and see who's already shipping against it.