All
Articles 133,316Blog Posts 137,787Tech Tutorials 34,567Research Papers 25,941News 18,856
⚡ AI Lessons

Dev.to · Maya Andersson
🧠 Large Language Models
⚡ AI Lesson
2w ago
Your LLM-as-judge has a position bias you are not measuring
If your pairwise judge sees answer A before answer B, it tends to prefer A. If you never swap the...

Dev.to · Maya Andersson
🧠 Large Language Models
⚡ AI Lesson
3w ago
We fixed the worst prompt variant. It got better. That doesn't mean the fix worked.
A pattern I've seen on more than one team: weekly eval run finishes, someone sorts the leaderboard,...

Dev.to · Maya Andersson
🧠 Large Language Models
⚡ AI Lesson
4w ago
I checked six LLM-as-judge tools against human labels. The scoreboard was the wrong thing to read.
Most LLM-as-judge comparisons rank tools by which one gives you a number fastest. That is the wrong...

Dev.to · Maya Andersson
🧠 Large Language Models
⚡ AI Lesson
1mo ago
We put confidence intervals on our LLM-judge scores. The error bars ate three weeks of "trend"
We track weekly agreement between an LLM judge and human labels (Cohen's kappa) on a sample of...

DeepCamp AI