All
Articles 138,257Blog Posts 141,986Tech Tutorials 35,879Research Papers 27,167News 19,394
⚡ AI Lessons

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
3w ago
Your LLM-as-judge disagrees with itself between runs
Same outputs, same judge, two runs, two scores. The gate flickered red then green on a branch with...

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
3w ago
LLM-as-judge disagrees with itself between runs
The flap I had a faithfulness gate on merge: judge scores every case, the mean has to clear 0.80....

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
3w ago
When an LLM answer is wrong, the trace is where you look. Some tools make that easy.
A user reports a hallucinated answer in prod. To fix it you need the full trace of that one request,...

Dev.to · Ethan Walker
☁️ DevOps & Cloud
⚡ AI Lesson
4w ago
our CI passed. Your agent isn't operator-ready.
Your CI passed. Your agent isn't operator-ready. We shipped a document-extraction agent to...

Dev.to · Ethan Walker
📐 ML Fundamentals
⚡ AI Lesson
1mo ago
91% pass rate. Gate green. Shipped. Worst regression we had all quarter.
The gate was a fixed 90% threshold on an intent-classification eval. The change came in at 91%,...

Dev.to · Ethan Walker
⚡ AI Lesson
2mo ago
Promptfoo is a CI gate, not an eval framework. Treating it like one cost us $4,200
Last Monday I logged into our billing dashboard and saw a $4,200 LangSmith spike from the weekend....
DeepCamp AI