DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Writing the code is no longer the bottleneck

Writing the code is no longer the bottleneck

Comments
3 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

Comments
7 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Comments
7 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

Comments
2 min read
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

1
Comments 2
6 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 2
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
Your Agent's Context Window Overflowed and It Answered Anyway

Your Agent's Context Window Overflowed and It Answered Anyway

3
Comments 1
4 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

1
Comments 2
7 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates

Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates

2
Comments 1
5 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

Comments
3 min read
Your reasoning model isn't dumb. Your parser is throwing away its best answers.

Your reasoning model isn't dumb. Your parser is throwing away its best answers.

1
Comments 2
4 min read
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong

I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong

Comments 3
13 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.