Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Writing the code is no longer the bottleneck
The Engineering Manager’s Desk
The Engineering Manager’s Desk
The Engineering Manager’s Desk
Follow
Aug 15
Writing the code is no longer the bottleneck
#
engineeringculture
#
evaluation
#
ai
#
techdebt
Comments
Add Comment
3 min read
Why AI Benchmarks Mean Less Than You Think
The AI Downside
The AI Downside
The AI Downside
Follow
Aug 15
Why AI Benchmarks Mean Less Than You Think
#
benchmarks
#
llms
#
evaluation
#
hype
Comments
Add Comment
6 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
Guatu
Guatu
Guatu
Follow
Aug 14
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
#
aiagents
#
knowledgegraph
#
evaluation
#
rag
Comments
Add Comment
7 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
Comments
Add Comment
7 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
Adha AK
Adha AK
Adha AK
Follow
Aug 4
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
#
llm
#
evaluation
#
testing
#
opensource
Comments
Add Comment
2 min read
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 12
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
#
ai
#
evaluation
#
dotnet
#
testing
1
 reaction
Comments
2
 comments
6 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
#
rag
#
llm
#
embeddings
#
evaluation
Comments
2
 comments
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format
Adha AK
Adha AK
Adha AK
Follow
Jul 30
OpenEval: Why LLM Evaluation Needs a Standard Format
#
llm
#
evaluation
#
ai
#
testing
Comments
Add Comment
1 min read
Your Agent's Context Window Overflowed and It Answered Anyway
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 12
Your Agent's Context Window Overflowed and It Answered Anyway
#
ai
#
agents
#
observability
#
evaluation
3
 reactions
Comments
1
 comment
4 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
1
 reaction
Comments
2
 comments
7 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
 reaction
Comments
Add Comment
6 min read
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 9
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates
#
ai
#
agents
#
evaluation
#
observability
2
 reactions
Comments
1
 comment
5 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
#
llms
#
codegeneration
#
selfrepair
#
evaluation
Comments
Add Comment
3 min read
Your reasoning model isn't dumb. Your parser is throwing away its best answers.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 7
Your reasoning model isn't dumb. Your parser is throwing away its best answers.
#
machinelearning
#
llm
#
evaluation
#
ai
1
 reaction
Comments
2
 comments
4 min read
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong
shakti tiwari
shakti tiwari
shakti tiwari
Follow
Aug 5
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong
#
aiagents
#
evaluation
#
localai
#
testing
Comments
3
 comments
13 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account