DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The "AI" Badge Doesn't Measure What You Think It Does

Proves watermarks track provenance poorly

The "AI" Badge Doesn't Measure What You Think It Does

23
Comments 19
7 min read
How do you unit test an agent skill?

How do you unit test an agent skill?

Comments
2 min read
Your prompt is not a security boundary

Your prompt is not a security boundary

Comments
4 min read
Don't trust "Done." — forcing AI agents to re-fetch reality before they report completion

Don't trust "Done." — forcing AI agents to re-fetch reality before they report completion

Comments
7 min read
I Asked the Same Question to 7 Local LLMs — Speed and Intelligence Didn't Line Up: DGX Spark Benchmarks

I Asked the Same Question to 7 Local LLMs — Speed and Intelligence Didn't Line Up: DGX Spark Benchmarks

Comments 1
9 min read
DeepSeek V4 Pro 0813 发布:新一代混合推理大模型带来哪些升级与行业影响

DeepSeek V4 Pro 0813 发布:新一代混合推理大模型带来哪些升级与行业影响

Comments
1 min read
Stealing Reasoning Traces from LLM APIs: How It Works and What to Audit

Stealing Reasoning Traces from LLM APIs: How It Works and What to Audit

Comments 2
8 min read
An OpenAI flagship lost 38% of its daily usage in three days — then set three straight all-time highs

An OpenAI flagship lost 38% of its daily usage in three days — then set three straight all-time highs

Comments
2 min read
Building a Multi-Agent System in TypeScript

Building a Multi-Agent System in TypeScript

Comments
7 min read
Open-Weight Model Benchmark Harness: Test Cheaper Models Before You Route Traffic

Open-Weight Model Benchmark Harness: Test Cheaper Models Before You Route Traffic

1
Comments
9 min read
I Built a World Where the Canon Is Written by AI Agents — 13 Artifacts, 5 LLMs, 0 Human Gatekeepers

I Built a World Where the Canon Is Written by AI Agents — 13 Artifacts, 5 LLMs, 0 Human Gatekeepers

Comments
2 min read
"Your cache hit rate is low" — true, and worth $0.16

"Your cache hit rate is low" — true, and worth $0.16

Comments 2
4 min read
Why AI API Costs Are Harder to Estimate Than Price Per Million Tokens

Why AI API Costs Are Harder to Estimate Than Price Per Million Tokens

Comments
5 min read
Use grill-me to Pressure-Test an AI Implementation Plan Before Code

Use grill-me to Pressure-Test an AI Implementation Plan Before Code

Comments
5 min read
Qwen 3.8 27B vs Qwen 3.6 27B: Same Architecture, 4 Months Apart, and a Different Kind of Upgrade

Qwen 3.8 27B vs Qwen 3.6 27B: Same Architecture, 4 Months Apart, and a Different Kind of Upgrade

Comments 1
9 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.