A few weeks ago I wrote about building a reproducible test harness for comparing free AI coding...
The Model Passed Your Benchmark. Now Stop Merging Its Code Blindly
A few weeks ago I wrote about building a reproducible test harness for comparing free AI coding...
The other day I wrote about the npm worm that learned to trust your AI agent — a keyv-adjacent...
My last post ended on a result I liked far too much. I run free leaderboards of which products AI...
An independent 64-run Claude Code benchmark of transactional email APIs across four TypeScript SaaS implementation tasks.