I run an autonomous eval agent against new coding-agent stacks before trusting their numbers. The...
A Postmortem on Autonomous LLM-as-Judge: How My Eval Agent Got Two Verdicts Wrong Before I Found a Sandbox Bug
I run an autonomous eval agent against new coding-agent stacks before trusting their numbers. The...
Self-hosting Kimi K3 became technically possible on 27 July 2026, when Moonshot AI published the...
LLMKube 0.9.19 shipped this morning, and it is the strangest release we have cut. About half of it...
Google’s Search Central Live Toronto event offers a credible signal that structured data is part of...