TL;DR: Our LLM-as-judge agreement (Cohen's kappa against human labels) swung between 0.41 and 0.63...
More eval traces will not stabilize your kappa. Stratify the ones you have
TL;DR: Our LLM-as-judge agreement (Cohen's kappa against human labels) swung between 0.41 and 0.63...
We started building OliverGraph to give teams and their AI agents shared context across GitHub,...
This is a submission for Weekend Challenge: Dog Days Edition What I Built PawSafe is an...
A working supervisor/executor split for Claude Code, one command for the real subscription and one for a proxy-routed cheap tier, and why they're deliberately asymmetric.