I've been running a small agentic eval harness against a local model and I'd like a sanity check on...
Agentic tool-use eval on a local 35B (Q8): trap-tool avoidance is solid, but I can't tell if my failures are the model or my harness
I've been running a small agentic eval harness against a local model and I'd like a sanity check on...
A retrieval pipeline built by hand in six files — the data structures behind each stage, the four bugs that cost me the most time, and the answer it made up that I left in the repo...
I built and open-sourced PacketVoyage—an Agent Skill & MCP server that turns boring traceroute...
Introduction For the past 10 days, I took part in the 10 Days of AI Voice Agents — #VoiceForBharat...