We built a benchmark, then caught it strangling the models it was grading

A couple of day ago I posted about OmnisBench, our open benchmark for LLM routing, specifically our...

Read Original

Related