ArXiv paper Mar 26

RenoBench: A Citation Parsing Benchmark

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often ...

Papers with Code paper Mar 26

Robust Reasoning Benchmark

While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their underlying reasoning processes remain highly overfit to standard textual form...

ArXiv paper Mar 24

Code Review Agent Benchmark

Software engineering agents have shown significant promise in writing code. As AI agents permeate code writing, and generate huge volumes of code automatically -- the matter of cod...