How do you build an expert-grade research QA benchmark without hand-authoring a single question?Mine the survey articles. Their section structure already encodes what a field asks and what a complete answer must contain, so a new benchmark distills 21K queries and grading rubrics from surveys across 75 fields. It doubles as a stress test: even the best system tops out at 75% rubric coverage, fully addressing under 11% of needed citations.https://benjaminhan.net/posts/20260625-researchqa/?utm_source=mastodon&utm_medium=social#LLMs #Evaluation #AI
Related
heretic: Tools for removing censorship in open weight LLMshttps://github.com/p-e-w/heretic #transformer #censorship #her...
heretic: Tools for removing censorship in open weight LLMshttps://github.com/p-e-w/heretic #transformer #censorship #heretic #llm #ai #+
Intel says PC market is ‘a tale of two kingdoms’ with mainstream ‘taking a beating’ — VP suggests a split between mainst...
Intel says PC market is ‘a tale of two kingdoms’ with mainstream ‘taking a beating’ — VP suggests a split between mainstream and enthusiast sockets across the industryIntel VP Robe...
📰 AI Fails to Deliver a 4-Day Work Week - Especially at AI Companies"An engineering director at Google said four years a...
📰 AI Fails to Deliver a 4-Day Work Week - Especially at AI Companies"An engineering director at Google said four years ago that AI would deliver a four-day work week by 2025," reme...