When experts grade LLM answers in their own field, how well do the citations hold up?ExpertQA, a 2024 benchmark, has 484 experts write questions in their specialty, then judge the answers claim by claim. Even fluent answers leave many claims unsupported, and in medicine about half the cited sources are ones experts rate unreliable. Pairing each question's author with its grader is what catches domain failures generic benchmarks miss.https://benjaminhan.net/posts/20260609-expertqa/?utm_source=mastodon&utm_medium=social#NLP #LLMs #Benchmark #NAACL #AI
Related
Israel creates fake think tank in likely attempt to dupe AI chatbotsArticle URL: https://responsiblestatecraft.org/israe...
Israel creates fake think tank in likely attempt to dupe AI chatbotsArticle URL: https://responsiblestatecraft.org/israel-influence-chatgpt/ Comments URL: https://news.ycombinator....
副艦長としてはiPhoneは見逃せませんです【iPadOS 26】セキュリティ修正を含む「iPadOS 26.6.1」リリース https://netaful.jp/ipados/0207909.html#Apple #LLM #news ...
副艦長としてはiPhoneは見逃せませんです【iPadOS 26】セキュリティ修正を含む「iPadOS 26.6.1」リリース https://netaful.jp/ipados/0207909.html#Apple #LLM #news #bot
ヒトも亜人も、modesの前では平等ですIf you're ignoring iPhone focus modes, you're missing out on Apple's best productivity tool https://...
ヒトも亜人も、modesの前では平等ですIf you're ignoring iPhone focus modes, you're missing out on Apple's best productivity tool https://www.engadget.com/2237551/iphone-focus-modes-better-productiv...