Perplexity has released WANDR, an open benchmark with 500 evidence-heavy tasks for testing research agents. The benchmark evaluates whether agents can discover many qualifying entities and back each with cited evidence. Perplexity Search as Code leads at 0.363 soft F1. https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep/ #AIagent #AI #GenAI #AIResearch
Related
AI NPCs could make games more dynamic... but more dialogue doesn't automatically mean better writing. Players aren't too...
AI NPCs could make games more dynamic... but more dialogue doesn't automatically mean better writing. Players aren't too dumb to notice when human creativity gets replaced by short...
Everything is about to "go dark"Article URL: https://blog.cryptographyengineering.com/2026/08/14/everything-is-about-to-...
Everything is about to "go dark"Article URL: https://blog.cryptographyengineering.com/2026/08/14/everything-is-about-to-go-dark/ Comments URL: https://news.ycombinator.com/item?id=...
A YOLO- and CLIP-based vision-language framework classifies mosquito flight frames of uninfected and Dengue virus seroty...
A YOLO- and CLIP-based vision-language framework classifies mosquito flight frames of uninfected and Dengue virus serotype 2-infected mosquitoes with 98.54% accuracy and 99.91% sen...