AgentSearchBench: A Benchmark for AI Agent Search in the Wild
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unli...
972 articles tagged with Benchmark
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unli...
OpenAI’s GPT-5.5 release is a planning signal for enterprise teams, not just a benchmark headline. Here is what to change in governance and rollout operations now. https://go.ainte...
Lei Li (@_TobiasLee)모델 출시가 많은 주간에 Claw-Eval도 업데이트되었으며, MiMo V2.5 Pro가 3위, MiMo V2.5가 5위로 올라섰다고 알린다. 다음 후보로 DeepSeek V4를 언급하며 최신 모델 벤치마크 흐름을 보여준다.https://x.com/_TobiasLee/status/204...
Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and ...
📰 NARS-Reasoning-v0.1: The First Executable Narsese Benchmark for Neuro-Symbolic AI in 2026A new benchmark called NARS-Reasoning-v0.1 translates natural language into executable Na...
Stop evaluating LLMs with vibes. Here's a practical framework for benchmarking open-source models against your API provider using real production data.
Full code, aggregated numbers (n=10 across 5 tasks and 5 transports), and a curated selection of 8...
【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated
At present, executable visual workflows have emerged as a mainstream paradigm in real-world industrial deployments, offering strong reliability and controllability. However, in cur...
PRL-Bench: LLMs Score Below 50% on End-to-End Physics Research TasksResearchers introduced PRL-Bench, a benchmark built from 100 recent Physical Review Letters papers, testing LLMs...
SocialGrid Benchmark Shows LLMs Fail at Deception, Score Below 60% on PlanningResearchers introduced SocialGrid, a multi-agent benchmark inspired by Among Us. It shows state-of-the...
Introducing the Open Ko-LLM Leaderboard: Leading the Korean LLM Evaluation Ecosystem #AINews #Shorts.
Mike Saleme — 2026-04-20 — views my own This week OpenAI released GPT-5.4-Cyber, positioned as the...
BEIJING: Robots outpace humans in half marathon, Bloomberg reports new AI benchmark. China’s robotic surge reshapes urban mobility and beyond.🚩 #Robotik #AI
📰 KWBench 2026: The First Benchmark for AI’s Unprompted Problem Recognition in Knowledge WorkKWBench, a groundbreaking benchmark for unprompted problem recognition in knowledge wor...
📢 Benchmark de LLMs auto-hébergés pour la sécurité offensive : résultats et observations📝 ## 🔍 ContextePublié le 14 avril 2026 sur le blog de TrustedSec par Brandon McGrath, cet ar...
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and t...
Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliability and vulnerabilities to ...
Claude's Opus 4.7 reclaims coding benchmark leadership with 87.6% on SWE-bench Verified, while Anthropic narrows OpenAI's enterprise lead to just 4.6 percentage points. Meanwhile, ...
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during de...
Delve faked compliance certificates for 494 companies. Now agents are faking benchmark scores. Same pattern, new layer. The only thing that catches both is behavioral telemetry.
GeoAgentBench: New Dynamic Benchmark Tests LLM Agents on 117 GIS ToolsA new benchmark, GeoAgentBench, evaluates LLM-based GIS agents in a dynamic sandbox with 117 tools. It introdu...
GPU server rental 2026 — H100 from $1.85/hr, B200 access via CoreWeave. Full benchmark + price comparison: https://server-rental-guide.pages.dev/ #AI #infrastructure