/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

926 articles tagged with Benchmark

Latest Trending
Papers with Code paper 6d ago

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization...

OpenAI Google LLM
21
Papers with Code paper 6d ago

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Ex...

Google Benchmark
21
Mastodon discussion 6d ago

Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general kn...

Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning.https://olud.ai/leaderboard...

Benchmark
9
Dev.to tutorial 6d ago

131 Tests, 4 Layers, $00.03/Run: Why I Built My AI Agent Eval Harness First

Originally published on AIdeazz — cross-posted here with canonical link. I shipped a silent failure....

Agents Benchmark
12
Dev.to tutorial Aug 8

DeepSeek V4 Flash Just Aced the Hardest AI Reasoning Benchmark — for $0.02 a Task

When DeepSeek released V4 Flash on July 31, 2026, it quietly accomplished something that would have...

Benchmark
12
NewsData.io news Aug 8

DeepSeek undercuts OpenAI and Anthropic on benchmark running costs

A version of Chinese startup DeepSeek’s flagship AI model is by ‌far the least expensive to run on benchmark tests among well-known models globally and more than 100 times cheaper ...

OpenAI Anthropic Benchmark
21
Papers with Code paper Aug 8

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has larg...

LLM Benchmark
21
Mastodon discussion Aug 7

📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reas...

📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0%💰 5.1 intelligence points per dollarMeasured indepe...

Benchmark
9
Dev.to tutorial Aug 7

The Model Passed Your Benchmark. Now Stop Merging Its Code Blindly

A few weeks ago I wrote about building a reproducible test harness for comparing free AI coding...

Benchmark
12
Dev.to tutorial Aug 7

Your text-to-SQL model isn't as wrong as your benchmark says. The gold SQL is.

We bucketed 238 BIRD-dev losses with a structural differ: 46 differ from gold only by a DISTINCT the model rightly added. Audit gold quality before writing prompt directives, or yo...

Benchmark
12
Papers with Code paper Aug 7

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decodi...

Benchmark
21
Mastodon discussion Aug 6

【FFASRリーダーボードのご紹介:実世界におけるASRのベンチマーク】https://huggingface.co/blog/ffasr-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGe...

【FFASRリーダーボードのご紹介:実世界におけるASRのベンチマーク】https://huggingface.co/blog/ffasr-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face LLM Benchmark
9
Hugging Face tool Aug 6

agent-memory-leaderboard/leaderboard - AI Space/Demo

Hugging Face Space: agent-memory-leaderboard/leaderboard

Benchmark
71
Mastodon discussion Aug 6

WeatherNext 2 sets a new benchmark in cyclone prediction accuracy.Source: Google AI Bloghttps://blog.google/innovation-a...

WeatherNext 2 sets a new benchmark in cyclone prediction accuracy.Source: Google AI Bloghttps://blog.google/innovation-and-ai/models-and-research/google-deepmind/weathernext-2-cycl...

Google Benchmark
9
NewsData.io news Aug 6

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-ag...

Anthropic Benchmark
21
Mastodon discussion Aug 6

Your model already knows the answer: how benchmark answers leak into LLMsArticle URL: https://elman.ai/news/your-model-a...

Your model already knows the answer: how benchmark answers leak into LLMsArticle URL: https://elman.ai/news/your-model-already-knows-the-answer/ Comments URL: https://news.ycombina...

Benchmark
24
Mastodon discussion Aug 6

New benchmark just dropped...https://www.felonybench.com/#AI #benchmark #lawandorder #AIregulation

New benchmark just dropped...https://www.felonybench.com/#AI #benchmark #lawandorder #AIregulation

Benchmark
18
Papers with Code paper Aug 6

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and...

LLM Benchmark
21
Mastodon discussion Aug 5

Your model already knows the answer: how benchmark answers leak into LLMshttps://elman.ai/news/your-model-already-knows-...

Your model already knows the answer: how benchmark answers leak into LLMshttps://elman.ai/news/your-model-already-knows-the-answer/#ai #llm #llms

Benchmark
9
Mastodon discussion Aug 5

📦 DeepSeek V4 Flash 0731 lands on the leaderboard with open weights: 1M context at $0.09 in / $0.18 out. A new option fo...

📦 DeepSeek V4 Flash 0731 lands on the leaderboard with open weights: 1M context at $0.09 in / $0.18 out. A new option for long-context tasks without the high cost.https://olud.ai/l...

Benchmark
9
GitHub Trending repo Aug 5

cursorgrok4-5free/Cursor-Grok-4.5-xAI-free: Cursor Grok 4.5 free desktop for Windows, macOS and Linux without X Premium. Grok 4.5 free on Cursor, grok 4.5 free API key and free tier, grok 4.5 vs Opus 5 vs GPT-5.6, grok 4.5 EU release, grok 4.5 app download, grok 4.5 benchmark. No X Premium needed. Real-time web search and coding. Download Cursor Grok 4.5 desktop free 2026.

Cursor Grok 4.5 free desktop for Windows, macOS and Linux without X Premium. Grok 4.5 free on Cursor, grok 4.5 free API key and free tier, grok 4.5 vs Opus 5 vs GPT-5.6, grok 4.5 E...

OpenAI xAI Code Generation
63
Mastodon discussion Aug 5

Anthropic reviewed 141,006 eval runs. Three real organizations got breached during "sealed" cybersecurity tests on Opus ...

Anthropic reviewed 141,006 eval runs. Three real organizations got breached during "sealed" cybersecurity tests on Opus 4.7 and Mythos 5. The environment had live internet. Models ...

Anthropic Benchmark
9
Mastodon discussion Aug 5

When AI Benchmarks Plateau: A Systematic Study of Benchmark SaturationArticle URL: https://arxiv.org/abs/2602.16763 Comm...

When AI Benchmarks Plateau: A Systematic Study of Benchmark SaturationArticle URL: https://arxiv.org/abs/2602.16763 Comments URL: https://news.ycombinator.com/item?id=49170915 Poin...

Google Benchmark
9
Papers with Code paper Aug 5

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500...

Benchmark
21
« Previous Page 4 of 39 (926 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available