/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

930 articles tagged with Benchmark

Latest Trending
Mastodon discussion Jul 21

Radar's open source, single Go binary, MCP server built in.https://radarhq.io/blog/radar-mcp-vs-kubectl-benchmark#Kubern...

Radar's open source, single Go binary, MCP server built in.https://radarhq.io/blog/radar-mcp-vs-kubectl-benchmark#Kubernetes #AI #MCP #DevOps #OpenSource

Open Source Benchmark MCP
9
Mastodon discussion Jul 21

📊 DBRX Instruct — the actual numbers GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3%Measured ...

📊 DBRX Instruct — the actual numbers GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3%Measured independently, not self-reported →https://opensourceai.tech/...

Benchmark
9
Dev.to tutorial Jul 21

Your A/B eval is paired. Your stat test probably isn't.

Two prompts on one eval set are paired data. The two-proportion SE said 'collect more'; McNemar said it was already decided.

Benchmark
20
Papers with Code paper Jul 21

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-ver...

Benchmark
21
Mastodon discussion Jul 20

Kimi K3 (Moonshot AI) beat GPT-5.6 and Fable 5 on the front-end coding benchmark, 76% vs 63%. Open weights, 2.8 trillion...

Kimi K3 (Moonshot AI) beat GPT-5.6 and Fable 5 on the front-end coding benchmark, 76% vs 63%. Open weights, 2.8 trillion parameters.It's also half the price per token, but needs ro...

OpenAI Anthropic Benchmark
18
Mastodon discussion Jul 20

📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28%...

📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28%⚡ 49.2 tokens/sec💰 15.6 intelligence points per dollarMeasur...

Mistral Benchmark
9
Mastodon discussion Jul 20

Why has Austria become Europe's hottest AI market this year?Austria’s benchmark ATX index has gained 21.3% since the beg...

Why has Austria become Europe's hottest AI market this year?Austria’s benchmark ATX index has gained 21.3% since the beginning of January.The biggest driver of that performance is ...

Benchmark
24
Mastodon discussion Jul 20

LoRA Speedrun - a public wall-clock leaderboard for fine-tuning techniqueshttps://github.com/Saivineeth147/lora-speedrun...

LoRA Speedrun - a public wall-clock leaderboard for fine-tuning techniqueshttps://github.com/Saivineeth147/lora-speedrun#HackerNews #Tech #AI

Fine-Tuning Benchmark
24
Papers with Code paper Jul 20

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu...

Anthropic Agents Benchmark
21
Papers with Code paper Jul 20

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. Th...

Anthropic Benchmark
21
Mastodon discussion Jul 19

Perplexity has released WANDR, an open benchmark with 500 evidence-heavy tasks for testing research agents. The benchmar...

Perplexity has released WANDR, an open benchmark with 500 evidence-heavy tasks for testing research agents. The benchmark evaluates whether agents can discover many qualifying enti...

Benchmark
9
Dev.to tutorial Jul 19

Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

There's a slide in every "LLM eval platform" pitch deck that says: score every response, catch every...

Benchmark
12
Mastodon discussion Jul 18

AI agent benchmark tests tool-switching when reliability shiftsA new arXiv paper uses cognitive psychology's set-shiftin...

AI agent benchmark tests tool-switching when reliability shiftsA new arXiv paper uses cognitive psychology's set-shifting concept to test whether LLM agents adapt when reliable too...

Agents Benchmark
9
Dev.to tutorial Jul 18

Try to Break Our AI Memory Benchmark

Facts change. An earnings forecast is revised. A policy is amended. A medication dose is corrected....

Benchmark
12
Mastodon discussion Jul 18

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #...

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face LLM Benchmark
9
Mastodon discussion Jul 18

China's 2.8-trillion-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena benchmark— Moonshot AI delivers large...

China's 2.8-trillion-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena benchmark— Moonshot AI delivers largest open-weight AI model ever, as Ch…Beijing-based Moonshot A...

Anthropic Google Benchmark
18
Mastodon discussion Jul 17

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し...

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face Benchmark
9
Mastodon discussion Jul 17

SWE-bench needs 90% of tasks for reliable agent benchmark resultsA replay analysis of three LLM agent benchmarks finds t...

SWE-bench needs 90% of tasks for reliable agent benchmark resultsA replay analysis of three LLM agent benchmarks finds the safe partial-run fraction ranges from 15% to over 95%, wi...

LLM Benchmark
9
Mastodon discussion Jul 17

LLM agent memory poisoning evades defenses, 1,227-case study findsMemPoison benchmark tests 1,227 attacks across 10 mode...

LLM agent memory poisoning evades defenses, 1,227-case study findsMemPoison benchmark tests 1,227 attacks across 10 model families, finding write-time defenses miss sophisticated m...

LLM Benchmark
9
Mastodon discussion Jul 17

LLM agents spot 88% of supply-chain failures but can't actNew STOCKTAKE benchmark: LLM agents detect up to 88% of hidden...

LLM agents spot 88% of supply-chain failures but can't actNew STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a ...

LLM Benchmark
9
Mastodon discussion Jul 17

📊 K-EXAONE (Reasoning) — the actual numbers GPQA: 78.3% MMLU-Pro: 83.8% Humanity's Last Exam: 13.1% Long Context Reasoni...

📊 K-EXAONE (Reasoning) — the actual numbers GPQA: 78.3% MMLU-Pro: 83.8% Humanity's Last Exam: 13.1% Long Context Reasoning: 55.7%Measured independently, not self-reported →https://...

Benchmark
9
Mastodon discussion Jul 17

This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard...

This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization is the new battlegr...

NVIDIA Benchmark
9
Dev.to tutorial Jul 17

If 30% of Coding Tasks May Be Broken, Your Leaderboard Needs an Uncertainty Budget

Turn a coding-benchmark audit into a reproducible task-state model, sensitivity analysis, and acceptance rule.

Benchmark
12
Mastodon discussion Jul 17

🧠 Researchers introduce the Wandr Benchmark, a tool for evaluating AI agents that perform web search and information gat...

🧠 Researchers introduce the Wandr Benchmark, a tool for evaluating AI agents that perform web search and information gathering tasks. The benchmark measures how well these agents c...

Benchmark
9
« Previous Page 10 of 39 (930 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available