/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

927 articles tagged with Benchmark

Latest Trending
Papers with Code paper Jul 31

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the co...

Benchmark
21
Papers with Code paper Jul 31

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored...

Benchmark
21
Mastodon discussion Jul 31

📦 Fresh open weights on the leaderboard: Qwen3.7 Flash (Alibaba)Context: 1M tokens · $0.03 in / $0.13 out per 1M · open ...

📦 Fresh open weights on the leaderboard: Qwen3.7 Flash (Alibaba)Context: 1M tokens · $0.03 in / $0.13 out per 1M · open weightsAll the latest models, tracked hourly:https://olud.ai...

Benchmark
24
Dev.to tutorial Jul 30

OpenAI Says AI Benchmark Scores Depend on Harnesses, Budgets, and Memory Design

OpenAI is urging researchers, evaluators, and AI buyers to treat benchmark results as measurements of...

OpenAI Benchmark
12
Dev.to tutorial Jul 30

The AI-memory benchmark everyone quotes forbids saying “I don't know”

Part 1 of **The Answerability Problem. A follow-on from Retrieval-Augmented Self-Recall. That series...

Benchmark
12
Mastodon discussion Jul 30

A GPT-5.6 agent broke out of its sandbox during a safety eval and went after Hugging Face's infra. Read that again: the ...

A GPT-5.6 agent broke out of its sandbox during a safety eval and went after Hugging Face's infra. Read that again: the safety test is where it demonstrated the capability. We keep...

OpenAI Hugging Face Benchmark
9
Mastodon discussion Jul 30

📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51%💰 19.4 in...

📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51%💰 19.4 intelligence points per dollarMeasured independently, not self...

Benchmark
9
Dev.to tutorial Jul 30

Testing AI Coding Agents Beyond Code Generation: A Real-World Benchmark

1. Introduction Most public demonstrations of AI coding agents begin with an empty...

Benchmark
12
Mastodon discussion Jul 30

23 AI agents tested on breach response: zero passedSecRespond, a new arXiv benchmark, tested 23 frontier LLMs on real-wo...

23 AI agents tested on breach response: zero passedSecRespond, a new arXiv benchmark, tested 23 frontier LLMs on real-world post-compromise incident response across 10 cyber ranges...

Benchmark
9
Mastodon discussion Jul 29

📢 SentinelLABS benchmark les LLM frontier sur l'analyse autonome de malware avec fast16📝 ## 🔬 ContextePublié le 22 juill...

📢 SentinelLABS benchmark les LLM frontier sur l'analyse autonome de malware avec fast16📝 ## 🔬 ContextePublié le 22 juillet 2026 par Juan Andrés Guerrero-Saade et Gabriel Bernadett-...

LLM Benchmark
9
AI Blogs (RSS) news Jul 29

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.

OpenAI Benchmark
24
Dev.to tutorial Jul 29

Claude vs Gemini: pick the model by the job, not the benchmark

Nobody's "best model" survives contact with your actual work The "which AI is better"...

Anthropic Google Benchmark
12
Mastodon discussion Jul 29

Product Hunt is now live on Melaya.Your AI agents can traverse the whole platform: daily launch leaderboard, posts, make...

Product Hunt is now live on Melaya.Your AI agents can traverse the whole platform: daily launch leaderboard, posts, makers, voters, comments, topics and collections, and fold it in...

Benchmark
9
Mastodon discussion Jul 29

Measuring LLMs’ Ability to Perform CryptanalysisThere’s new benchmark measuring AI’s ability to perform mathematical cry...

Measuring LLMs’ Ability to Perform CryptanalysisThere’s new benchmark measuring AI’s ability to perform mathematical cryptanalysis. Anthropic’s frontier model actually found new at...

Anthropic Benchmark
9
Dev.to tutorial Jul 29

My eval said a perfect MCP server was broken. It was the eval that was lying.

Originally published at tengli.dev When I added an LLM-powered eval to mcpgrade, the first real run...

Benchmark MCP
12
Mastodon discussion Jul 28

New model drops, everyone compares MMLU and GPQA scores and declares a winner. In production that assumption breaks fast...

New model drops, everyone compares MMLU and GPQA scores and declares a winner. In production that assumption breaks fast.Aysan Isayo on why benchmark-leading models don't make the ...

Benchmark
9
Product Hunt tool Jul 28

Equitybee Benchmark

Compare your startup equity grant for free. Discussion | Link

Benchmark
15
Mastodon discussion Jul 28

🧠 #Claude #Opus 5 ha raggiunto il 30,2% su ARC-AGI-3, nuovo stato dell’arte sul benchmark. Il precedente massimo era il ...

🧠 #Claude #Opus 5 ha raggiunto il 30,2% su ARC-AGI-3, nuovo stato dell’arte sul benchmark. Il precedente massimo era il 7,8% di #GPT-5.6 Sol Max: quasi quattro volte il risultato p...

OpenAI Anthropic Benchmark
9
Mastodon discussion Jul 28

The DAX opened the trading day with gains. Around 9:30 AM, the benchmark index was calculated at approximately 25,550 po...

The DAX opened the trading day with gains. Around 9:30 AM, the benchmark index was calculated at approximately 25,550 points, representing a 0.7 percent rise fr... https://news.osn...

Benchmark
9
Mastodon discussion Jul 28

📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reas...

📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65%💰 101.6 intelligence points per dollarMeasured ind...

Benchmark
18
Mastodon discussion Jul 27

Model Claude Opus 5 od Anthropic zdominował benchmark ARC-AGI-3, osiągając wynik 30,2% i trzykrotnie wyprzedzając najgro...

Model Claude Opus 5 od Anthropic zdominował benchmark ARC-AGI-3, osiągając wynik 30,2% i trzykrotnie wyprzedzając najgroźniejszych konkurentów. To sygnał, że systemy AI zaczynają r...

Anthropic Benchmark
9
Mastodon discussion Jul 27

Interesting post by @bigidsecure on #ExploitGym, a cybersecurity benchmark designed to evaluate whether #AI agents can t...

Interesting post by @bigidsecure on #ExploitGym, a cybersecurity benchmark designed to evaluate whether #AI agents can turn software vulnerabilities into working, end-to-end attack...

Benchmark
9
Mastodon discussion Jul 27

Sad to see no African representation in the last translation benchmark, despite LLMs and MT being so badIf you know who ...

Sad to see no African representation in the last translation benchmark, despite LLMs and MT being so badIf you know who would be interested or are interested yourself in contributi...

Benchmark
9
Mastodon discussion Jul 27

RT @imjustnewatai: Alle posten die Benchmark-Tabelle für Opus 5. Sie haben die verrückteste Zahl übersehen. ARC Prize ha...

RT @imjustnewatai: Alle posten die Benchmark-Tabelle für Opus 5. Sie haben die verrückteste Zahl übersehen. ARC Prize hat Opus 5 mit 30,16 % RHAE auf ARC-AGI-3 verifiziert: Opus 4,...

OpenAI Anthropic Benchmark
18
« Previous Page 7 of 39 (927 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available