/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

926 articles tagged with Benchmark

Latest Trending
Mastodon discussion Aug 4

Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter ...

Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter model scores come from internal testing only. Independent ve...

Benchmark
9
Dev.to tutorial Aug 4

Surviving the Shai-Hulud: Why Agent Eval Harnesses and Local LLMs Are the New Supply Chain Defense

How local LLMs and specialized evaluation harnesses create a defense-in-depth strategy against the chaos of unmanaged AI agent supply chains.

Benchmark
12
Mastodon discussion Aug 4

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturationhttps://arxiv.org/abs/2602.16763Comments: https://...

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturationhttps://arxiv.org/abs/2602.16763Comments: https://news.ycombinator.com/item?id=49170915#HackerNews #AI #Benchm...

Benchmark
9
Hacker News discussion Aug 4

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Benchmark
64
Mastodon discussion Aug 4

Homebench – Benchmark local LLMs for speed, memory, and qualityhttps://github.com/david-g-3654/homebench#github #llm #ll...

Homebench – Benchmark local LLMs for speed, memory, and qualityhttps://github.com/david-g-3654/homebench#github #llm #llms

LLM Benchmark
9
Mastodon discussion Aug 4

Spirit AI, a Chinese physical AI startup, briefly overtook Nvidia to top the global RoboArena benchmark with its Spirit ...

Spirit AI, a Chinese physical AI startup, briefly overtook Nvidia to top the global RoboArena benchmark with its Spirit v1.6 model, raising questions about benchmark integrity in t...

NVIDIA Benchmark
9
Hacker News discussion Aug 4

Homebench – Benchmark local LLMs for speed, memory, and quality

Homebench – Benchmark local LLMs for speed, memory, and quality

Benchmark
64
Mastodon discussion Aug 4

📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reason...

📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3%⚡ 217.9 tokens/sec💰 22.4 intelligence points per d...

Benchmark
9
Papers with Code paper Aug 4

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on t...

Benchmark
21
Mastodon discussion Aug 4

Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measur...

Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measuring capability, it's leaking answers into the pretraining se...

Benchmark
9
Mastodon discussion Aug 3

Alibaba open-sources a Max-class Qwen model with coding benchmark gains, OpenAI's super PAC is caught running an AI-gene...

Alibaba open-sources a Max-class Qwen model with coding benchmark gains, OpenAI's super PAC is caught running an AI-generated news site to shape policy debates, plus a language mod...

OpenAI Benchmark
9
Mastodon discussion Aug 3

📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7%💰 ...

📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7%💰 10 intelligence points per dollarMeasured independently, not...

Benchmark
9
Papers with Code paper Aug 3

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation

Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating them on specific tasks requires large, high-quality, multi-modal ...

Benchmark
21
Dev.to tutorial Aug 2

I Built an Agent Eval Harness. Real Agents Broke the Clean Version of the Story

Two weeks ago, I published "Why Agent Evaluation Is Harder Than Model Evaluation." The core argument:...

Benchmark
33
Mastodon discussion Aug 2

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."https://frogs.vaguespac.es/#ai

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."https://frogs.vaguespac.es/#ai

Benchmark
9
Hacker News discussion Aug 2

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."

Benchmark
65
Mastodon discussion Aug 2

Meta AI tests a second agent as a memory coach to keep long tasks on track, improving benchmark scores by up to 8.3%.Sou...

Meta AI tests a second agent as a memory coach to keep long tasks on track, improving benchmark scores by up to 8.3%.Source: The Decoder AIhttps://the-decoder.com/meta-ai-uses-a-se...

Meta Benchmark
9
Mastodon discussion Aug 2

I don't work with chatbots, but this is a good article about how Airbnb used Eval Driven Development (EDD - kind of like...

I don't work with chatbots, but this is a good article about how Airbnb used Eval Driven Development (EDD - kind of like Test Driven Development (TDD)) to get a handle on testing n...

Benchmark
9
Mastodon discussion Aug 2

Composio a benchmarké le coût par tâche des #agents harnesses, sur les prix listes de Kimi K3 : #Hermes Agent à 0,39 $, ...

Composio a benchmarké le coût par tâche des #agents harnesses, sur les prix listes de Kimi K3 : #Hermes Agent à 0,39 $, #Claude Code à 1,47 $. Soit 3,7x plus cher.Le coût médian co...

Anthropic Benchmark
9
Mastodon discussion Aug 2

🧠 Researchers examine MUD (a benchmark environment) as a tool for evaluating AI systems and identify how LLM judges can ...

🧠 Researchers examine MUD (a benchmark environment) as a tool for evaluating AI systems and identify how LLM judges can become distorted in ways that standard aggregate metrics lik...

LLM Benchmark
9
Mastodon discussion Aug 2

RT @__tinygrad__: Die Benchmark-Ergebnisse liegen vor: GLM-5.2 auf einer $160k Tinybox Pro V2 Black erreicht 119 Token p...

RT @__tinygrad__: Die Benchmark-Ergebnisse liegen vor: GLM-5.2 auf einer $160k Tinybox Pro V2 Black erreicht 119 Token pro Sekunde im Einzelbetrieb und 917 Token pro Sekunde im Agg...

Benchmark
9
GitHub Trending repo Aug 2

ryan-lewisgxrc6658/aion-gridload-model-eval: Structured machine learning workflow for dataset preparation, model training, evaluation, and grid load forecasting, with organized notebooks and saved model artifacts to support repeatable data science experiments.

Structured machine learning workflow for dataset preparation, model training, evaluation, and grid load forecasting, with organized notebooks and saved model artifacts to support r...

Benchmark
48
Mastodon discussion Aug 2

When Llama 3.1 Tulu3 405B hits 71.6% on MMLU-Pro but only 3.5% on Humanity's Last Exam, the gap shows even top open mode...

When Llama 3.1 Tulu3 405B hits 71.6% on MMLU-Pro but only 3.5% on Humanity's Last Exam, the gap shows even top open models struggle with truly hard reasoning—https://olud.ai/leader...

Meta Benchmark
9
Dev.to tutorial Aug 2

Your New Eval Rule Is Untested Code Guarding Production

You wrote a new eval. It caught the failure you just saw in production. You shipped it as a gate....

Benchmark
12
« Previous Page 5 of 39 (926 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available