/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

973 articles tagged with Benchmark

Latest Trending
NewsData.io news Apr 10

Inaugural State of AI in Gaming Report Establishes Benchmark for AI Practice and Policy Across the Global Gambling Industry

The UNLV International Gaming Institute's (IGI) AI Research Hub (AiR Hub) in collaboration with KPMG LLP, the U.S. audit, tax and advisory firm, today released The State of AI in G...

Benchmark
21
Mastodon discussion Apr 9

I expanded the #Vera benchmark to six models across three providers (Anthropic, OpenAI, Moonshot). Kimi K2.5 got every V...

I expanded the #Vera benchmark to six models across three providers (Anthropic, OpenAI, Moonshot). Kimi K2.5 got every Vera problem right while only managing 86% in Python and 91% ...

OpenAI Anthropic Benchmark
18
Papers with Code paper Apr 9

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environments remains severely limited....

Multimodal Benchmark
21
Papers with Code paper Apr 9

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and v...

Benchmark
21
NewsData.io news Apr 9

HappyHorse-1.0 Crowned #1 Open-Source AI Video Generator, Tops Artificial Analysis Global Leaderboard

Open Source Benchmark
21
Mastodon discussion Apr 8

Organizational dynamics at Meta yielded a leaderboard for AI usage because it is a legible compromise for leadership. It...

Organizational dynamics at Meta yielded a leaderboard for AI usage because it is a legible compromise for leadership. It turns the work into a tokenized game to force an "agentic" ...

Benchmark
24
Dev.to tutorial Apr 8

A Postmortem on Autonomous LLM-as-Judge: How My Eval Agent Got Two Verdicts Wrong Before I Found a Sandbox Bug

I run an autonomous eval agent against new coding-agent stacks before trusting their numbers. The...

LLM Benchmark
12
Mastodon discussion Apr 8

📰 GLM-5.1 on Huawei Cloud: The 2026 AI Benchmark with Open-Weight PerformanceGLM-5.1 has launched on Huawei Cloud, offer...

📰 GLM-5.1 on Huawei Cloud: The 2026 AI Benchmark with Open-Weight PerformanceGLM-5.1 has launched on Huawei Cloud, offering users access through multiple integrated products. The m...

Benchmark
18
Mastodon discussion Apr 8

📰 GLM-5.1 Outperforms Opus4.6: Open Source AI Breakthrough Sets New Benchmark (2026)The open source model GLM-5.1 has un...

📰 GLM-5.1 Outperforms Opus4.6: Open Source AI Breakthrough Sets New Benchmark (2026)The open source model GLM-5.1 has unexpectedly surpassed Opus4.6 in performance benchmarks, trig...

Open Source Benchmark
24
Papers with Code paper Apr 8

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, perio...

Benchmark
21
Dev.to tutorial Apr 7

What Gemma 4's multi-token prediction head actually means for your eval pipeline

Gemma 4 dropped with a multi-token prediction (MTP) head and immediately every benchmark thread on...

Google Benchmark
28
GitHub Trending repo Apr 7

moatifbutt/gen-color-bench: Benchmark for evaluating color understanding in text-to-image models | CVPR 2026

Benchmark for evaluating color understanding in text-to-image models | CVPR 2026

Image Generation Benchmark
35
Dev.to tutorial Apr 7

I published my benchmark scores. Your turn.

Back in March I released agent-egress-bench, a test corpus for evaluating security tools that sit...

Benchmark
12
Mastodon discussion Apr 7

I just read that Meta has an internal leaderboard for LLM token usage, and I really can't imagine a less thoughtful, les...

I just read that Meta has an internal leaderboard for LLM token usage, and I really can't imagine a less thoughtful, less useful status metric. Even assuming that the tokens are be...

LLM Benchmark
38
Mastodon discussion Apr 7

A 600-run benchmark by Yusuke Endoh tested #ClaudeCode across 13 #ProgrammingLanguages by implementing a simplified Git....

A 600-run benchmark by Yusuke Endoh tested #ClaudeCode across 13 #ProgrammingLanguages by implementing a simplified Git.Key takeaways:⇨ Ruby, Python & JavaScript were the fastest, ...

Benchmark
24
Papers with Code paper Apr 7

Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents

Large language models are increasingly deployed as autonomous agents executing multi-step workflows in real-world software environments. However, existing agent benchmarks suffer f...

Benchmark
21
Papers with Code paper Apr 7

MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts

Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured bio...

Benchmark
21
Mastodon discussion Apr 7

Microsoft udostępnia trzy autorskie modele MAI w ramach platformy Foundry. Wśród nowości znalazł się lider benchmarków t...

Microsoft udostępnia trzy autorskie modele MAI w ramach platformy Foundry. Wśród nowości znalazł się lider benchmarków transkrypcji oraz narzędzia do błyskawicznej syntezy mowy i f...

Microsoft Benchmark
18
Mastodon discussion Apr 6

📰 LTX-Video 2.3 Benchmark: 11 Models Tested for AI Video GenerationLTX-Video 2.3 emerges as a leading AI video generatio...

📰 LTX-Video 2.3 Benchmark: 11 Models Tested for AI Video GenerationLTX-Video 2.3 emerges as a leading AI video generation framework, with 11 distinct models tested across hardware ...

Benchmark
9
Mastodon discussion Apr 6

📰 Gemma 4 Leads 2026 AI Leaderboard with $0.20 Inference Cost and 100% Survival RateGemma 4, a 31-billion-parameter mode...

📰 Gemma 4 Leads 2026 AI Leaderboard with $0.20 Inference Cost and 100% Survival RateGemma 4, a 31-billion-parameter model, has shattered benchmarks by achieving 1,144% median ROI a...

OpenAI Google Benchmark
9
Mastodon discussion Apr 6

📰 AI Memory Benchmark 2026: Chinese Teen Coders Break Records with 98.7% Referential Resolution Acc...A new AI memory be...

📰 AI Memory Benchmark 2026: Chinese Teen Coders Break Records with 98.7% Referential Resolution Acc...A new AI memory benchmark, pioneered by teenage developers from China, has ach...

Benchmark
18
Mastodon discussion Apr 6

Is AI just a benchmark race? Google’s Gemini 3.1 Pro says yes, but 1 million medical images processed in Galicia say the...

Is AI just a benchmark race? Google’s Gemini 3.1 Pro says yes, but 1 million medical images processed in Galicia say there’s a much deeper story. From doctor's offices to insurance...

Google Benchmark
18
Mastodon discussion Apr 6

Google DeepMind just scored 85% on ARC-AGI-2 — the most difficult general reasoning benchmark in AI.Previous best? 54%.T...

Google DeepMind just scored 85% on ARC-AGI-2 — the most difficult general reasoning benchmark in AI.Previous best? 54%.This kind of leap doesn't happen often. General reasoning is ...

Google Benchmark
18
Mastodon discussion Apr 6

new writeup: why eval pipelines matter more than model selection.teams with good evals ship 3x faster. teams without the...

new writeup: why eval pipelines matter more than model selection.teams with good evals ship 3x faster. teams without them are flying blind, afraid to change anything because they c...

Benchmark
30
« Previous Page 37 of 41 (973 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available