/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

973 articles tagged with Benchmark

Latest Trending
Papers with Code paper Mar 30

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer ...

Benchmark
21
Mastodon discussion Mar 30

https://winbuzzer.com/2026/03/30/arc-agi-3-offers-2m-ai-matching-human-reasoning-benchmark-xcxwbn/ARC-AGI-3 Offers $2M f...

https://winbuzzer.com/2026/03/30/arc-agi-3-offers-2m-ai-matching-human-reasoning-benchmark-xcxwbn/ARC-AGI-3 Offers $2M for AI Matching Human Reasoning#AI #ARCAGI #ARCAGI3 #AGI #AIB...

Benchmark
18
Papers with Code paper Mar 30

GEditBench v2: A Human-Aligned Benchmark for General Image Editing

Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks ...

Benchmark
21
Mastodon discussion Mar 30

Information Fusion: Privacy-aware detection of fake identity documents: methodology, benchmark, and improved algorithms ...

Information Fusion: Privacy-aware detection of fake identity documents: methodology, benchmark, and improved algorithms (FakeIDet2) . “Researchers are now trying to develop methods...

Benchmark
18
NewsData.io news Mar 30

Starcloud Raises $170M Series A at $1.1bn Valuation Led by Benchmark and EQT Ventures

REDMOND, Wash.--(BUSINESS WIRE)-- #AI--Starcloud, the company building data centers in space, today announced it has raised a $170 million Series A, at a $1.1 billion valuation. Ac...

Benchmark
21
Papers with Code paper Mar 30

MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, ...

Benchmark
21
Papers with Code paper Mar 30

LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models

Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings,...

Benchmark
21
Mastodon discussion Mar 29

🧠 ARC Prize Foundation ha introdotto ARC-AGI-3, un benchmark progettato per valutare come gli agenti #AI apprendono, non...

🧠 ARC Prize Foundation ha introdotto ARC-AGI-3, un benchmark progettato per valutare come gli agenti #AI apprendono, non cosa ricordano.👉 I dettagli: https://www.linkedin.com/posts...

Benchmark
18
Mastodon discussion Mar 28

📰 AI Solves the Reproducibility Crisis? 2026 Benchmark Reveals 78% Success in Automated Reproducibi...Can AI automate co...

📰 AI Solves the Reproducibility Crisis? 2026 Benchmark Reveals 78% Success in Automated Reproducibi...Can AI automate computational reproducibility? A new benchmark is emerging to ...

Benchmark
9
Mastodon discussion Mar 28

📰 Claude AI Tops 2026 Truthfulness Benchmark: Least Hallucinatory AI Outperforms ChatGPT & GeminiClaude AI has emerged a...

📰 Claude AI Tops 2026 Truthfulness Benchmark: Least Hallucinatory AI Outperforms ChatGPT & GeminiClaude AI has emerged as the least bullshit-y large language model according to a n...

OpenAI Anthropic Google
18
Mastodon discussion Mar 28

🤖 Claude is the least bullshit-y AIJust found this “bullshit benchmark,” and sort of shocked by the divergence of Anthro...

🤖 Claude is the least bullshit-y AIJust found this “bullshit benchmark,” and sort of shocked by the divergence of Anthropic’s models from other major models (ChatGPT and Gemini). I...

OpenAI Anthropic Google
9
Mastodon discussion Mar 28

Il benchmark ARC-AGI 3 svela il limite dei modelli AI attuali. Gemini 3.1 Pro si ferma allo 0,37% rispetto alla baseline...

Il benchmark ARC-AGI 3 svela il limite dei modelli AI attuali. Gemini 3.1 Pro si ferma allo 0,37% rispetto alla baseline umana. La capacità di adattamento a compiti inediti resta l...

Google Benchmark
18
Mastodon discussion Mar 27

📰 Open Source Speech Recognition Model Beats Whisper in 2026 ASR Benchmark with 12% Lower WERCohere has released an open...

📰 Open Source Speech Recognition Model Beats Whisper in 2026 ASR Benchmark with 12% Lower WERCohere has released an open source speech recognition model that outperforms industry l...

OpenAI Open Source Benchmark
9
Papers with Code paper Mar 27

PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is su...

Benchmark
21
Papers with Code paper Mar 27

Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To a...

Benchmark
21
Dev.to tutorial Mar 27

Why your word error rate (WER) benchmark might be lying to you

At AssemblyAI, we've spent years helping customers evaluate speech-to-text performance. Word Error...

Benchmark
12
Mastodon discussion Mar 27

https://winbuzzer.com/2026/03/27/cohere-open-source-transcribe-model-tops-asr-leaderboard-xcxwbn/Cohere's Open-Source Tr...

https://winbuzzer.com/2026/03/27/cohere-open-source-transcribe-model-tops-asr-leaderboard-xcxwbn/Cohere's Open-Source Transcribe Model Tops ASR Leaderboard#AI #Cohere #CohereTransc...

Open Source Benchmark
18
Mastodon discussion Mar 27

EnterpriseArena Benchmark Reveals LLM Agents Fail at Long-Horizon CFO-Style Resource AllocationResearchers introduced En...

EnterpriseArena Benchmark Reveals LLM Agents Fail at Long-Horizon CFO-Style Resource AllocationResearchers introduced EnterpriseArena, a 132-month enterprise simulator, to test LLM...

LLM Benchmark
18
Dev.to tutorial Mar 27

MagiC v0.4: Embedded Server Mode + Real Benchmark Numbers

We shipped MagiC v0.4 today. Two things worth talking about: embedded server mode and the benchmark...

Benchmark
12
Papers with Code paper Mar 27

QuitoBench: A High-Quality Open Time Series Forecasting Benchmark

Time series forecasting is critical across finance, healthcare, and cloud computing, yet progress is constrained by a fundamental bottleneck: the scarcity of large-scale, high-qual...

Benchmark
21
ArXiv paper Mar 26

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly f...

Benchmark
18
Papers with Code paper Mar 26

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly f...

Benchmark
21
Mastodon discussion Mar 26

Aktuelle KI-Modelle scheitern beim ARC-AGI-3-Benchmark für interaktives Reasoning mit Erfolgsquoten unter 0,4 Prozent.Di...

Aktuelle KI-Modelle scheitern beim ARC-AGI-3-Benchmark für interaktives Reasoning mit Erfolgsquoten unter 0,4 Prozent.Die Modelle scheitern an visuellen Transferleistungen, die unt...

Benchmark
18
ArXiv paper Mar 26

Can Users Specify Driving Speed? Bench2Drive-Speed: Benchmark and Baselines for Desired-Speed Conditioned Autonomous Driving

End-to-end autonomous driving (E2E-AD) has achieved remarkable progress. However, one practical and useful function has been long overlooked: users may wish to customize the desire...

Benchmark
18
« Previous Page 39 of 41 (973 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available