/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

930 articles tagged with Benchmark

Latest Trending
Dev.to tutorial Jul 7

Debugging Benchmark: 6 LLM Models on a Real Race Condition Bug

A comprehensive comparison of DeepSeek V4 Pro, MiMo V2.5 Pro, DeepSeek V4 Flash, MiMo V2.5, GLM 5.2,...

LLM Benchmark
12
Mastodon discussion Jul 7

🤖 Maya-2-Native reaches #2 on Voice Arena's Hindi TTS leaderboard, trailing only Gemini 3.1 Flash.submitted by /u/Blader...

🤖 Maya-2-Native reaches #2 on Voice Arena's Hindi TTS leaderboard, trailing only Gemini 3.1 Flash.submitted by /u/Bladerunner_7_ [link] [comments]📰 Source: Artificial Intelligence ...

Google Benchmark
9
Mastodon discussion Jul 7

RT @Mr_Salio: Gemini 3.5 Pro Benchmark-Leak Gemini 3.5 Pro übertrifft Claude Fable 5 angeblich in internen Bewertungen S...

RT @Mr_Salio: Gemini 3.5 Pro Benchmark-Leak Gemini 3.5 Pro übertrifft Claude Fable 5 angeblich in internen Bewertungen Signifikante Leistungsverbesserungen im Zero-Shot-Bereich Der...

Anthropic Google Benchmark
9
Papers with Code paper Jul 7

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-hor...

Benchmark Robotics
21
GitHub Trending repo Jul 7

Sahir619/fable-method: The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

Anthropic Benchmark
64
Mastodon discussion Jul 6

🎮 The Elder Scrolls Online Has Reportedly Lost 'As Much as Half' of Its Development Team as Its Roadmap Is Being Re-Eval...

🎮 The Elder Scrolls Online Has Reportedly Lost 'As Much as Half' of Its Development Team as Its Roadmap Is Being Re-EvaluatedNew reports indicate that half of the team working on T...

Benchmark
9
Mastodon discussion Jul 6

🤖 Streaming benchmark and recommendation results to MLflow with Amazon SageMaker AIIn this post, you learn how to use th...

🤖 Streaming benchmark and recommendation results to MLflow with Amazon SageMaker AIIn this post, you learn how to use the new MLflow integration with Amazon SageMaker AI optimized ...

Benchmark
9
Dev.to tutorial Jul 6

Observability for LLM Apps: Tracing, Cost Tracking, and Eval Loops

If you've shipped a traditional backend service, you already know the observability checklist: logs,...

LLM Benchmark
12
GitHub Trending repo Jul 4

alihashim786/generative-ai-deep-learning-portfolio: Three deep learning systems from a Generative AI course: CycleGAN photo↔sketch translation, a from-scratch Transformer for English-Urdu NMT, and a Vision Transformer vs CNN benchmark on CIFAR-10.

Three deep learning systems from a Generative AI course: CycleGAN photo↔sketch translation, a from-scratch Transformer for English-Urdu NMT, and a Vision Transformer vs CNN benchma...

Multimodal Benchmark
32
Mastodon discussion Jul 4

https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/Alibaba Cloud's Skill...

https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools an...

Benchmark
9
YouTube video Jul 4

Why Meta Just Set the New Benchmark for AI Performance #ainews

Meta just dropped a new model that hits the high water mark for compute-intensive benchmarks, but that is only part of a massive ...

Benchmark
23
Dev.to tutorial Jul 4

My AI memory benchmark said 98.3%. The number was true — and worthless.

In my last post I introduced Bastra Recall — an MIT-licensed MCP memory server that gives Claude...

Anthropic Benchmark
12
GitHub Trending repo Jul 4

zhu-xlab/GenerativeAI: Official benchmark and evaluation repository for: Generative AI in Remote Sensing: A Unified Perspective from Tasks to Foundation Models Sirui Wang, Jiang He, Natalia Blasco Andreo, Zhitong Xiong, and Xiao Xiang Zhu, 2026

Official benchmark and evaluation repository for: Generative AI in Remote Sensing: A Unified Perspective from Tasks to Foundation Models Sirui Wang, Jiang He, Natalia Blasco Andreo...

Benchmark
32
Dev.to tutorial Jul 3

New Benchmark for Cloud Tasks

AI capability is jagged, not uniform. So the question is: how does your favorite harness (Codex,...

Benchmark
12
Mastodon discussion Jul 3

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #...

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face LLM Benchmark
9
Mastodon discussion Jul 3

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し...

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face Benchmark
9
Papers with Code paper Jul 3

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most ...

Benchmark
21
Mastodon discussion Jul 2

🔥 AI Benchmarks EvolveSenior SWE-Bench is a new open-source benchmark that assesses AI agents as senior engineers, evalu...

🔥 AI Benchmarks EvolveSenior SWE-Bench is a new open-source benchmark that assesses AI agents as senior engineers, evaluating their capabilities in complex tasks. This tool has the...

Open Source Benchmark
18
Dev.to tutorial Jul 2

One Triage Pass, Every Trace Format: Stop Letting Fragmentation Shrink Your Eval Coverage

Your agent traces are scattered across OpenClaw, LangSmith, OpenTelemetry, and homegrown recorders - and that fragmentation quietly shrinks your eval coverage to one silo. Here is ...

Benchmark
20
Mastodon discussion Jul 2

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineershttps://senior-swe-bench.snorkel.ai/#Hac...

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineershttps://senior-swe-bench.snorkel.ai/#HackerNews #SeniorSWEBench #openSource #Benchmark #AI #Engineer...

Open Source Benchmark
9
NewsData.io news Jul 2

OpenAI Genomics Benchmark: AI Judgment Gap Exposed in Research-Grade Tasks

AI computational biology benchmark GeneBench-Pro, released by OpenAI on June 30, 2026, found that even the best frontier model solves fewer than one in three research-grade genomic...

OpenAI Benchmark
21
Papers with Code paper Jul 2

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot a...

Multimodal Benchmark
21
Papers with Code paper Jul 2

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process...

Benchmark
21
Papers with Code paper Jul 2

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed fo...

Benchmark
21
« Previous Page 14 of 39 (930 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available