/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

926 articles tagged with Benchmark

Latest Trending
Mastodon discussion 2d ago

VibeLifeBench introduces a benchmark of 200 long-horizon tasks to test if LLM agents can act proactively and persistentl...

VibeLifeBench introduces a benchmark of 200 long-horizon tasks to test if LLM agents can act proactively and persistently in a changing simulated world. Seven leading models all sc...

LLM Benchmark
9
Mastodon discussion 2d ago

New benchmark finds automated evaluation of tool-using LLM agents often unreliable; GPT-4o-mini surpasses heuristic judg...

New benchmark finds automated evaluation of tool-using LLM agents often unreliable; GPT-4o-mini surpasses heuristic judging, and runtime interceptors cut hallucinations by 24 perce...

LLM Multimodal Benchmark
9
Dev.to tutorial 2d ago

New Benchmark for Evaluating Long-Horizon Agents in Online Environments

RealReplicaBench offers developers a new tool for benchmarking agents in high-fidelity replicas of real online services, highlighting critical tradeoffs in AI training.

Benchmark
12
Mastodon discussion 2d ago

Lookspan's eval feature runs the first 100 items of a dataset. The panel said 'runs up to 100 items synchronously' β€” the...

Lookspan's eval feature runs the first 100 items of a dataset. The panel said 'runs up to 100 items synchronously' β€” the same sentence whether your dataset had five items or five h...

Benchmark
9
Dev.to tutorial 2d ago

From a half-finished training run to a reproducible eval

So our last training pass on Qwen3-Omni-30B-A3B-Instruct β€” fine-tuning it on Barbados newspapers β€”...

Benchmark
12
Mastodon discussion 2d ago

πŸ“Š GLM-4.6 (Reasoning) β€” the actual numbers GPQA: 78% MMLU-Pro: 82.9% Humanity's Last Exam: 14.5% Long Context Reasoning:...

πŸ“Š GLM-4.6 (Reasoning) β€” the actual numbers GPQA: 78% MMLU-Pro: 82.9% Humanity's Last Exam: 14.5% Long Context Reasoning: 55.3%πŸ’° 30.4 intelligence points per dollarMeasured independ...

Benchmark
9
Dev.to tutorial 2d ago

When your benchmark is wrong and your model is right

When your benchmark is wrong and your model is right We fine-tuned a 30-billion-parameter model on...

Benchmark
12
NewsData.io news 2d ago

Explained: Sarvam AI's Indic benchmark to test voice AI across 22 Indian languages

This launch addresses a major gap felt in the Indian AI ecosystem, especially voice models. AI voice tools have made rapid advancements in Western languages, but diverse, multi-spe...

Benchmark
21
Mastodon discussion 2d ago

New TAF-MED benchmark shows 71.6% of LLM medical conversations contain unsafe responses, with 61.4% of initially safe in...

New TAF-MED benchmark shows 71.6% of LLM medical conversations contain unsafe responses, with 61.4% of initially safe interactions later collapsing to unsafe. Conversational safety...

LLM Benchmark
9
Dev.to tutorial 2d ago

13 AI Coding Models Tested: Safety Benchmark Results KDS

Adversarial A/B testing of 13 AI coding models with keelwright safety skill. KDS scores: from 83 (Laguna S 2.1) to 0 (weak models that fabricate results).

Benchmark
12
Mastodon discussion 3d ago

πŸ‘‰ "PT2PR: A Benchmark for Multimodal Patent-to-Product Retrieval" by Lia Shahnazaryan & Stefan Heindorf, also at #CIKM20...

πŸ‘‰ "PT2PR: A Benchmark for Multimodal Patent-to-Product Retrieval" by Lia Shahnazaryan & Stefan Heindorf, also at #CIKM2026 πŸ‘ Congratulations to all authors β€” in bocca al lupo! 🐺(th...

Multimodal Benchmark
9
Mastodon discussion 3d ago

Xiaomi's MiLM Plus has released PROVE, a new perception-aligned benchmark for video object removal. The metrics RC-S and...

Xiaomi's MiLM Plus has released PROVE, a new perception-aligned benchmark for video object removal. The metrics RC-S and RC-T evaluate spatial coherence and temporal consistency wi...

Benchmark
18
Mastodon discussion 3d ago

In Hong Kong, HKU introduces RoboDojo, a unified benchmark for embodied AI to standardise robot reliability testingSourc...

In Hong Kong, HKU introduces RoboDojo, a unified benchmark for embodied AI to standardise robot reliability testingSource: University of Hong Kong Presshttp://www.hku.hk/press/news...

Benchmark Robotics
9
Mastodon discussion 3d ago

Qwen3.8 Max trails Claude Fable 5 by just 4 points on today's benchmark, yet costs 8x less per 1M output tokens. That ga...

Qwen3.8 Max trails Claude Fable 5 by just 4 points on today's benchmark, yet costs 8x less per 1M output tokens. That gap is the whole story for budget-conscious buildersβ€”see where...

Anthropic Benchmark
9
Dev.to tutorial 3d ago

We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

Our CEO Rob Imbeault published a piece on LinkedIn this week about a result our team posted: 99.95%...

Benchmark
20
Papers with Code paper 3d ago

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despit...

Multimodal Benchmark
21
Dev.to tutorial 3d ago

My fine-tuned model scored 100%... The benchmark was lying

I fine-tuned Mistral 7B on my laptop to detect personal data in log lines and support messages. On my...

Benchmark
12
Dev.to tutorial 3d ago

Part 3: Build the Eval Set Before the Agent Exists

Part 3 of a series building a support-ticket agent with no framework. Previous: Part 2 (tool...

Benchmark
12
GitHub Trending repo 3d ago

KasraAhmadi/job-eval: Score your resume against any number of jobs locally β€” free, explainable, CPU or GPU. Powered by a 156M ModernBERT cross-encoder.

Score your resume against any number of jobs locally β€” free, explainable, CPU or GPU. Powered by a 156M ModernBERT cross-encoder.

Benchmark AI Hardware
39
Mastodon discussion 3d ago

πŸ“Š Granite 4.0 1B β€” the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6%M...

πŸ“Š Granite 4.0 1B β€” the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6%Measured independently, not self-reported β†’https://olud.ai/le...

Benchmark
9
Mastodon discussion 3d ago

MasDrift benchmark shows centralized multi-agent hierarchies complete 93.9-98.6% of tasks but allow unauthorized actions...

MasDrift benchmark shows centralized multi-agent hierarchies complete 93.9-98.6% of tasks but allow unauthorized actions in 2.7-19.8% of cases, while peer networks lag in completio...

Benchmark
9
Dev.to tutorial 3d ago

Build a Reproducible AI Tool Benchmark in Python

Most AI tools look impressive in a polished demo. The harder question is whether they are reliable,...

Benchmark
12
Mastodon discussion 3d ago

New benchmark MSEval shows multi-agent coding outcomes depend as much on team topology as model strength, with topology ...

New benchmark MSEval shows multi-agent coding outcomes depend as much on team topology as model strength, with topology shifts altering scores by over 30 points and doubling runtim...

Benchmark
9
Mastodon discussion 4d ago

New benchmark RoboGraph defines task-state horizon to expose gaps in long-horizon embodied agent performanceSource: arXi...

New benchmark RoboGraph defines task-state horizon to expose gaps in long-horizon embodied agent performanceSource: arXiv cs.ROhttps://arxiv.org/abs/2608.08036#MachineLearning

Benchmark
9
« Previous Page 2 of 39 (926 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available