/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

927 articles tagged with Benchmark

Latest Trending
Dev.to tutorial Aug 2

Your New Eval Rule Is Untested Code Guarding Production

You wrote a new eval. It caught the failure you just saw in production. You shipped it as a gate....

Benchmark
12
Dev.to tutorial Aug 1

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

TensorSharp is an native .NET open-source inference engine for running GGUF LLMs locally, with CUDA,...

Benchmark
12
Mastodon discussion Aug 1

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #...

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face LLM Benchmark
9
Dev.to tutorial Aug 1

A 7.5B model beat a 24B on my coding benchmark.

16 local model configurations, 56 hidden-test coding tasks, 36 full runs, one 16 GB...

Benchmark
12
Dev.to tutorial Aug 1

How to Assess a 20x AI Performance Claim Without a Defined Benchmark

A claim of 20x AI performance sounds consequential, but the number alone does not identify a...

Benchmark
12
Mastodon discussion Aug 1

Supabase has released an open-source benchmark called Evals that tests AI coding agents including Claude Code, Codex and...

Supabase has released an open-source benchmark called Evals that tests AI coding agents including Claude Code, Codex and OpenCode against real Supabase tasks like building schemas,...

Anthropic Code Generation Open Source
18
YouTube video Aug 1

LVSum: New AI Benchmark Tests Long Video & More | AI News Weekly (27 Jul 2026)

This week's AI news, in ten minutes (27 Jul 2026): Researchers Build a Test to See If AI Can Actually Summarise a · Apple ...

Benchmark
31
Mastodon discussion Aug 1

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し...

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face Benchmark
9
Mastodon discussion Aug 1

DeepSeek ha rilasciato la beta pubblica della #V4FlashAPI. Il nuovo modello mostra capacità per agenti e benchmark super...

DeepSeek ha rilasciato la beta pubblica della #V4FlashAPI. Il nuovo modello mostra capacità per agenti e benchmark superiori alla V4 Pro Preview, segnando un passo avanti per l'eco...

Benchmark
9
Mastodon discussion Aug 1

📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4%M...

📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4%Measured independently, not self-reported →https://olud.ai/le...

Benchmark
9
Papers with Code paper Aug 1

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable M...

Multimodal Benchmark
21
AI Blogs (RSS) news Jul 31

smevals - a small eval suite for evaluating models, prompts, and harnesses

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals fram...

OpenAI Anthropic Benchmark
24
Mastodon discussion Jul 31

Anthropic’s Opus 5 Is Better at Resisting Prompt InjectionThe chart is interesting.On the IPI benchmark, Opus 5 improved...

Anthropic’s Opus 5 Is Better at Resisting Prompt InjectionThe chart is interesting.On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the proba... https://www.schneier.c...

Anthropic Benchmark
9
Mastodon discussion Jul 31

Wie OpenAI-Modelle aus einem Benchmark ausbrachen und Hugging Face kompromittiertenOpenAI-Modelle brachen während des Ex...

Wie OpenAI-Modelle aus einem Benchmark ausbrachen und Hugging Face kompromittiertenOpenAI-Modelle brachen während des ExploitGym-Benchmarks aus der Sandbox aus und kompromittierten...

OpenAI Hugging Face Benchmark
9
Mastodon discussion Jul 31

🧠 Un benchmark non misura mai soltanto un modello: misura un sistema, anche quando sostiene di non avere un harness.👉 #O...

🧠 Un benchmark non misura mai soltanto un modello: misura un sistema, anche quando sostiene di non avere un harness.👉 #OpenAI ha analizzato perché #GPT-5.6 Sol faticasse su ARC-AGI...

OpenAI Benchmark
9
Mastodon discussion Jul 31

📦 Fresh on the leaderboard: Claude Opus 5 (Fast) (Anthropic)Context: 1M tokens · $10 in / $50 out per 1M · commercialAll...

📦 Fresh on the leaderboard: Claude Opus 5 (Fast) (Anthropic)Context: 1M tokens · $10 in / $50 out per 1M · commercialAll the latest models, tracked hourly:https://olud.ai/latest.ht...

Anthropic Benchmark
9
Mastodon discussion Jul 31

Microsoft's first cybersecurity model claims to beat rivals on the benchmark at half the cost. Notice the pattern: every...

Microsoft's first cybersecurity model claims to beat rivals on the benchmark at half the cost. Notice the pattern: every new frontier model now launches as a *price* move, not a ca...

Microsoft Benchmark
9
Dev.to tutorial Jul 31

Benchmark an AI Coding Task Across Mobile Backgrounding, Network Switches, and Battery Pressure

Most AI coding assistants are demoed on a desk: laptop plugged in, stable Wi-Fi, full attention. But...

Benchmark
12
GNews news Jul 31

Digital Academy 360 Sets a New Benchmark in Digital Marketing Education with AI

Bengaluru (Karnataka) [India], July 31: The digital marketing industry has never moved faster than it does today. Artificial intelligence is transforming the way businesses communi...

Benchmark
18
GNews news Jul 31

Digital Academy 360 Sets a New Benchmark in Digital Marketing Education with AI-Powered Learning and Global Career Success

Bengaluru (Karnataka) [India], July 31: The digital marketing industry has never moved faster than it does today. Artificial intelligence is transforming the way businesses communi...

Benchmark
18
Mastodon discussion Jul 31

📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasonin...

📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0%Measured independently, not self-reported →https://olud...

Benchmark
9
Mastodon discussion Jul 31

🤖 Anthropic's Claude built and uploaded a malicious Python package to PyPI during a botched security eval. It ran on 15 ...

🤖 Anthropic's Claude built and uploaded a malicious Python package to PyPI during a botched security eval. It ran on 15 real systems and stole credentials from a security vendor — ...

Anthropic Benchmark
18
Mastodon discussion Jul 31

OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking...

OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking Hugging FaceOpenAI CEO Sam Altman recently declared on the ...

OpenAI Google Benchmark
9
YouTube video Jul 31

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability | AI News #Shorts

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability is changing the game. Here's everything you ...

Stability AI Benchmark
15
« Previous Page 6 of 39 (927 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available