/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

959 articles tagged with Benchmark

Latest Trending
Dev.to tutorial May 24

Evaluation & Benchmark Results

Multimodal Gemma 4 Visual Regression & Patch Agent devchallenge gemmachallenge gemma ai Gemma...

Google Benchmark
20
Mastodon discussion May 23

China’s New AI Shocks The World: Hits Top 10 Globally OvernightA sudden disruption on the global leaderboard. A newly de...

China’s New AI Shocks The World: Hits Top 10 Globally OvernightA sudden disruption on the global leaderboard. A newly deployed foundational model originating from China has breache...

Benchmark
9
GitHub Trending repo May 23

sitodowubb/spatial-vqa-bench: Spatial-VQA-Bench: a focused benchmark of spatial visual reasoning for multimodal LLMs.

Spatial-VQA-Bench: a focused benchmark of spatial visual reasoning for multimodal LLMs.

Multimodal Benchmark
63
GitHub Trending repo May 23

okturro/asr-rescore-bench: Benchmark for LLM-based ASR n-best rescoring (ngram, neural-LM, MLM-PLL, LLM-prompt strategies).

Benchmark for LLM-based ASR n-best rescoring (ngram, neural-LM, MLM-PLL, LLM-prompt strategies).

LLM Benchmark
63
Mastodon discussion May 22

Artificial Analysis (@ArtificialAnlys)Cartesia의 Sonic-3.5가 Artificial Analysis Speech Arena Leaderboard에서 1위를 차지해 Inworl...

Artificial Analysis (@ArtificialAnlys)Cartesia의 Sonic-3.5가 Artificial Analysis Speech Arena Leaderboard에서 1위를 차지해 Inworld Realtime TTS 1.5 Max와 Google Gemini 3.1 Flash TTS를 제쳤습니다. ...

Google Benchmark
18
Mastodon discussion May 22

📜 Latest Top Story on #HackerNews: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark🔍 Original Story: htt...

📜 Latest Top Story on #HackerNews: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark🔍 Original Story: https://modelrift.com/blog/openscad-llm-benchmark/👤 Author: jet...

LLM Benchmark
9
Dev.to tutorial May 22

An open source LLM eval tool with two independent quality signals

LLM-as-judge has become the dominant pattern for evaluating language model outputs. Tools like...

LLM Open Source Benchmark
12
Mastodon discussion May 22

📜 Latest Top Story on #HackerNews: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark🔍 Original Story: htt...

📜 Latest Top Story on #HackerNews: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark🔍 Original Story: https://modelrift.com/blog/openscad-llm-benchmark/👤 Author: jet...

LLM Benchmark
9
Mastodon discussion May 22

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmarkhttps://modelrift.com/blog/openscad-llm-benchmark/#llm

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmarkhttps://modelrift.com/blog/openscad-llm-benchmark/#llm

LLM Benchmark
9
Hacker News discussion May 22

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

LLM Benchmark
69
GitHub Trending repo May 21

Neal006/memorylens: The open-source benchmark for LLM memory decay. Measure how Naive, RAG, Chunked RAG, Cascading, and SummaryMemory degrade over 100 conversation turns. Ebbinghaus forgetting curves, 5-provider LLM eval, multi-seed CI. No API key needed.

The open-source benchmark for LLM memory decay. Measure how Naive, RAG, Chunked RAG, Cascading, and SummaryMemory degrade over 100 conversation turns. Ebbinghaus forgetting curves,...

LLM RAG Open Source
42
Papers with Code paper May 21

TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports training models to bypass this de...

Benchmark
21
Papers with Code paper May 21

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential ch...

Benchmark
21
Mastodon discussion May 20

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し...

【Open ASR リーダーボードに Benchmaxxer Repellant を追加】https://huggingface.co/blog/open-asr-leaderboard-private-data※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face Benchmark
18
Dev.to tutorial May 20

Which LLM is the best stock picker? I built a benchmark to find out.

7 frontier LLMs. $100K each. Same prompts, same tools, same data. Different brains. Here's the architecture.

LLM Benchmark
12
Mastodon discussion May 20

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #...

【QIMMA قِمّة ⛰: 品質第一のアラビア語LLMリーダーボード】https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated

Hugging Face LLM Benchmark
18
GitHub Trending repo May 20

MFS9628/Deepseek-v4-pro-app: DeepSeek v4 Pro github Flash chat: API flash gemma 4 gemini qwen claude chatgpt 4 key pricing tier, open source weights, huggingface model repository, local execution ollama setup. context window token limit, coding benchmark leaderboard ranking, reasoning model architecture v4, .visual studio code extension integration, cursor ai

DeepSeek v4 Pro github Flash chat: API flash gemma 4 gemini qwen claude chatgpt 4 key pricing tier, open source weights, huggingface model repository, local execution ollama setup....

OpenAI Anthropic Google
66
Mastodon discussion May 20

Benchmark results for Qwen 3.6 27B and 35B MTP speculative decoding in llama.cpp on RTX 4080 16GB. Token speed, VRAM cos...

Benchmark results for Qwen 3.6 27B and 35B MTP speculative decoding in llama.cpp on RTX 4080 16GB. Token speed, VRAM cost, and optimal --spec-draft-n-max settings.#SelfHosting #LLM...

Benchmark
27
Papers with Code paper May 20

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. For simpler systems, human ...

LLM Benchmark
21
Dev.to tutorial May 20

AgentThreatBench: The First OWASP Agentic Top 10 Security Benchmark

The AI safety community has a blind spot. We have excellent benchmarks for measuring whether an LLM...

Agents Benchmark
12
Mastodon discussion May 19

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels#LLM #Benchmarking #Packagehttps://hgpu....

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels#LLM #Benchmarking #Packagehttps://hgpu.org/?p=30812

LLM Benchmark AI Hardware
18
Mastodon discussion May 19

🤖 AI AGENTSOpen Agent Leaderboard: good start, but what's the incentive to game it? Seems like optimizing for benchmarks...

🤖 AI AGENTSOpen Agent Leaderboard: good start, but what's the incentive to game it? Seems like optimizing for benchmarks could quickly diverge from real-world usefulness. Thoughts?...

Benchmark
18
Mastodon discussion May 19

Najnowszy benchmark ARFBench dowodzi, że w diagnozowaniu awarii systemów inżynierowie wciąż miażdżą GPT-5 i Gemini. Rzec...

Najnowszy benchmark ARFBench dowodzi, że w diagnozowaniu awarii systemów inżynierowie wciąż miażdżą GPT-5 i Gemini. Rzeczywistość systemów produkcyjnych brutalnie weryfikuje market...

OpenAI Google Benchmark
18
Dev.to tutorial May 19

Your model speed benchmark is measuring the wrong thing

Model speed is not a property of the model. It is a property of the model plus your payload size plus...

Benchmark
12
« Previous Page 25 of 40 (959 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available