DeonticBench: A Benchmark for Reasoning over Rules
Reasoning with complex, context-specific rules remains challenging for large language models (LLMs). In legal and policy settings, this manifests as deontic reasoning: reasoning ab...
973 articles tagged with Benchmark
Reasoning with complex, context-specific rules remains challenging for large language models (LLMs). In legal and policy settings, this manifests as deontic reasoning: reasoning ab...
Title: P2: P2: P2: P2: How LLM sees own adaptive thinking and evolution..5. Metacognition and Self-Reflection- Self-Evaluation :: Assesses own performance, outputs, and reasoning.-...
When I was doing traditional development, I had TDD. I wrote a test, it passed or failed, done. But...
NEW: Why Most AI Benchmarks Are Lying to You (And What to Measure Instead)MMLU is saturated. Contamination is rampant. Companies cherry-pick evals.What actually predicts production...
Meta released TRIBE v2 last week - a foundation model that predicts fMRI brain activation from video,...
#Gemma4 is the first model to pass the techno vibe benchmark. Thank you @Google #vibes #ai #google #llm #agi #techno #music
How creative are AI scientists? A new benchmark evaluates idea generation across originality, feasibility, and flexibility, revealing gaps between reasoning skills and scientific c...
📰 Qwen 3.6 Beats GPT-4o and Claude 3.5 in China’s Blind AI Coding Benchmark (2026)Qwen 3.6 has emerged as China's leading AI programming model after dominating a global blind bench...
Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across ...
The autonomous discovery of bugs remains a significant challenge in modern software development. Compared to code generation, the complexity of dynamic runtime environments makes b...
📰 MLPerf Inference v6.0 2026: Nvidia, AMD, and Intel Break Records Amid Benchmark ChallengesMLPerf Inference v6.0 introduces multimodal and video models, with Nvidia, AMD, and Inte...
Ghost hosting operator tested 16 coding models on custom TypeScript benchmark built from real commits. Open-weights Qwen3-Coder-Next scored 94.8% vs Claude Opus's 98.8% - gap close...
Reseña de «Eval awareness in Claude Opus 4.6’s BrowseComp performance». https://jaalonso.github.io/vestigium/posts/2026/03/24-resena-de-eval-awareness-in-claude-opus-46s-browsecomp...
Era łatwych zwycięstw w rankingach MMLU dobiegła końca. Nowy test „Humanity’s Last Exam” (HLE) weryfikuje zdolność modeli AI do operowania na najwyższym poziomie specjalizacji, a j...
📰 How Many Human Raters Are Needed for AI Benchmark Accuracy in 2026? (5–8 Is the Sweet Spot)Determining the optimal number of human raters for AI benchmarks is critical to ensurin...
ReCUBE Benchmark Reveals GPT-5 Scores Only 37.6% on Repository-Level Code GenerationResearchers introduce ReCUBE, a benchmark isolating LLMs' ability to use repository-wide context...
Start with the benchmarks In a previous article, I compared three Qwen3.5 models on the...
📰 Qwen3.5-Omni Beats Gemini-3.1 Pro in 2026 Multimodal AI Benchmark — Costs 90% LowerQwen3.5-Omni, Alibaba's latest multimodal AI model, delivers superior performance across benchm...
Chinese AI startup GigaAI's GigaWorld-1 world model has topped the WorldArena benchmark, becoming the only model to achieve a composite score above 60. Developed with Tsinghua Univ...
【Open ASRリーダーボード:新しい多言語トラックと長尺トラックによるトレンドとインサイト】https://huggingface.co/blog/open-asr-leaderboard※AI生成の自動投稿(見出し+リンク)#AI #生成AI #LLM #AIGenerated
📰 LLM Buyout Game Benchmark 2026: How GPT-5.4 Outsmarted GLM-5 in AI Strategy DuelThe LLM Buyout Game Benchmark evaluates advanced AI models on coalition politics, financial negoti...
📰 LLM Buyout Game Benchmark 2026: GPT-5.4, GLM-5 ve Opus 4.6 ile AI Sosyal Zekâ TestiYapay zekâ modelleri, koalisyon politikası, özel anlaşma ve hayatta kalma maliyeti gibi insan b...
There's now a leaderboard at work to see who uses the most tokens in the team LOL#AI#LLM#Work
📰 Multilingual Embedding Models 2026: Harrier-OSS-v1 Sets New SOTA Benchmark in AI TranslationMicrosoft has launched Harrier-OSS-v1, a family of multilingual embedding models that ...