📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17%Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html#LLM #Benchmarks #OpenSource #AI
Related
I'm not sure how "autonomous" these attacks and sandbox escapes really are, but they are certainly leading to real conse...
I'm not sure how "autonomous" these attacks and sandbox escapes really are, but they are certainly leading to real consequences!https://www.theregister.com/security/2026/08/14/auto...
Android 17 QPR2 greatly expands Dynamic Color theming on PixelWith Android 17 QPR2 Beta 3’s surprise Friday drop, Google...
Android 17 QPR2 greatly expands Dynamic Color theming on PixelWith Android 17 QPR2 Beta 3’s surprise Friday drop, Google is introducing expanded Dynamic Color theming options for P...
"Open-source DeepResearch" reaches 55.15% on GAIA validation, up from 46%, with code agents outperforming JSON. #DeepRes...
"Open-source DeepResearch" reaches 55.15% on GAIA validation, up from 46%, with code agents outperforming JSON. #DeepResearch #AI #OpenSource #AgenticAI https://huggingface.co/blog...