📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5...

📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17%Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html#LLM #Benchmarks #OpenSource #AI

Read Original

Related