When Llama 3.1 Tulu3 405B hits 71.6% on MMLU-Pro but only 3.5% on Humanity's Last Exam, the gap shows even top open mode...

When Llama 3.1 Tulu3 405B hits 71.6% on MMLU-Pro but only 3.5% on Humanity's Last Exam, the gap shows even top open models struggle with truly hard reasoning—https://olud.ai/leaderboard.html#LLM #Benchmarks #OpenSource #AI

Read Original

Related