📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7...

📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7%Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html#LLM #Benchmarks #OpenSource #AI

Read Original

Related