A 7.5B model beat a 24B on my coding benchmark.

16 local model configurations, 56 hidden-test coding tasks, 36 full runs, one 16 GB...

Read Original

Related