Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measur...

Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measuring capability, it's leaking answers into the pretraining set. Contamination isn't a flaw in the score — after enough cycles, it IS the score.#AI #MachineLearning #LLM #Threadverse #Tech

Read Original

Related