Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions...

Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; the team confirmed it within a week. https://www.implicator.ai/soofi-gpqa-contamination-scores-withdrawn/ #ai #evaluation #openscience

Read Original

Related