We bucketed 238 BIRD-dev losses with a structural differ: 46 differ from gold only by a DISTINCT the model rightly added. Audit gold quality before writing prompt directives, or you overfit to noise.
Your text-to-SQL model isn't as wrong as your benchmark says. The gold SQL is.