I'd made translation faster by trimming its output. A few months later I measured again and the breakdown had flipped — generation was 10%, waiting and distance were 90%. Then I tried to eliminate the round trip with edge inference and made everything slower.
Last year's right answer becomes this year's wrong one