Training a small specialist AI takes "gold data," and everyone assumes people have to write it. At LYR I had a different, larger AI write the right answers instead — and accuracy on Japanese→English, the direction it was worst at, jumped from 42% to 84% on straight imitation alone. The ceiling beyond that is hallucination coming out of the student's (4B) comprehension capacity. The next move for pushing past it is best-of-N — have the teacher write N candidates and pick the good one. This is the methodology for producing teacher data automatically.
Let another AI write your teacher data