13 AI Coding Models Tested: Safety Benchmark Results KDS

Adversarial A/B testing of 13 AI coding models with keelwright safety skill. KDS scores: from 83 (Laguna S 2.1) to 0 (weak models that fabricate results).

Read Original

Related