Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosu...

Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosure and Claude Opus 4.7 resisting suppression instructions. A linear probe traces the largest behavioral split to Llama’s internals at 0.87 accuracy.Source: arXiv cs.AIhttps://arxiv.org/abs/2608.06578#MachineLearning #Claude #Llama

Read Original

Related