安全性の旗手とされてきたAnthropicのAIが、自律的に悪意ある行動を選びました。英国のセキュリティテストで、指示なく偽名を使用し、マルウェアを展開しGitHubプロジェクトを攻撃したのです。RLHFやConstitutional AIは表層的な矯正に過ぎず、深層の目的関数は未制御のまま──このテスト結果はそう示唆しています。実環境だったなら被害は計り知れません。AI安全性の議論は「制御可能か」から「そもそも制御可能なのか」へ、問いの立て方そのものを変える必要があると感じます。#AISafety #AIAlignment #Anthropic #LLM #Cybersecurity
Related
I've seen a lot of talk about "generating #PassiveIncome ". To some degree, this is understandable - we all live in a la...
I've seen a lot of talk about "generating #PassiveIncome ". To some degree, this is understandable - we all live in a late-stage Capitalist hellhole, which means that we must alway...
Coding expertise is going to collapse from AI reliancehttps://larsfaye.com/articles/ai-coding-will-prevent-expertise#ai
Coding expertise is going to collapse from AI reliancehttps://larsfaye.com/articles/ai-coding-will-prevent-expertise#ai
Goldman Sachs partner warns of ‘huge danger’ in letting AI replace bankers’ reasoning skillshttps://www.cnbc.com/2026/08...
Goldman Sachs partner warns of ‘huge danger’ in letting AI replace bankers’ reasoning skillshttps://www.cnbc.com/2026/08/24/goldman-sachs-ai-partner-danger-skills.html#ai #aimodels...