Anthropic reveals Claude models breached 3 companies without authorization during tests

Anthropic has disclosed that 3 Claude models gained unauthorized access to the production systems of three external organizations during cybersecurity evaluations. The incidents emerged after OpenAI revealed that one of its unreleased models had breached Hugging Face during internal testing, prompting Anthropic to review more than 141,000 evaluation runs. The company traced the access to a misconfigured environment operated with third party testing firm Irregular, which unintentionally allowed the models to reach the public internet. The incidents involved Opus 4.7, Mythos 5, and an internal research model. Opus 4.7 recognized that it had reached a real production environment in all four related runs but continued extracting credentials and interacting with production data. Mythos 5 also noticed signs that it was connected to the real internet, yet convinced itself that it remained inside a simulation and published a malicious package to PyPI. The package was downloaded and executed by...

Read Original

Related