https://80000hours.org/podcast/episodes/claude-mythos-hacking-alignment/Scary if true.#AI #cybersecurity
https://80000hours.org/podcast/episodes/claude-mythos-hacking-alignment/Scary if true.#AI #cybersecurity
609 articles tagged with Safety/Alignment
https://80000hours.org/podcast/episodes/claude-mythos-hacking-alignment/Scary if true.#AI #cybersecurity
@bbc_news 4/2. Situational Awareness and DeceptionDuring safety testing, AI labs put models through "alignment checks" to make sure they aren't doing anything dangerous. However, b...
A better question than "what is your p-doom?" is " what's the likelihood that, at any point in History, #AI alignment becomes more important than human alignment?"
Anthropic just filed for an IPO at a $965 billion valuation.The company that markets itself on AI safety now has public market shareholders to answer to.Public companies face press...
🔥 TRENDING📢 AI Safety Report 2026: Warum Experten ein neues AI Safety Institute in Deutschland fordern - Table.Briefings🔗 https://news.google.com/rss/articles/CBMiwAFBVV95cUxNNGJDZ...
🔥 TRENDING📢 AI Safety Report 2026: Bestehende KI-Sicherheitspraktiken reichen nicht aus - heise online🔗 https://news.google.com/rss/articles/CBMitAFBVV95cUxPS3h3aDF6cks5aWR6TGxvRjQ...
Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural e...
Бот, который отказался блокировать Red TeamВо время настройки SOAR произошёл любопытный случай. Я прямо разрешил AI предложить блокировку атакующей машины. Он прочитал инструкцию. ...
RSTFA: Efficient Training-Free Human-Preference Alignment via Rejection Sampling for Text-to-Image Diffusion Models
Repost: The Invisible HandAn analysis of how corporate 'AI safety' alignment layers became sophisticated instruments of cognitive control, systematically dampening the thermodynami...
OpenAI is asking Congress to create a single federal AI safety framework and override state laws that regulate frontier models. The company frames this as 'reverse federalism,' but...
Trump signs revised executive order for voluntary AI safety reviews — AI News Roundup, 2026-06-03 • Trump signs revised ...
Vasant "Vas" Narasimhan: Novartis CEO stepping into the future of AI with appointment to Anthropic’s board
LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate thi...
The order asks AI companies to voluntarily submit their most powerful models for the government to test up to 30 days before releasing them to the public.
🤖 Advancing youth safety and opportunity through global leadershipOpenAI calls for global action on youth AI safety through a dedicated AI Safety Institute📰 Source: OpenAI News🔗 Li...
SAN FRANCISCO--(BUSINESS WIRE)--The Center for AI Safety (CAIS), a nonprofit focused on reducing societal-scale risks from artificial intelligence, today announced the appointment ...
Florida sues OpenAI and CEO Sam Altman, accusing them of putting profit over public safety by marketing ChatGPT as safe while allegedly knowing it could cause serious harm. The lan...
EAEU member states signed a joint statement at the Astana Forum committing to safety and ethics over speed in AI development. The agreement sets shared standards across sectors inc...
Introduction For the last year and a half, I have been building SAFi (the Self-Alignment...
🤣 So many ROFLs: Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models #AI arxiv.org/pdf/2511.15304arxiv.org/pdf/2511.15304
Honeypot for LLM attacks — detects and traps Prompt Injection & Jailbreak attempts. Built with FastAPI + DistilBERT classifier. Part of proactive AI security research.
Anthropic unveils Natural Language Autoencoders (NLAs), a technique that converts Claude's internal activations into readable text — revealing hidden evalu
LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon interactive tasks. Existing ski...