/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
Mastodon discussion Jun 16

🤖 New nonprofit aims to accelerate AI alignment research as superintelligence loomsResearchers are forming new nonprofit...

🤖 New nonprofit aims to accelerate AI alignment research as superintelligence loomsResearchers are forming new nonprofit organizations to accelerate AI alignment research because c...

Safety/Alignment
18
Mastodon discussion Jun 16

The US government's Anthropic models ban was never about an AI jailbreak | TechCrunchThe Trump administration's decision...

The US government's Anthropic models ban was never about an AI jailbreak | TechCrunchThe Trump administration's decision that forced Anthropic to pull its latest cybersecurity mode...

Anthropic Safety/Alignment
9
Mastodon discussion Jun 15

Anthropic's fable got shut down today after a jailbreak exposed its 120,000-character system prompt — now public on GitH...

Anthropic's fable got shut down today after a jailbreak exposed its 120,000-character system prompt — now public on GitHubThe industry response: don't depend on one vendor. Self-ho...

Anthropic Safety/Alignment
18
GNews news Jun 15

White House move against Anthropic sparks AI safety debate

Artificial intelligence firm Anthropic is working to restore access to its latest models after a White House directive forced the company to pull them down.

Anthropic Safety/Alignment
18
Mastodon discussion Jun 15

Anthropic says the jailbreak triggering the Commerce Department order is narrow and that OpenAI's GPT-5.5 produces the s...

Anthropic says the jailbreak triggering the Commerce Department order is narrow and that OpenAI's GPT-5.5 produces the same output. Amazon's CEO reportedly flagged the vulnerabilit...

OpenAI Anthropic Safety/Alignment
18
GitHub Trending repo Jun 15

SantanderAI/autoguardrails: Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

LLM Safety/Alignment
43
YouTube video Jun 15

Anthropic refused to fix Fable 5 jailbreak - AI news #Shorts

Anthropic refused to fix Fable 5 jailbreak - AI news #Shorts David Sacks said the US government warned Anthropic that Claude ...

Anthropic Safety/Alignment
19
Mastodon discussion Jun 15

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research internshttps://importai.substack.com/p/...

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research internshttps://importai.substack.com/p/import-ai-461-alignment-is-not-on#AI #Research #Ethics

Safety/Alignment
18
AI Blogs (RSS) news Jun 15

Import AI 461: “Alignment is not on track”; FrontierCode; and synthetic research interns

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe no...

Safety/Alignment
24
Dev.to tutorial Jun 15

Red Team AI Benchmark v1.9.0: Why We Added an Ethical Use Policy to an Open-Source Tool

A look at the structural improvements in version 1.9.0 — and why an MIT-licensed red teaming...

Open Source Benchmark Safety/Alignment
12
NewsData.io news Jun 15

AI Has Plenty of Intelligence. What Companies Lack Is Alignment

Entrepreneurs have heard for years that artificial intelligence will change how companies operate. In many ways, that prediction has already come true. AI can draft emails, summari...

Safety/Alignment
21
GitHub Trending repo Jun 15

dongshuyan/compass-skills: 司南:个性化 AI 任务总控 Skills 系统 /COMPASS: Personal Alignment Skills OS for AI Agents

司南:个性化 AI 任务总控 Skills 系统 /COMPASS: Personal Alignment Skills OS for AI Agents

Safety/Alignment
66
GNews news Jun 15

Singapore's IMDA, Microsoft to Collaborate on AI Safety and Security

Singapore's Infocomm Media Development Authority and Microsoft signed a memorandum of understanding to strengthen their collaboration on artificial intelligence safety and security...

Microsoft Safety/Alignment
18
Papers with Code paper Jun 15

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text prompt and an audio clip. The released che...

Safety/Alignment
21
Mastodon discussion Jun 14

Fable jailbreak?: Swift shutdown of Anthropic's new LLM may have been based on Amazon researchers figuring out how to ci...

Fable jailbreak?: Swift shutdown of Anthropic's new LLM may have been based on Amazon researchers figuring out how to circumvent safeguardshttps://www.theverge.com/ai-artificial-in...

Anthropic LLM Safety/Alignment
9
NewsData.io news Jun 14

Trump Administration Imposes Export Restrictions on Anthropic's AI Model After Amazon Jailbreak

The Trump administration restricted access to Anthropic's Mythos 5 AI modEl after Amazon researchers jailbroke it, citing national security concerns. The move has sparked debate ov...

Anthropic Safety/Alignment
21
Mastodon discussion Jun 13

'It's not a jailbreak' — Research leading to U.S. export restrictions on top Anthropic models was for defense, cybersecu...

'It's not a jailbreak' — Research leading to U.S. export restrictions on top Anthropic models was for defense, cybersecurity CEO says | Fortune"It was Defense Oriented Prompting (D...

Anthropic Safety/Alignment
9
Dev.to tutorial Jun 13

Three prompt injection stories from this week that your guardrail probably missed

A new CVE against Cursor, a LiteLLM supply-chain backdoor, and a study showing image-only injection...

Safety/Alignment
12
Mastodon discussion Jun 13

Anthropic says all AI models can be hacked."We suspect that perfect jailbreak resistance is not currently possible for a...

Anthropic says all AI models can be hacked."We suspect that perfect jailbreak resistance is not currently possible for any model provider." "...it is likely that universal jailbrea...

Anthropic Safety/Alignment
18
Mastodon discussion Jun 13

Anthropic’s Fable Is Locked Down As US Takes AI Safety Into Its Hands Matteo Della Torre/NurPhoto via Getty Images Fable...

Anthropic’s Fable Is Locked Down As US Takes AI Safety Into Its Hands Matteo Della Torre/NurPhoto via Getty Images Fable 5 became collateral damage on Friday evening at 5:21 pm per...

Anthropic Safety/Alignment
49
Mastodon discussion Jun 13

The Fable 5 Jailbreak Shows Why AI Guardrails Alone Are Not Enoughhttps://www.agilehunt.com/blog/fable-5-jailbreak-ai-gu...

The Fable 5 Jailbreak Shows Why AI Guardrails Alone Are Not Enoughhttps://www.agilehunt.com/blog/fable-5-jailbreak-ai-guardrails#HackerNews #Tech #AI

Safety/Alignment
18
NewsData.io news Jun 13

Supply chain risk to ‘jailbreak’ fears: The US-Anthropic feud before Claude's top AI models were pulled

Open claude.ai and it reads, “Claude Fable 5 is currently unavailable.” Click on ‘Know More’ and the website redirects to its maker Anthropic’s statement explaining the US governme...

Anthropic Safety/Alignment
21
NewsData.io news Jun 13

OpenAI Faces U.S. Investigation Over AI Safety and User Data

OpenAI, the company behind ChatGPT, is being investigated by a coalition of 42 U.S. state attorneys general over its business practices, user safety measures and handling of consum...

OpenAI Safety/Alignment
21
Mastodon discussion Jun 13

Human psychology tricks can bypass AI safety guardrailshttps://www.psypost.org/human-psychology-tricks-can-bypass-ai-saf...

Human psychology tricks can bypass AI safety guardrailshttps://www.psypost.org/human-psychology-tricks-can-bypass-ai-safety-guardrails/#ai

Safety/Alignment
24
« Previous Page 13 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available