/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
NewsData.io news Jun 20

White House, Anthropic Move Toward AI Safety Rulebook After Clash Over Powerful New Models

The White House and Anthropic are working to establish what could become one of the first formal frameworks for evaluating security vulnerabilities in advanced artificial intellige...

Anthropic Safety/Alignment
21
Mastodon discussion Jun 19

AI Regulation Should Be Rational, Not RetaliatoryThe Trump administration’s approach to AI safety, particularly the gene...

AI Regulation Should Be Rational, Not RetaliatoryThe Trump administration’s approach to AI safety, particularly the generative AI models that regularly grab headlines, has been hap...

Anthropic Safety/Alignment
9
Dev.to tutorial Jun 19

Your LLM guardrail speaks English. Your attacker doesn't.

I found this out the embarrassing way by " _attacking my own system _". I maintain FIE, an...

LLM Safety/Alignment
20
Papers with Code paper Jun 19

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an imp...

LLM Safety/Alignment
21
Mastodon discussion Jun 18

⚖️ AI Regulation Should Be Rational, Not RetaliatoryThe Trump administration’s approach to AI safety, particularly the g...

⚖️ AI Regulation Should Be Rational, Not RetaliatoryThe Trump administration’s approach to AI safety, particularly the generative AI models that regularly grab headlines, has been ...

Safety/Alignment
18
Mastodon discussion Jun 18

Jailbreak #AI — from House of El. Griezelige werkelijkheid.https://youtu.be/R4nFEQb7kZo

Jailbreak #AI — from House of El. Griezelige werkelijkheid.https://youtu.be/R4nFEQb7kZo

Safety/Alignment
18
Dev.to tutorial Jun 18

ASR-generated subtitles vs forced alignment: why script-first captions fail less

A mistake I keep seeing in subtitle tools is simple but expensive: someone already has an approved...

Safety/Alignment
12
Dev.to tutorial Jun 18

I put 6 LLM guardrail tools inline and measured what they cost me. Here is the latency-vs-recall tradeoff.

An input guardrail runs on every request. Too slow and you rip it out; fast but blind and you get...

LLM Safety/Alignment
12
Mastodon discussion Jun 18

El lado del mal - Hacking AI: Jailbreak, Prompt Injection, Hallucinations & Misalignment. How to Hack Digital Services B...

El lado del mal - Hacking AI: Jailbreak, Prompt Injection, Hallucinations & Misalignment. How to Hack Digital Services Based on LLMs & AI Agents (English Edition) https://www.ellad...

Safety/Alignment
27
Dev.to tutorial Jun 18

LLM Prompt Injection & Guardrail Security

A recall reference built from working through a 7-layer prompt-injection challenge. Focus: how each...

LLM Safety/Alignment
12
Mastodon discussion Jun 18

The Commerce Department ordered Anthropic to shut down Fable 5 and Mythos 5 after a disputed jailbreak finding. The core...

The Commerce Department ordered Anthropic to shut down Fable 5 and Mythos 5 after a disputed jailbreak finding. The core tension: governments can block model access, but cannot man...

Anthropic Safety/Alignment
18
YouTube video Jun 17

US administration demands uncircumventable jailbreak protections from Anthropic | AI News Roundup

US administration demands uncircumventable jailbreak protections from Anthropic — AI News Roundup, 2026-06-17 • US ...

Anthropic Safety/Alignment
15
GNews news Jun 17

China pushes for AI safety as G7 summit wraps up without Beijing

China included artificial intelligence in its global governance whitepaper published Wednesday, and said Beijing was speeding up cooperation efforts.

Safety/Alignment
18
Mastodon discussion Jun 17

Great paper on how research in adversarial #ML and #AI safety has stalled for current-gen LLMs:https://arxiv.org/abs/250...

Great paper on how research in adversarial #ML and #AI safety has stalled for current-gen LLMs:https://arxiv.org/abs/2502.02260#Research #AISafety #LLM #Security

Safety/Alignment
24
Papers with Code paper Jun 17

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we i...

Safety/Alignment Robotics
21
Mastodon discussion Jun 16

unitedforclimate.blogspot.com/2024/10/lets... general erosion of trust and alignment with Putin's interests. NOTE: Verif...

unitedforclimate.blogspot.com/2024/10/lets... general erosion of trust and alignment with Putin's interests. NOTE: Verify AI-generated content critically #AI #Perplexity #DeepAI #C...

Safety/Alignment
18
Mastodon discussion Jun 16

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcherAccording to the one person...

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcherAccording to the one person who actually read the research paper#AI https://www.theregi...

Safety/Alignment
18
Mastodon discussion Jun 16

CISA SCuBA Gear M365ScubaGear is a no-cost assessment tool that verifies M365 tenant configuration alignment to the poli...

CISA SCuBA Gear M365ScubaGear is a no-cost assessment tool that verifies M365 tenant configuration alignment to the policies described in SCuBA's secure configuration baselines. #C...

Safety/Alignment
30
Mastodon discussion Jun 16

AI safety critic fired right before massive IPO - I'll give you three guesses which trillionaire was involved and the fi...

AI safety critic fired right before massive IPO - I'll give you three guesses which trillionaire was involved and the first two don't count. https://techcrunch.com/2026/06/10/xai-f...

Safety/Alignment
30
Mastodon discussion Jun 16

🤖 AI safety debate shifts toward resilience over regulationPolicymakers are increasingly focusing on improving AI resili...

🤖 AI safety debate shifts toward resilience over regulationPolicymakers are increasingly focusing on improving AI resilience rather than implementing extraordinary government inter...

Safety/Alignment
18
Hacker News discussion Jun 16

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak

Safety/Alignment
66
Mastodon discussion Jun 16

🤖 Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailb...📝 The “jailbreak”...https://www.theregister...

🤖 Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailb...📝 The “jailbreak”...https://www.theregister.com/security/2026/06/15/feds-freaked-over-fable-5-after-sim...

Safety/Alignment
9
Dev.to tutorial Jun 16

RLHF vs DPO vs IPO vs KTO: which alignment method should you use

A practical comparison of RLHF, DPO, IPO, and KTO — what each method actually does under the hood, how their data and compute requirements differ, and when to pick one over the oth...

Safety/Alignment
12
Mastodon discussion Jun 16

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcherAccording to the one person...

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcherAccording to the one person who actually read the research paperhttps://www.theregister...

Safety/Alignment
9
« Previous Page 12 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available