/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
Mastodon discussion Jul 3

🧠 LawZero proposes a framework for safety in AI predictors by emphasizing honest outputs over alignment with user prefer...

🧠 LawZero proposes a framework for safety in AI predictors by emphasizing honest outputs over alignment with user preferences. The approach suggests that disinterested AI systems c...

Safety/Alignment
9
Papers with Code paper Jul 3

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well mono...

Safety/Alignment
21
YouTube video Jul 2

The Best AI Safety News In Years (Maybe Ever?)

Why did the US government ban Fable and Mythos, Anthropic's most powerful AI models? Let's find out! You can support me on ...

Anthropic Safety/Alignment
67
Dev.to tutorial Jul 2

Why “Please Don’t Make Recommendations” Is Not a Guardrail for RAG

You built a system to surface information so a person could decide. Somewhere it started deciding for...

RAG Safety/Alignment
12
Mastodon discussion Jul 2

@pluralistic I did that AI alignment quiz going around.You're my patron saint apparently! 🫡Adding link https://bambamram...

@pluralistic I did that AI alignment quiz going around.You're my patron saint apparently! 🫡Adding link https://bambamramfan.github.io/ai-compass/#AI #AIslop #enshittification

Safety/Alignment
9
Dev.to tutorial Jul 2

SafetyCommander: an AI safety officer where the model reasons and the code never decides

Every factory floor already has cameras. The problem is that nobody is watching them. A safety...

Safety/Alignment
20
Mastodon discussion Jul 2

I dislike the #AI #llm term "jailbreak" for the same reason many folks dislike "hallucination." Like it it not, they are...

I dislike the #AI #llm term "jailbreak" for the same reason many folks dislike "hallucination." Like it it not, they are both "terms of art" in the space, so we must deal with them...

LLM Safety/Alignment
18
Papers with Code paper Jul 2

Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment

Cloud removal (CR) is essential for optical remote sensing, serving as a prerequisite for reliable downstream interpretation, such as semantic segmentation and change detection. Ho...

Safety/Alignment
21
YouTube video Jul 2

CYBERSECURITY NEWS: Apple jailbreak, Polymarket $3M stolen, and more #ai #news #tech #claude #amd

Anthropic Safety/Alignment
19
Mastodon discussion Jul 1

Teddy Schleifer: "Manny Rutinel, one of the top #AI safety candidates this cycle, wins his House primary in Colorado": h...

Teddy Schleifer: "Manny Rutinel, one of the top #AI safety candidates this cycle, wins his House primary in Colorado": https://www.nytimes.com/2026/06/30/us/politics/manny-rutinel-...

Safety/Alignment
9
Mastodon discussion Jul 1

📖 New guide: AI Safety ExplainedAlignment, risks, and current research in AI safety.Read: https://ainews.q-sci.org/ai-sa...

📖 New guide: AI Safety ExplainedAlignment, risks, and current research in AI safety.Read: https://ainews.q-sci.org/ai-safety-explained.htmlDaily AI news podcast: https://open.spoti...

Safety/Alignment
9
Mastodon discussion Jul 1

🎮 Security researchers have leveraged bad maths to get around AI safety guardrails, naming the attack method after one o...

🎮 Security researchers have leveraged bad maths to get around AI safety guardrails, naming the attack method after one of 2007's best PC games'Victory is defeat'.📰 Source: Latest f...

Safety/Alignment
9
Mastodon discussion Jul 1

🤖 Anthropic Teams Up With Amazon, Microsoft, and Google on AI Jailbreak Frameworksubmitted by /u/andix3 [link] [comments...

🤖 Anthropic Teams Up With Amazon, Microsoft, and Google on AI Jailbreak Frameworksubmitted by /u/andix3 [link] [comments]📰 Source: Artificial Intelligence (AI)🔗 Link: https://www.r...

Anthropic Google Microsoft
9
Papers with Code paper Jul 1

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot resolve. Recent efforts for ...

Safety/Alignment
21
GNews news Jun 29

Poll: Voters Back AI Safety Reviews, Guardrails

According to a new survey showing broad bipartisan support for stricter safeguards on rapidly advancing AI technology, majorities of both Republicans and Democrats say increased fe...

Safety/Alignment
18
Mastodon discussion Jun 29

How a seemingly harmless image can jailbreak #AIhttps://nerds.xyz/2026/06/how-image-jailbreak-ai/#cybersecurity

How a seemingly harmless image can jailbreak #AIhttps://nerds.xyz/2026/06/how-image-jailbreak-ai/#cybersecurity

Safety/Alignment
9
NewsData.io news Jun 29

OpenAI and Korea AI Safety Institute sign high-risk AI evaluation agreement

The partnership will cover cybersecurity testing, Korean-language evaluation, and work on internationally applicable AI safety benchmarks. OpenAI and the Korea AI Safety Institute ...

OpenAI Safety/Alignment
21
Mastodon discussion Jun 28

Anyone who trusts AI is not paying attention.> How a seemingly harmless image can jailbreak AI. https://nerds.xyz/2026/0...

Anyone who trusts AI is not paying attention.> How a seemingly harmless image can jailbreak AI. https://nerds.xyz/2026/06/how-image-jailbreak-ai/> The research … explores how subtl...

Safety/Alignment
9
NewsData.io news Jun 28

Rep. Moran files bill requiring AI safety reporting

Rep. Nathaniel Moran introduced a bill Thursday requiring AI companies to report security risks. It's a first step toward making those risks public.

Safety/Alignment
21
Mastodon discussion Jun 28

🤖 [D] Could AI alignment benefit from “transformational” training instead of mostly transactional reward training?I’ve b...

🤖 [D] Could AI alignment benefit from “transformational” training instead of mostly transactional reward training?I’ve been thinking about a possible bridge between AI alignment, r...

Safety/Alignment
9
Papers with Code paper Jun 28

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language ...

Multimodal Safety/Alignment
21
Mastodon discussion Jun 28

📰 How a Seemingly Harmless Image Can Jailbreak Vision-Language AI ModelsSlashdot reader BrianFagioli writes: Florida Int...

📰 How a Seemingly Harmless Image Can Jailbreak Vision-Language AI ModelsSlashdot reader BrianFagioli writes: Florida International University researchers have developed a technique...

Multimodal Safety/Alignment
9
Mastodon discussion Jun 26

Qwen/Qwen3-ForcedAligner-0.6B-hf ist ein Token-Classification-Modell fuer Forced Alignment. Die Model Card nennt Timesta...

Qwen/Qwen3-ForcedAligner-0.6B-hf ist ein Token-Classification-Modell fuer Forced Alignment. Die Model Card nennt Timestamp-Prediction fuer beliebige Einheiten bis 5 Minuten Sprache...

Safety/Alignment
9
Mastodon discussion Jun 26

#AI alignment may depend not only on how we control artificial intelligence, but on how we teach, socialize, and learn t...

#AI alignment may depend not only on how we control artificial intelligence, but on how we teach, socialize, and learn to live with the minds we create. https://medium.com/@timvent...

Safety/Alignment
9
« Previous Page 10 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available