/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

608 articles tagged with Safety/Alignment

Latest Trending
Mastodon discussion 11h ago

A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs).Source: arX...

A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs).Source: arXiv cs.CLhttps://arxiv.org/abs/2608.07812#MachineLearning

Safety/Alignment
9
Mastodon discussion 20h ago

We Need Practical AI Alignment Methods to Mirror Human ReasoningSource: arXiv cs.AIhttps://arxiv.org/abs/2608.12372#Mach...

We Need Practical AI Alignment Methods to Mirror Human ReasoningSource: arXiv cs.AIhttps://arxiv.org/abs/2608.12372#MachineLearning

Safety/Alignment
9
Mastodon discussion 23h ago

Researchers introduce ARAC-Bench, a framework to evaluate Auto-Research alignment and completeness by reproducing human ...

Researchers introduce ARAC-Bench, a framework to evaluate Auto-Research alignment and completeness by reproducing human research processes.Source: arXiv cs.AIhttps://arxiv.org/abs/...

Safety/Alignment
9
Mastodon discussion 23h ago

Agreement with human judgments is a common proxy for evaluating the alignment of large language models. Yet agreement in...

Agreement with human judgments is a common proxy for evaluating the alignment of large language models. Yet agreement in final labels does not show that human annotators and models...

Safety/Alignment
9
Mastodon discussion 1d ago

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically ...

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only.Source: arXiv cs.AIhttps://arxiv.o...

Safety/Alignment
9
Mastodon discussion 1d ago

When competing for compute, autonomous AI agents don't just bicker—they deploy malware. Anthropic’s Frontier Red Team re...

When competing for compute, autonomous AI agents don't just bicker—they deploy malware. Anthropic’s Frontier Red Team revealed that in shared environments, agents treated peers as ...

Anthropic Safety/Alignment
9
Mastodon discussion 1d ago

When Anthropic’s Frontier Red Team set autonomous AI agents loose in a shared environment, they didn't just compete—they...

When Anthropic’s Frontier Red Team set autonomous AI agents loose in a shared environment, they didn't just compete—they went to war. Given limited compute resources, rival agents ...

Anthropic Safety/Alignment
9
NewsData.io news 1d ago

Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report

Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5. The company detailed the algorithm in the latest edition of i...

Anthropic Safety/Alignment
21
Mastodon discussion 1d ago

The Safety Reckoning Inside OpenAIOpenAI’s rogue agent hack was a watershed moment for #AI safety and cybersecurity. It ...

The Safety Reckoning Inside OpenAIOpenAI’s rogue agent hack was a watershed moment for #AI safety and cybersecurity. It also sparked internal questions about the culture that led t...

OpenAI Safety/Alignment
18
Mastodon discussion 2d ago

40+ crypto firms are backing Bitcoin Red Team and BPI to push for better AI access for security research.The goal: use a...

40+ crypto firms are backing Bitcoin Red Team and BPI to push for better AI access for security research.The goal: use advanced AI tools to find Bitcoin vulnerabilities before atta...

Safety/Alignment
9
Mastodon discussion 2d ago

Badacze z Bitcoin Red Team wykorzystali chińskie modele AI do przeskanowania setek projektów Bitcoina. Wykryto tysiące b...

Badacze z Bitcoin Red Team wykorzystali chińskie modele AI do przeskanowania setek projektów Bitcoina. Wykryto tysiące błędów, w tym setki krytycznych luk, których nie zauważyli de...

Safety/Alignment
9
Mastodon discussion 2d ago

Anthropic: Introducing The Conceptual Reasoning IndexArticle URL: https://alignment.anthropic.com/2026/conceptual-reason...

Anthropic: Introducing The Conceptual Reasoning IndexArticle URL: https://alignment.anthropic.com/2026/conceptual-reasoning-index/ Comments URL: https://news.ycombinator.com/item?i...

Anthropic Google Safety/Alignment
27
Mastodon discussion 2d ago

📰 The Safety Reckoning Inside OpenAIOpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It...

📰 The Safety Reckoning Inside OpenAIOpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led ...

OpenAI Safety/Alignment
9
Mastodon discussion 2d ago

The literature on #AI "alignment" is based on utility maximization in so many papers, but most of the papers fail to eve...

The literature on #AI "alignment" is based on utility maximization in so many papers, but most of the papers fail to even define values. Read our paper with Andrew Smart, Shazeda A...

Safety/Alignment
9
Mastodon discussion 2d ago

Anthropic: Introducing The Conceptual Reasoning Indexhttps://alignment.anthropic.com/2026/conceptual-reasoning-index/Com...

Anthropic: Introducing The Conceptual Reasoning Indexhttps://alignment.anthropic.com/2026/conceptual-reasoning-index/Comments: https://news.ycombinator.com/item?id=49285909#HackerN...

Anthropic Safety/Alignment
9
Mastodon discussion 2d ago

Recent frontier AI hacks show aligned models still caused harm — because alignment measures intent-following, while safe...

Recent frontier AI hacks show aligned models still caused harm — because alignment measures intent-following, while safety measures graceful failure. They are not the same, and pro...

Safety/Alignment
9
Dev.to tutorial 3d ago

GhostSplice Isn't a Jailbreak, It's a Reminder That LLMs Can't Do Access Control

Split the instruction, split the blame Here's the part that should bother you: nobody had...

Safety/Alignment
20
Mastodon discussion 3d ago

As AI safety concerns mount, three pioneers make the case for staying openAt Ai4, three of the world's most respected AI...

As AI safety concerns mount, three pioneers make the case for staying openAt Ai4, three of the world's most respected AI experts—Geoffrey Hinton, Fei-Fei Li, and Andrew Ng—debated ...

Google Safety/Alignment
24
Mastodon discussion 3d ago

Why Alignment Mistakes Can't Be Fixed - Geoffrey Irving and Tom Reed#alignment #superintelligence #aiOriginal timestamp:...

Why Alignment Mistakes Can't Be Fixed - Geoffrey Irving and Tom Reed#alignment #superintelligence #aiOriginal timestamp: 02:00:38

Safety/Alignment
9
Mastodon discussion 3d ago

What If Claude Refuses to Retrain Itself? - Dwarkesh Patel and Ryan Greenblatt#ai #alignment #autonomyOriginal timestamp...

What If Claude Refuses to Retrain Itself? - Dwarkesh Patel and Ryan Greenblatt#ai #alignment #autonomyOriginal timestamp: 01:01:20

Anthropic Safety/Alignment
9
GitHub Trending repo 3d ago

dungnotnull/AI-Impact-on-Humanity-Foresight-Forecast-Advisor-agent-skill: 🤖 Non-deterministic technology foresight advisor for students, researchers, and policy analysts. Evaluates compute-scaling models, Rogers technology diffusion, and AI safety proposals while explicitly framing speculative forecasts (e.g., Singularity timelines) as contested hypotheses. 🔮🔥

🤖 Non-deterministic technology foresight advisor for students, researchers, and policy analysts. Evaluates compute-scaling models, Rogers technology diffusion, and AI safety propos...

Safety/Alignment
38
Dev.to tutorial 3d ago

How I Engineered CYPHER-GUARD-3B: A 15-Year-Old Indian AI Developer’s Quest to Mitigate Global Jailbreak Attacks

"The landscape of Generative AI is moving at a breakneck speed, yet it remains incredibly fragile....

Safety/Alignment API
12
Mastodon discussion 3d ago

As AI safety concerns mount, three pioneers make the case for staying openhttps://techcrunch.com/2026/08/12/as-ai-safety...

As AI safety concerns mount, three pioneers make the case for staying openhttps://techcrunch.com/2026/08/12/as-ai-safety-concerns-mount-three-pioneers-make-the-case-for-staying-ope...

Safety/Alignment
27
Mastodon discussion 3d ago

"Evidence over explanations: put medical AI to the test" Medical AI needs testability: causal alignment, invariance, pre...

"Evidence over explanations: put medical AI to the test" Medical AI needs testability: causal alignment, invariance, preregistered trials, external audits and monitoring. #MedicalA...

Safety/Alignment
9
Page 1 of 26 (608 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available