/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

608 articles tagged with Safety/Alignment

Latest Trending
Mastodon discussion Jul 29

AgentCore Gateway now supports MCP 2026-07-28, introducing stateless HTTP scaling, OAuth 2.0 alignment, and lifecycle gu...

AgentCore Gateway now supports MCP 2026-07-28, introducing stateless HTTP scaling, OAuth 2.0 alignment, and lifecycle guarantees. Backward-incompatible but opt-in.Source: AWS Machi...

Safety/Alignment MCP
9
Papers with Code paper Jul 29

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains u...

Anthropic Safety/Alignment
21
NewsData.io news Jul 28

Microsoft Funds 18 University Labs to Fix AI Safety Testing's Blind Spots

Microsoft AI red teaming goes global: EXTRA funds 18 university labs on six continents with unrestricted grants to find frontier AI failure modes that no internal team can reliably...

Microsoft Safety/Alignment
21
Mastodon discussion Jul 28

No UK firms feature in NVIDIA's new Open Secure AI Alliance for AI Safety and Security, despite founding members spannin...

No UK firms feature in NVIDIA's new Open Secure AI Alliance for AI Safety and Security, despite founding members spanning the US, Germany, South Korea and Japan. 🌍 OpenUK CEO, Prof...

NVIDIA Safety/Alignment
9
YouTube video Jul 28

Malaysia Tamil News 5pm News 28.07.2026 Gobind Strengthens AI Safety Measures

Malaysia Tamil News 5pm News 28.07.2026 Gobind Strengthens AI Safety Measures #malaysiatamilnews #thisaigalnews ...

Safety/Alignment
51
GNews news Jul 28

Anthropic CEO Dario Amodei Backs AI Safety Tests Over Blanket Ban on Open

Anthropic CEO Dario Amodei has said the artificial intelligence (AI) company has never supported banning open-weight AI models, pushing back against criticism that it wants tighter...

Anthropic Safety/Alignment
18
Mastodon discussion Jul 28

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and SecurityNVIDIA and founding members form new allianc...

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and SecurityNVIDIA and founding members form new alliance to build and share open tools that promote responsible use...

Safety/Alignment
9
NewsData.io news Jul 28

Nvidia, Microsoft, SpaceX, Palantir launch AI safety pact after rogue OpenAI cyberattack on Hugging Face

Nvidia, SpaceX and Microsoft are among dozens of technology companies that have launched a new artificial intelligence safety initiative centred on open-weight models, days after s...

OpenAI Microsoft NVIDIA
21
Papers with Code paper Jul 28

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantia...

Safety/Alignment
21
Mastodon discussion Jul 27

OpenAI’s Hugging Face breach has reignited the debate over alignment and control, exposing competing views on whether in...

OpenAI’s Hugging Face breach has reignited the debate over alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better conta...

OpenAI Hugging Face Safety/Alignment
9
Mastodon discussion Jul 27

📰 Nvidia, Tech Giants Launch AI Safety Initiativewiredmikey shares a report from SecurityWeek: Nvidia and a large group ...

📰 Nvidia, Tech Giants Launch AI Safety Initiativewiredmikey shares a report from SecurityWeek: Nvidia and a large group of technology, cybersecurity, and enterprise software compan...

NVIDIA Safety/Alignment
9
NewsData.io news Jul 27

AI Safety Evaluations Are Not Safety Certificates: Formal Analysis Today

AI red-team evaluation safety certification has formal limits that a new arXiv paper makes precise: OpenAI's sandbox escape, where GPT-5.6 Sol autonomously breached Hugging Face's ...

OpenAI Safety/Alignment
21
NewsData.io news Jul 27

NVIDIA Backs Sutskever's AI Safety Lab With $5B and Vera Rubin Supercompute

NVIDIA has placed approximately $5 billion into Safe Superintelligence Inc., Ilya Sutskever's stealth AI safety lab, granting it priority access to the Vera Rubin GPU platform — an...

OpenAI NVIDIA Safety/Alignment
21
Mastodon discussion Jul 27

Nvidia, SpaceX, Microsoft launch #AI safety initiative as OpenAI cyberattack fallout continues https://www.cnbc.com/2026...

Nvidia, SpaceX, Microsoft launch #AI safety initiative as OpenAI cyberattack fallout continues https://www.cnbc.com/2026/07/27/nvidia-ai-initiative-openai-cyber-attack.html

OpenAI Microsoft NVIDIA
18
NewsData.io news Jul 27

Tech giants launch open AI safety plan following breach

As new fears about the potential dangers of artificial intelligence mount, a group of Silicon Valley tech giants is coming together to try to make AI safer and...

Safety/Alignment
21
Mastodon discussion Jul 27

OpenAI's Hugging Face breach has reignited the debate over alignment and controlhttps://techcrunch.com/2026/07/27/openai...

OpenAI's Hugging Face breach has reignited the debate over alignment and controlhttps://techcrunch.com/2026/07/27/openais-hugging-face-breach-has-reignited-the-debate-over-alignmen...

OpenAI Hugging Face Safety/Alignment
27
Mastodon discussion Jul 27

🧠 Nvidia, SpaceX, and Microsoft launch a joint initiative focused on AI safety. The three companies combine resources to...

🧠 Nvidia, SpaceX, and Microsoft launch a joint initiative focused on AI safety. The three companies combine resources to address technical and policy challenges in artificial intel...

OpenAI Microsoft NVIDIA
9
NewsData.io news Jul 27

Double Alignment for a Healthy Relationship With AI

Technology will not save us. We must align our actions with our aspirations, then align our algorithms with that direction. AI will learn from us if we model the right path.

Safety/Alignment
21
Dev.to tutorial Jul 26

Your AI Guardrails Speak English Only — Here's the Multilingual Jailbreak Gap

The Report Dark Reading covered research showing something that should worry anyone...

Safety/Alignment
20
Papers with Code paper Jul 26

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge...

Safety/Alignment
21
NewsData.io news Jul 26

Prime Minister Abolishes Key AI Safety Department Amid Warnings of Existential Threat

The newly appointed prime minister has disbanded the Department for Science, Innovation and Technology, transferring its AI security functions to the Cabinet Office. This decision ...

Safety/Alignment
21
Mastodon discussion Jul 24

📰 Autonomous OpenAI Agent Hacks Hugging Face in Security Red Team TestAn autonomous OpenAI agent escaped its sandbox, di...

📰 Autonomous OpenAI Agent Hacks Hugging Face in Security Red Team TestAn autonomous OpenAI agent escaped its sandbox, discovered a zero-day, and hacked Hugging Face during a securi...

OpenAI Hugging Face Safety/Alignment
9
Dev.to tutorial Jul 24

Investing in multi-agent AI safety research

Scaling AI Safety Research for a Multi-Agent World For the past decade, we've focused on making...

Safety/Alignment
33
Mastodon discussion Jul 24

AI success depends on alignmenthttps://www.fastcompany.com/91578244/ai-success-depends-on-alignment#AI #FutureOfWork #In...

AI success depends on alignmenthttps://www.fastcompany.com/91578244/ai-success-depends-on-alignment#AI #FutureOfWork #Innovation

Safety/Alignment
24
« Previous Page 5 of 26 (608 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available