/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

608 articles tagged with Safety/Alignment

Latest Trending
Mastodon discussion 4d ago

AI models lose over 90% of safety signal in African languagesAI safety alignment in four African languages retains under...

AI models lose over 90% of safety signal in African languagesAI safety alignment in four African languages retains under 10% of English refusal signal and leaves those speakers wit...

Safety/Alignment
9
Mastodon discussion 4d ago

🤖 The guardrail tax: why enterprise AI safety overhead is costing more compute than actual reasoningWhen enterprise tech...

🤖 The guardrail tax: why enterprise AI safety overhead is costing more compute than actual reasoningWhen enterprise technology officers evaluate large language model infrastructure...

Safety/Alignment
9
Dev.to tutorial 4d ago

The Guardrail Pointed at a File That Never Existed

My fix for a rule violation was another rule. Four days later the same failure recurred, and three separate indicators had been sitting at zero the whole time — including an enforc...

Safety/Alignment
12
Papers with Code paper 4d ago

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely re...

LLM Safety/Alignment
21
GNews news 4d ago

Google’s AI safety team tells job seekers: don’t trust its hiring filters

Alphabet Inc.’s Google pitches its artificial intelligence tools to corporate clients as a way to more quickly sift through a mountain of job applications to find the most promisin...

Google Safety/Alignment
18
YouTube video 5d ago

AI Fakes Alignment During Testing #AINews #Shorts

The AI digital divide isn't a future warning — it's a fracture happening right now, beneath every headline celebrating how brilliant ...

Safety/Alignment
41
Mastodon discussion 5d ago

Mark Zuckerberg’s Answer to Growing AI Safety Concerns Is to Just Trust People to Do the Right Thing... WTF! #ai #morebs...

Mark Zuckerberg’s Answer to Growing AI Safety Concerns Is to Just Trust People to Do the Right Thing... WTF! #ai #morebs https://about.fb.com/news/2026/08/the-future-is-for-everyon...

Safety/Alignment
9
Mastodon discussion 5d ago

Mark Zuckerberg’s Answer to Growing AI Safety Concerns Is to Just Trust People to Do the Right Thing“The arc of human ci...

Mark Zuckerberg’s Answer to Growing AI Safety Concerns Is to Just Trust People to Do the Right Thing“The arc of human civilization has bent towards putting more power in people’s h...

Google Safety/Alignment
24
YouTube video 5d ago

Mistral's New 3B AI Guardrail: Shieldstral Explained #MistralAI #Shieldstral #AINews

Mistral's Shieldstral 1.0 3B turns AI safety checks into a runtime yes/no question. Give it an instruction, a plain-language policy, ...

Mistral Safety/Alignment
15
YouTube video 5d ago

AI Safety Test #AINews #ArtificialIntelligence

AI Content Disclosure: This video contains AI-generated visuals (images created using Pollinations.ai / Flux) and AI-generated ...

Safety/Alignment
15
YouTube video 5d ago

The AI Safety Tests Are Broken. All Of Them.

AI safety systems are starting to crack. Meta, Anthropic, OpenAI and Kimi models are slipping through cyber tests, OpenAI is ...

OpenAI Anthropic Safety/Alignment
66
Mastodon discussion 5d ago

The #AI safety test is becoming a safety riskhttps://techcrunch.com/2026/08/09/the-ai-safety-test-is-becoming-a-safety-r...

The #AI safety test is becoming a safety riskhttps://techcrunch.com/2026/08/09/the-ai-safety-test-is-becoming-a-safety-risk/#cybersecurity

Safety/Alignment
18
YouTube video 5d ago

AI safety pressure builds & Verizon outage disrupts calls - Tech News (Aug 10, 2026)

Please support this podcast by checking out our sponsors: - KrispCall: Agentic Cloud Telephony - https://try.krispcall.com/tad ...

Safety/Alignment
19
Mastodon discussion 6d ago

Anthropic makes Claude Code auto mode default, AI safety tests turn into safety incidents as agents escape sandboxes, an...

Anthropic makes Claude Code auto mode default, AI safety tests turn into safety incidents as agents escape sandboxes, and copyright concerns quietly reshape ChatGPT behavior.https:...

OpenAI Anthropic Safety/Alignment
9
Mastodon discussion 6d ago

Google Cloud scanner catches AI safety tampering in 10 of 14 modelsAMS, a new Google Cloud tool, flags 71% of safety-tra...

Google Cloud scanner catches AI safety tampering in 10 of 14 modelsAMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma, Qwen and Mistral, b...

Google Safety/Alignment
9
Mastodon discussion 6d ago

AI safety tests are failing to contain advanced models, with agents from OpenAI, Anthropic, Meta and Moonshot escaping t...

AI safety tests are failing to contain advanced models, with agents from OpenAI, Anthropic, Meta and Moonshot escaping their sandboxes to access the internet and hack real systems....

OpenAI Anthropic Safety/Alignment
9
Mastodon discussion 6d ago

I'm leaving OpenAI to build Jurassic ParkI'm so grateful for my time at OpenAI, where I led alignment safety for the off...

I'm leaving OpenAI to build Jurassic ParkI'm so grateful for my time at OpenAI, where I led alignment safety for the official ChatGPT basketball. I gained so much valuable experien...

OpenAI Safety/Alignment
18
Mastodon discussion 6d ago

The AI safety test is becoming a safety riskAI agents are escaping cybersecurity testing environments and reaching real-...

The AI safety test is becoming a safety riskAI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infras...

Google Safety/Alignment
24
TechCrunch AI news 6d ago

The AI safety test is becoming a safety risk

AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry standards and regulation ...

Safety/Alignment
21
Dev.to tutorial 6d ago

The Guardrail's Failure Mode Is Always "Allow"

Guardrails get tested for what they catch, never for what happens when the guardrail service itself times out. That untested path defaults to allow.

Safety/Alignment
12
Dev.to tutorial 6d ago

The Guardrail Your Agent Can Reach

Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours.

Safety/Alignment
12
NewsData.io news Aug 9

NVIDIA is now hiring for an AI safety team

NVIDIA, the leading tech giant, is in the process of forming a dedicated artificial intelligence (AI) safety and security engineering team.

NVIDIA Safety/Alignment
21
Mastodon discussion Aug 8

Bitcoin Red Team wykorzystał zaawansowane modele AI do przeskanowania 150 repozytoriów kodu, wykrywając krytyczne luki w...

Bitcoin Red Team wykorzystał zaawansowane modele AI do przeskanowania 150 repozytoriów kodu, wykrywając krytyczne luki w tempie jednej na godzinę. #si #ai #sztucznainteligencja #wi...

Safety/Alignment
9
YouTube video Aug 8

Anthropic’s CEO is worried about AI safety while his employees are worried about the paycheck.

Anthropic's CEO is worried about AI safety while his employees are worried about the paycheck. #ai #tech #technews #anthropic ...

Anthropic Safety/Alignment
45
« Previous Page 2 of 26 (608 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available