/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
NewsData.io news Jun 13

OpenAI Faces Multi-State Investigation Over AI Safety and User Harm Concerns

A coalition of US state attorneys general is investigating OpenAI amid growing concerns over AI safety, accountability, and user harm.

OpenAI Safety/Alignment
21
NewsData.io news Jun 13

US government issues AI recall, forcing Anthropic to pull Fable 6 and Mythos 5 over jailbreak panic

The fast-moving world of artificial intelligence just hit its first massive regulatory brick wall, and it is unlike anything we have seen before. In an unprecedented move, Anthropi...

Anthropic Safety/Alignment
21
NewsData.io news Jun 12

UCLA Health launches center to evaluate AI safety and implementation in health care

UCLA Health launched the INOVAi Center on June 11, 2026, to test AI safety in clinical care. Trials show AI scribes cut physician note-writing time and reduce exhaustion.

Safety/Alignment
21
Mastodon discussion Jun 12

This also implies the Alignment Problem of keeping #AI to human goals is intractable without constant human input, as mo...

This also implies the Alignment Problem of keeping #AI to human goals is intractable without constant human input, as model collapse will always pull towards the unaligned minima o...

Safety/Alignment
9
GNews news Jun 12

Microsoft Signs MoU to Collaborate With Singapore Regulator on AI Safety, Security

Microsoft has signed a memorandum of understanding to collaborate with Singapore's Infocomm Media Development Authority on artificial intelligence safety and security, the latter s...

Microsoft Safety/Alignment
18
GitHub Trending repo Jun 12

bingook/bingo: Bingo - AI-powered Red Team Terminal (DeepSeek/Claude/GPT/GLM)

Bingo - AI-powered Red Team Terminal (DeepSeek/Claude/GPT/GLM)

Anthropic Safety/Alignment
63
NewsData.io news Jun 12

Former xAI engineer sues company over AI safety concerns, claims he was laid off for speaking up

A former xAI engineer has accused Elon Musk's artificial intelligence company of dismissing him after he repeatedly warned about the risks posed by Grok. The lawsuit paints a pictu...

xAI Safety/Alignment
21
Mastodon discussion Jun 11

đź“° A former xAI engineer has filed a lawsuit against the company and SpaceX, alleging he was fired for raising AI safety ...

đź“° A former xAI engineer has filed a lawsuit against the company and SpaceX, alleging he was fired for raising AI safety concerns about Grok days before SpaceX's historic IPO.đź”— http...

xAI Safety/Alignment
9
Mastodon discussion Jun 11

🤖 Guided Model Alignment Frameworks Gain Traction in AI ResearchResearchers are increasingly focusing on inference time ...

🤖 Guided Model Alignment Frameworks Gain Traction in AI ResearchResearchers are increasingly focusing on inference time alignment methods to improve the performance of large langua...

Safety/Alignment
18
Dev.to tutorial Jun 11

Why Your Next.js SaaS Needs a Production AI Agent Guardrail Architecture, Not Just a Prompt

I've spent the last year building production AI pipelines for SaaS platforms. The prompts were solid....

Agents Safety/Alignment
12
NewsData.io news Jun 11

Former XAI Engineer Sues Elon Musks SpaceX For Firing Over AI Safety Concerns

A former engineer at Elon Musk’s xAI who now heads a think tank focused on AI safety filed a lawsuit claiming he was fired from the SpaceX subsidiary for raising concerns about the...

xAI Safety/Alignment
21
Mastodon discussion Jun 10

🤖 AI red teaming comes of age📝 When Ram Shankar Siva Kumar launched Microsoft’s AI red team in 2019, the discipline bare...

🤖 AI red teaming comes of age📝 When Ram Shankar Siva Kumar launched Microsoft’s AI red team in 2019, the discipline barely existed. “The running jok...https://www.csoonline.com/art...

Microsoft Safety/Alignment
9
Papers with Code paper Jun 10

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile, Grammar-Constrained Decoding...

Safety/Alignment
21
YouTube video Jun 9

AI Safety Expert: These People Will Only Survive Till 2030

Make yourself and your family AI-scam proof, step by step → https://neuralnutshell.com Roman Yampolsky, who coined the term ...

Safety/Alignment
65
NewsData.io news Jun 9

Center for AI Safety Names Former Robinhood and Meta Executive Rochelle Nadhiri as Head of Public Engagement

SAN FRANCISCO--(BUSINESS WIRE)--The Center for AI Safety (CAIS), a nonprofit focused on reducing societal-scale risks from artificial intelligence, today announced the appointment ...

Safety/Alignment
21
Papers with Code paper Jun 9

The Role of Feedback Alignment in Self-Distillation

Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this impr...

Safety/Alignment
21
Papers with Code paper Jun 9

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spac...

Safety/Alignment
21
Mastodon discussion Jun 8

Survey reveals 80% would jailbreak their Kindle before letting Amazon winAndroid Authity readers want control over their...

Survey reveals 80% would jailbreak their Kindle before letting Amazon winAndroid Authity readers want control over their pre-2012 Kindle devices.https://www.androidauthority.com/su...

Google Safety/Alignment
9
Papers with Code paper Jun 8

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon kno...

Safety/Alignment
21
Papers with Code paper Jun 8

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support toward excessive validation or ...

Safety/Alignment
21
Dev.to tutorial Jun 7

AI in SDLC: Why I Stopped Optimizing for Code Generation and Started Optimizing for Alignment

Over the past few months I built an AI-assisted delivery framework — not to write code faster, but to...

Safety/Alignment
12
NewsData.io news Jun 7

MIT symposium examines AI alignment, education and the limits of machine reasoning

MIT researchers gathered April 30 to examine what happens when AI logic doesn't match human reasoning. A central warning: replacing institutions with AI before understanding how th...

Safety/Alignment
21
Dev.to tutorial Jun 7

Evals Are Alignment Enforcement: Why Your Safety Strategy Needs Runtime Checks

The Argument The AI safety conversation is dominated by two camps: the alignment...

Safety/Alignment
20
Papers with Code paper Jun 7

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion featu...

Safety/Alignment
21
« Previous Page 14 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available