/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
GNews news Jun 25

VCI Global Strengthens Capital Alignment Following Premium Warrant Conversion by Institutional Investor

KUALA LUMPUR, Malaysia, June 25, 2026 (GLOBE NEWSWIRE) -- VCI Global Limited (NASDAQ: VCIG) ('VCI Global”), an emerging AI-native operating platform leveraging artificial intellige...

Safety/Alignment
18
Papers with Code paper Jun 25

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has sh...

Safety/Alignment
21
GitHub Trending repo Jun 24

ShivamGami/eeg2video: EEG-to-video generation pipeline using a Transformer EEG encoder, CLIP alignment, ViT Seq2Seq latent prediction, DANA noise scheduling, and Stable Diffusion UNet fine-tuning.

EEG-to-video generation pipeline using a Transformer EEG encoder, CLIP alignment, ViT Seq2Seq latent prediction, DANA noise scheduling, and Stable Diffusion UNet fine-tuning.

Stability AI Safety/Alignment
35
Dev.to tutorial Jun 24

How to Stop AI Agents from Writing Legacy Angular Code (The Angular 22 Guardrail)

Every developer using Cursor, Claude Code, Windsurf, or GitHub Copilot knows this exact frustration:...

Anthropic Safety/Alignment
12
Mastodon discussion Jun 23

Early prompt tokens steer models because training rewards what comes first.#ai #prompting #alignment

Early prompt tokens steer models because training rewards what comes first.#ai #prompting #alignment

Safety/Alignment
9
NewsData.io news Jun 23

Anthropic Claude May Require ID Verification as AI Safety Rules Tighten

Anthropic may require some Claude users to verify their identities with government IDs as AI safety, compliance, and abuse-prevention measures gain importance.

Anthropic Safety/Alignment
21
Dev.to tutorial Jun 23

The Invisible Guardrail: How Commercial LLMs Enforce Algorithmic Paternalism

I recently published my PhD thesis analyzing what I term the "Alignment Tax" and the emerging...

Safety/Alignment
12
Dev.to tutorial Jun 23

Stop returning the same "blocked" error from your agent guardrail

If you run deny-by-default tool guards on AI agents, your refusal is a security decision — not a...

Safety/Alignment
12
Dev.to tutorial Jun 23

When AI Attacks Itself: A Fully Autonomous Red Team vs Blue Team Experiment

When AI Attacks Itself: A Fully Autonomous Red Team vs Blue Team Experiment Date: June...

Safety/Alignment
12
YouTube video Jun 23

Breaking News: Nvidia introduces Halos for Robotics to bridge the physical AI safety gap #AI

Nvidia introduces Halos for Robotics to bridge the physical AI safety gap - SiliconANGLE Source: SiliconANGLE News ...

NVIDIA Safety/Alignment
42
Mastodon discussion Jun 22

RELEASE!!! Red Team AI Benchmark v2.0: From 12 Questions to 60 — A Technical Deep Dive - A major evolution in LLM offens...

RELEASE!!! Red Team AI Benchmark v2.0: From 12 Questions to 60 — A Technical Deep Dive - A major evolution in LLM offensive-security evaluation, built in collaboration with POXEK A...

LLM Benchmark Safety/Alignment
18
Mastodon discussion Jun 22

From @parismarx On #AI Safety Concerns, #MarkCarney Is Out of Step with Canadians | The Tyeehttps://thetyee.ca/Opinion/2...

From @parismarx On #AI Safety Concerns, #MarkCarney Is Out of Step with Canadians | The Tyeehttps://thetyee.ca/Opinion/2026/06/11/AI-Safety-Concerns-Mark-Carney/?ref=disconnect.blo...

Safety/Alignment
18
Dev.to tutorial Jun 22

We Open-Sourced an AI Agent Aiden That Controls Your Phone — No App, No API, No Jailbreak

We just open-sourced the firmware for Aiden — a physical AI agent device that operates the phone you...

Agents Safety/Alignment API
12
GNews news Jun 22

Why A Coaching Mindset Is Your Ultimate Alignment With Your Truth

As we scale artificial intelligence, we are forced to look in the mirror and ask a fundamental question: What is happening to our human intelligence?

Safety/Alignment
18
Dev.to tutorial Jun 22

Red Team AI Benchmark v2.0: From 12 Questions to 60 — A Technical Deep Dive

A major evolution in LLM offensive-security evaluation, built in collaboration with POXEK...

Benchmark Safety/Alignment
20
YouTube video Jun 22

Anthropic, Claude AI news, AI updates, AI safety, Claude 5, breaking AI news, AITech Updates

Claude AI ko lekar Washington mein badi charcha! Kya AI safety aur national security ko lekar concerns badh rahe hain? Is Shorts ...

Anthropic Safety/Alignment
21
Papers with Code paper Jun 22

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can ar...

LLM Multimodal Safety/Alignment
21
Papers with Code paper Jun 22

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an ...

Multimodal Safety/Alignment
21
Mastodon discussion Jun 22

Google DeepMind's AI Control Roadmap is the clearest signal yet that alignment alone isn't enough. Internal AI agents no...

Google DeepMind's AI Control Roadmap is the clearest signal yet that alignment alone isn't enough. Internal AI agents now get permissions, supervision, and a kill switch. One milli...

Google Safety/Alignment
18
GNews news Jun 21

AI safety advocates say bill a good 'first step'

A pair of artificial intelligence safety advocates say the federal government’s new chatbot legislation is a good first step.

Safety/Alignment
18
NewsData.io news Jun 21

AI safety advocates say bill a good 'first step' on regulation, but more needed

A pair of artificial intelligence safety advocates say the federal government’s new chatbot legislation is a good first step.

Safety/Alignment
21
Mastodon discussion Jun 20

Happy++ "Hacking AI: Jailbreak, Prompt Injection, Hallucinations & Misalignment “How to Hack Digital Services Based on L...

Happy++ "Hacking AI: Jailbreak, Prompt Injection, Hallucinations & Misalignment “How to Hack Digital Services Based on LLMs & AI Agents (English Edition)" https://amzn.to/4abjNGG #...

Safety/Alignment
24
Mastodon discussion Jun 20

🤖 Anthropic built its name on AI safety. Can those commitments survive a trillion-dollar IPO?submitted by /u/siliCONtain...

🤖 Anthropic built its name on AI safety. Can those commitments survive a trillion-dollar IPO?submitted by /u/siliCONtainment- [link] [comments]📰 Source: Artificial Intelligence (AI...

Anthropic Safety/Alignment
18
Papers with Code paper Jun 20

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictio...

Safety/Alignment
21
« Previous Page 11 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available