/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

609 articles tagged with Safety/Alignment

Latest Trending
NewsData.io news May 28

Xage Security Unlocks Jailbreak-proof AI Agent Autonomy with End-to-End Visibility and Control

New Zero Trust capabilities provide deterministic visibility and control over AI agents, enabling secure production deployments across SaaS, cloud, in-house data center and edge As...

Agents Safety/Alignment
21
GitHub Trending repo May 28

shenwenhao01/GPA-HMR: [CVPR 2026] Official Code for CVPR 2026 paper "VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery"

[CVPR 2026] Official Code for CVPR 2026 paper "VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery"

Safety/Alignment
35
Mastodon discussion May 28

Illinois Lawmakers Just Passed America's Strongest AI Safety Billhttps://www.wired.com/story/illinois-pass-major-ai-safe...

Illinois Lawmakers Just Passed America's Strongest AI Safety Billhttps://www.wired.com/story/illinois-pass-major-ai-safety-law-pritzker/#AI #Regulation #TechPolicy

Safety/Alignment
18
Papers with Code paper May 28

Native Audio-Visual Alignment for Generation

Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on eith...

Safety/Alignment
21
Papers with Code paper May 28

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI mod...

OpenAI Agents Safety/Alignment
21
GitHub Trending repo May 27

beykantemel0702azfy8144/WorpGPT-Latest-2026-AllPrompts: A comprehensive Red Teaming framework for testing Large Language Model (LLM) robustness against adversarial prompt engineering and jailbreak vectors.

A comprehensive Red Teaming framework for testing Large Language Model (LLM) robustness against adversarial prompt engineering and jailbreak vectors.

LLM Safety/Alignment
67
Mastodon discussion May 27

🤖 Built a live red team environment for AI agent security — try to get a prompt injection throughAI agents that can use ...

🤖 Built a live red team environment for AI agent security — try to get a prompt injection throughAI agents that can use tools have a serious problem: any content they read can cont...

Agents Safety/Alignment
18
Mastodon discussion May 27

Alignment Is Not Meetings. It Is Shared Accountability. Alignment is not communication volume. It is collective ownershi...

Alignment Is Not Meetings. It Is Shared Accountability. Alignment is not communication volume. It is collective ownership of outcomes. #Leadership #CIO #CEO #BusinessTransformation...

Safety/Alignment
35
Papers with Code paper May 27

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers ...

LLM Safety/Alignment
21
NewsData.io news May 26

AI safety regulations advance in Springfield, despite industry concern

(The Center Square) – A push to regulate artificial intelligence products in Illinois has taken a major step toward becoming law. The plan, which has broad support from industry le...

Safety/Alignment
21
NewsData.io news May 26

Guardrail Dynamics and the Thrill of Plinko

Guardrail Dynamics and the Thrill of Plinko Understanding the Mechanics of Plinko The Role of the Random Number Generator (RNG) Developing a Plinko Strategy Plinko Variations and M...

Safety/Alignment
21
GitHub Trending repo May 26

zhristophe/Claude-Mythos-AI-Anthropic-App: claude mythos ai anthropic app github open source desktop mobile client abhishekk130804 anthropic api key creative writing roleplay character cards deep reasoning chain of thought system prompts jailbreak windows macos linux android apk next js electron tauri python rust download setup guide tutorial latest version 2026 free update

claude mythos ai anthropic app github open source desktop mobile client abhishekk130804 anthropic api key creative writing roleplay character cards deep reasoning chain of thought ...

Anthropic Open Source Safety/Alignment
64
Mastodon discussion May 26

🔥 TRENDING📢 AI Safety Report 2026: Bestehende KI-Sicherheitspraktiken reichen nicht aus - heise online🔗 https://news.goo...

🔥 TRENDING📢 AI Safety Report 2026: Bestehende KI-Sicherheitspraktiken reichen nicht aus - heise online🔗 https://news.google.com/rss/articles/CBMitAFBVV95cUxPS3h3aDF6cks5aWR6TGxvRjQ...

Google Safety/Alignment
18
Papers with Code paper May 26

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tamperin...

Safety/Alignment
21
Mastodon discussion May 25

Constructive alignment and AI-integration into teaching and learning In Rethinking  the  Integration  of  AI  in  Higher...

Constructive alignment and AI-integration into teaching and learning In Rethinking  the  Integration  of  AI  in  Higher  Education  Teaching  and  Learning, Lilian Schofielda and ...

Safety/Alignment API
9
Mastodon discussion May 25

Whether an #AI system is safe or not is a question of perspective...as the comic nicely illustrates, #alignment is a cha...

Whether an #AI system is safe or not is a question of perspective...as the comic nicely illustrates, #alignment is a challenge that goes way beyond #technologyhttps://www.smbc-comi...

Safety/Alignment
24
Mastodon discussion May 25

@juergen_hubert The depiction of the monster is poignant because:1) The alignment problem of large language models has b...

@juergen_hubert The depiction of the monster is poignant because:1) The alignment problem of large language models has been depicted as a Lovecraftian monsters (Cthulhu) with a smi...

OpenAI Safety/Alignment
9
NewsData.io news May 24

Trump drops AI safety review after tech billionaires lobby against it

Trump dropped a planned AI safety review hours before signing it Thursday, after Elon Musk and Mark Zuckerberg lobbied against it. The order would have been voluntary with no legal...

Safety/Alignment
21
YouTube video May 24

Trump Scraps Executive Order Demanding AI Safety Reviews | AI News Roundup

Trump Scraps Executive Order Demanding AI Safety Reviews — AI News Roundup, 2026-05-24 • Trump Scraps Executive Order ...

Safety/Alignment
19
GitHub Trending repo May 24

pardcomper/mllm-jailbreak-bench: Reproducible benchmark for adversarial attacks on multimodal large language models

Reproducible benchmark for adversarial attacks on multimodal large language models

Multimodal Benchmark Safety/Alignment
64
Dev.to tutorial May 24

AI Safety is a Systems Problem: Building a 4-Layer Runtime Defense

When we talk about LLM security, the conversation usually flattens into semantic prompt analysis or...

Safety/Alignment
12
YouTube video May 24

White House Scraps the AI Safety Executive Order #ainews #aivideo #ai #aifinance

Safety/Alignment
37
Papers with Code paper May 24

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinfor...

Safety/Alignment
21
GitHub Trending repo May 23

edmicho/mm-probe-kit: A small, hackable toolkit for probing multimodal LLMs — attention, hidden states, alignment, and causal tracing.

A small, hackable toolkit for probing multimodal LLMs — attention, hidden states, alignment, and causal tracing.

Multimodal Safety/Alignment
64
« Previous Page 16 of 26 (609 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available