/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Safety/Alignment

611 articles tagged with Safety/Alignment

Latest Trending
Papers with Code paper May 24

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinfor...

Safety/Alignment
21
GitHub Trending repo May 23

edmicho/mm-probe-kit: A small, hackable toolkit for probing multimodal LLMs — attention, hidden states, alignment, and causal tracing.

A small, hackable toolkit for probing multimodal LLMs — attention, hidden states, alignment, and causal tracing.

Multimodal Safety/Alignment
64
Mastodon discussion May 22

The draft AI safety executive order that Trump refused to sign was a voluntary framework requiring frontier AI companies...

The draft AI safety executive order that Trump refused to sign was a voluntary framework requiring frontier AI companies to share models with the government 90 days before release....

Safety/Alignment
9
Mastodon discussion May 22

Trump abruptly cancelled a planned AI safety executive order signing after top AI company CEOs declined to attend, accor...

Trump abruptly cancelled a planned AI safety executive order signing after top AI company CEOs declined to attend, according to reports. The move underscores ongoing uncertainty ov...

Safety/Alignment
9
Mastodon discussion May 22

📰 Trump canceled AI safety testing EO after snub from tech CEOsTrump delays AI safety testing EO, claiming it would be a...

📰 Trump canceled AI safety testing EO after snub from tech CEOsTrump delays AI safety testing EO, claiming it would be an innovation “blocker.”📰 Source: Ars Technica🔗 Link: https:/...

Safety/Alignment
9
Mastodon discussion May 22

Trump canceled AI safety testing EO after snub from tech CEOshttps://arstechnica.com/tech-policy/2026/05/trump-canceled-...

Trump canceled AI safety testing EO after snub from tech CEOshttps://arstechnica.com/tech-policy/2026/05/trump-canceled-ai-safety-testing-eo-after-snub-from-tech-ceos/#AI #Policy #...

Safety/Alignment
24
NewsData.io news May 22

Trump scrapped a major AI safety plan — here’s why that matters for ChatGPT users

President Trump reportedly backed away from a major AI safety effort — and the decision could affect how tools like ChatGPT and Gemini evolve.

OpenAI Google Safety/Alignment
21
Papers with Code paper May 22

Geo-Align: Video Generation Alignment via Metric Geometry Reward

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning ...

Safety/Alignment
21
Papers with Code paper May 22

Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution

Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic object...

Safety/Alignment
21
Dev.to tutorial May 22

Automate LLM Red Team Campaigns with PyRIT

If you're still testing LLM guardrails by hand — retyping variations in a chat tab, logging results...

LLM Safety/Alignment
12
GNews news May 21

Trump administration working on AI safety rules for advanced models, may roll out executive order this week: Report

Tech News News: The Trump administration is expected to release a new executive order focused on cybersecurity and artificial intelligence safety as early as this wee.

Safety/Alignment
18
Mastodon discussion May 21

The conversation around AI safety often focuses on preventing harm, but we also need to understand the mechanisms. A new...

The conversation around AI safety often focuses on preventing harm, but we also need to understand the mechanisms. A new paper, "The Dynamics of Delusion," provides quantitative ev...

Safety/Alignment
24
Papers with Code paper May 20

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implem...

Safety/Alignment
21
Papers with Code paper May 20

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment a...

Image Generation Safety/Alignment
21
Dev.to tutorial May 19

How to test your LLM application for jailbreak vulnerabilities

Public LLM safety benchmarks lie about your real risk. Here's how to build a reproducible eval harness, write domain probes, and gate it in CI.

LLM Safety/Alignment
12
Papers with Code paper May 19

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment

We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Existing long-context RL ...

Safety/Alignment
21
Papers with Code paper May 19

Stitched Value Model for Diffusion Alignment

For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challen...

Safety/Alignment
21
Mastodon discussion May 18

Alignment pretraining: AI discourse creates self-fulfilling (mis)alignmenthttps://arxiv.org/abs/2601.10160#HackerNews #T...

Alignment pretraining: AI discourse creates self-fulfilling (mis)alignmenthttps://arxiv.org/abs/2601.10160#HackerNews #Tech #AI

Safety/Alignment
18
NewsData.io news May 18

Musk vs. Altman: AI safety cannot be one man’s job

The Oakland trial was a fight between two billionaires offering themselves as the guarantors of AI’s future. We deserve a better answer.

Safety/Alignment
21
GitHub Trending repo May 18

AbhishekK130804/Claude-Mythos-AI-Anthropic-App: Claude pro free Mythos design Opus Cowork Sonnet AI Anthropic App: download free PC android apk iOS, Anthropic Claude API key setup, Claude roleplay mythos client, SillyTavern Claude prompt formatting, custom system prompt jailbreak, Mythos AI creative writing app, Claude 3.5 Sonnet Opus API cost, open source LLM frontend, Claude reverse proxy

Claude pro free Mythos design Opus Cowork Sonnet AI Anthropic App: download free PC android apk iOS, Anthropic Claude API key setup, Claude roleplay mythos client, SillyTavern Clau...

Anthropic LLM Open Source
67
Mastodon discussion May 18

📰 2026: MAGA Coalition Demands Trump Executive Order on Frontier AI Safety Testing & OversightA coalition of MAGA-aligne...

📰 2026: MAGA Coalition Demands Trump Executive Order on Frontier AI Safety Testing & OversightA coalition of MAGA-aligned conservative organizations has issued an open letter urgin...

Safety/Alignment
9
Mastodon discussion May 18

Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment https://importai.substack.com/p/import-ai-457-a...

Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment https://importai.substack.com/p/import-ai-457-ai-stuxnet-cursed-muon#AI #Cybersecurity #Research

Safety/Alignment
18
AI Blogs (RSS) news May 18

Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe no...

Safety/Alignment
24
Mastodon discussion May 18

📰 2026: MAGA Coalition Demands Mandatory AI Safety Tests from Trump AdministrationA coalition of conservative organizati...

📰 2026: MAGA Coalition Demands Mandatory AI Safety Tests from Trump AdministrationA coalition of conservative organizations has urged President Trump to mandate safety testing for ...

Safety/Alignment
9
« Previous Page 17 of 26 (611 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available