Mastodon discussion Apr 15

Anthropic (@AnthropicAI)Anthropic Fellows의 새로운 연구로, 약한 AI 모델이 강한 모델의 학습을 감독하는 ‘Automated Alignment Researcher’ 실험이 소개됐다....

Anthropic (@AnthropicAI)Anthropic Fellows의 새로운 연구로, 약한 AI 모델이 강한 모델의 학습을 감독하는 ‘Automated Alignment Researcher’ 실험이 소개됐다. Claude Opus 4.6이 정렬 연구의 실험 속도와 탐색 범위를 높일 수 있음을 보여주는 의미 있는 연...

GitHub Trending repo Apr 14

baojudezeze/Qwen-dpo: Training code for Diffusion-DPO applied to the Qwen Image-2512 model. This implementation builds on the training framework provided by zk1009 and follows the methodology described in the paper “Diffusion Model Alignment Using Direct Preference Optimization”.

Training code for Diffusion-DPO applied to the Qwen Image-2512 model. This implementation builds on the training framework provided by zk1009 and follows the methodology described ...