iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that only-latter timestep updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Community
RL fine-tuning of diffusion models (DDPO-style) usually buys reward at the cost of diversity and mode collapse. iADD keeps both: sparse, incrementally chosen timestep updates plus training-time Feynman-Kac branch/resample. Rarity on hard prompts 66.85% vs 24.81% for DDPO; combined with GRPO it lifts DanceGRPO's CLIPScore 0.3886 -> 0.4126. Works on Stable Diffusion, vanishing-point correction and 3D scene synthesis.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration (2026)
- DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance (2026)
- CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models (2026)
- Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View (2026)
- Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization (2026)
- ExploreNet: Learning Where to Explore in Diffusion GRPO (2026)
- CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.01789 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper