SOURCE-LINKED INTELLIGENCE
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO…
Read original source ↗ Open in workspace
- recordType
- article
- region
- Global
Evidence & attribution
- Apple Machine Learning Research · 2026-09-16T00:00:00.000Z
First collected: 2026-09-19T20:26:46.936Z. This is not the publication date.