Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

Linwei Dong1,2 Ruoyu Guo2† Ge Bai2 Zehuan Yuan2 Yawei Luo1* Changqing Zou1,3
1Zhejiang University 2Bytedance Inc. 3Zhejiang Lab
† Project Leader *Corresponding Author
Teaser image

Samples from 4-NFE student model distilled through our methods.
GDMD showcases outstanding image generation, delivering ultra-realistic visuals and profound concept understanding.

Abstract

Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices quality for sampling speed. While integrating Reinforcement Learning (RL) into distillation offers potential, a naive fusion of these two objectives relies on suboptimal raw sample evaluation. This sample-based scoring creates inherent conflicts with the distillation trajectory and produces unreliable rewards due to the noisy nature of early-stage generation. To overcome these limitations, we propose GDMD, a novel framework that redefines the reward mechanism by prioritizing distillation gradients over raw pixel outputs as the primary signal for optimization. By reinterpreting the DMD gradients as implicit target tensors, our framework enables existing reward models to directly evaluate the quality of distillation updates. This gradient-level guidance functions as an adaptive weighting that synchronizes the RL policy with the distillation objective, effectively neutralizing optimization divergence. Empirical results show that GDMD sets a new SOTA for few-step generation. Specifically, our 4-step models outperform the quality of their multi-step teacher and substantially exceed previous DMDR results in GenEval and human-preference metrics, exhibiting strong scalability potential.

Radar Vision

Performance of GDMD. Left: Head-to-head comparison with DMDR on the GenEval benchmark. Right: By integrating multiple reward models, GDMD significantly boosts performance across all evaluated benchmarks (including other unseen rewards), significantly outperforming existing methods.

Method

We propose a paradigm shift in DMD-RL distillation: leveraging optimization gradients as the primary reward signal, rather than raw sample pixels. Based on this insight, we propose GDMD, a novel framework that treats the gradients derived from the DMD objective as the guiding target for RL-based optimization.

Distribution Matching (a) Illustration of divergence between RL updates and DMD updates. RL optimization diverges from few-step distillation in certain extreme phases. (b) Compared to DMDR without cold start. DMDR is more prone to reward hacking (background artifacts), and unseen metrics (such as Pick Score) do not show significant improvement. (c) Visualization of the initial phase and visualization of reward hacking in DMDR. GDMD has not exhibited any significant hacking.
Pipeline Overview of the DMDR and GDMD pipelines. Top: DMDR employs a sample-based scoring approach, simply combining the losses from DMD and RL. Bottom: GDMD achieves gradient-based RL optimization through its carefully designed components: Distillation-aware Gradient Collection (DaGC), Implicit Gradient Scoring (IGS), and Negative-aware Preference Optimization (NaPO).

Main Visualization

Qualitative comparison against teacher and existing models. Using identical noise inputs, our method outperforms others in both quality and prompt alignment, showing strong performance.

RL Visualization

Visual comparison of different tasks (Counting, Position, Attribute Binding, Colors and Texts) across different methods: Teacher, DMD and DMD + NFT.

BibTeX


        @article{dong2026guiding,
          title={Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning},
          author={Dong, Linwei and Guo, Ruoyu and Bai, Ge and Yuan, Zehuan and Luo, Yawei and Zou, Changqing},
          journal={arXiv preprint arXiv:2604.19009},
          year={2026}
        }