Gorio Tech Blog search

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models | Summary

|

Contents

This article explains the key points of D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models.

  • 2026-05-06 (arXiv), Tech Report (D-OPSD)
  • Jiang, Dengyang, Jin, Xin, Liu, Dongyang, Wang, Zanyi, Zheng, Mingzhe, Du, Ruoyi, Yang, Xiangpeng, Wu, Qilong, Li, Zhen, Gao, Peng, et al.
  • The Hong Kong University of Science and Technology, Z-Image Team, Alibaba Group, University of California, San Diego, The Chinese University of Hong Kong
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • D-OPSD addresses supervised adaptation of step-distilled image generators while preserving their few-step generation quality. Standard flow-matching fine-tuning trains on noised target images rather than the states visited by the model’s own sampler, creating a training–inference mismatch.
  • D-OPSD generates text-conditioned student trajectories and matches student velocities to an EMA teacher evaluated on those same states. The teacher receives image–text conditioning through a modified multimodal encoder, allowing paired images to supply supervision without replacing the student’s sampled states or requiring an external reward function.
  • Experiments on Z-Image-Turbo 6B and FLUX.2-klein 4B cover small-dataset LoRA customization and full fine-tuning on 25K anime images. D-OPSD improves few-step quality and adaptation relative to the tested baselines, but general-capability benchmarks still decline from the base models, training iterations cost more, and adaptation can fail when the multimodally conditioned teacher does not capture the target concept.

1 Introduction

The central problem is how to introduce target-image information without replacing the states in the student’s few-step rollout. Online RL offers on-policy optimization, but designing a suitable reward is difficult when customization data consist of only a few image–text pairs.

  • Figure 1 shows that Z-Image-Turbo, sampled with 8 steps, can preserve aspects of reference concepts and styles when conditioned on image–text features rather than text-only features.
  • The paper uses this conditioning behavior to transfer on-policy self-distillation from language models to diffusion models.
  • The project URL is https://vvvvvjdy.github.io/d-opsd; the supplied PDF identifies the report as arXiv:2605.05204v3.

The wolf plushie and sunglasses examples show that text-only prompts do not specify the reference object’s particular appearance. Adding reference-image features brings outputs closer to that appearance, while the port and tiger examples show transfer of stylistic characteristics. These selected examples motivate the richer-context teacher.

Text-only and image–text conditioned generations from Z-Image-Turbo using 8 steps, with reference examples for concepts and styles.
Text-only and image–text conditioned generations from Z-Image-Turbo using 8 steps, with reference examples for concepts and styles.

2 Method

The method separates supervision from sampling: target images enrich the teacher’s conditioning, while the student samples through the original text-to-image pathway. Both branches predict velocities at student-generated states, and only the student receives gradients from their alignment loss.

2.1 Background

Vanilla SFT regresses toward a ground-truth flow-matching velocity on states constructed by noising target images. D-OPSD instead defines both the optimization states and the supervisory predictions on the current student’s trajectory, reducing the mismatch with few-step inference.

  • The motivation concerns adaptation of already step-distilled generators, not training diffusion models from scratch.
  • Online RL also uses generated trajectories, but D-OPSD replaces reward-based supervision with multimodally conditioned self-distillation.

2.2 D-OPSD

For each training pair ((x_0,y)), D-OPSD constructs conditions cs = ftext(y) and ct = fmm(y, x0). The student sees only the prompt, while the teacher sees features encoding both the prompt and the target image; Figure 2 summarizes this separation.

  • The target image enters through the conditioning representation, not through a noisy latent state substituted for the student’s sample.
  • The teacher’s diffusion weights originate from the same base model and are maintained through EMA; its multimodal encoder construction is detailed in Appendix D.

The model defines the ODE (\frac{dx_t}{dt}=v_\theta(x_t,t,c)), using the schedule 1 = t_K > t_{K−1} > … > t_0 = 0. Starting from Gaussian noise, the student follows the same few-step solver used at inference, and teacher and student velocities are evaluated at each shared trajectory state.

  • The formulation uses the native sampling budget, such as 4 or 8 steps, rather than an independently chosen training-time denoising process.
  • The teacher does not determine the states visited during the student rollout.

Equation 6 minimizes L_D-OPSD = E_(x0,y)[(1/K) Σ_(k=1)^K ||u_k^s − sg(u_k^t)||_2^2], where sg denotes stop-gradient. This is mean-squared velocity alignment on on-policy states, not a token-level KL divergence or an explicit distribution-matching objective.

  • The paper interprets velocity alignment as pulling the student’s conditional generation dynamics toward those of the teacher under richer context.
  • This motivates the analogy to language-model OPSD, but the report does not establish a formal equivalence between velocity MSE and a divergence between complete image distributions.

Algorithm 1 initializes both diffusion branches from the base model, samples a minibatch and Gaussian starting states, accumulates alignment losses across the few-step schedule, updates the student, and then updates the teacher through EMA. Transitions between sampled states are detached, and teacher predictions receive no gradients.

  • The teacher branch is discarded after training, so deployment retains the original text-only conditioning pipeline and sampling budget.
  • Preservation of few-step generation quality is an empirical result, not a mathematical guarantee that optimization leaves the generation dynamics unchanged.

The diagram separates text-conditioned sampling from image–text teacher conditioning. Both prediction branches consume student-generated states, gradients update the student, and EMA updates the teacher; reference images influence supervision without replacing rollout states.

D-OPSD training architecture with text-conditioned student sampling, multimodal teacher conditioning, velocity alignment, and EMA updates.
D-OPSD training architecture with text-conditioned student sampling, multimodal teacher conditioning, velocity alignment, and EMA updates.

3 Experiment

The experiments test two adaptation regimes: learning individual subjects or styles through LoRA and shifting a fully fine-tuned model toward an anime domain. Evaluation separates target-image similarity, generalization to new prompts, few-step image quality, and retention of general prompt-following capabilities.

3.1 Experimental Setup

The base models are Z-Image-Turbo 6B and FLUX.2-klein 4B, evaluated with their original inference settings across the main comparisons. LoRA baselines include Vanilla SFT, SFT + LoRA on distilled, Dreambooth, and PSO; full-finetuning baselines include Vanilla SFT and PSO.

  • The small-data setting uses DreamBooth data and additional stylized examples; Table 1 averages over 30 concept classes and 10 style classes.
  • Full fine-tuning uses an in-house dataset of 25K high-quality anime images.
  • The original foundation models’ exact distillation recipes are unavailable, so Appendix C evaluates an open-source DMD two-stage alternative rather than reproducing those recipes.

DINO-D and LPIPS-D measure similarity to target images; VLM-J and CLIP-S assess concept/style consistency and prompt alignment on novel prompts. Quality-S and Aesthetic-S come from an internal VLM-based reward model used only for evaluation, while GenEval and DPG assess retained general capabilities.

  • DINO-D uses DINOv3-ViT-S-plus, LPIPS-D uses VGG, VLM-J uses Qwen3-VL-8B-Instruct, and CLIP-S uses DFN-CLIP-H. For CLIP-S, the special DreamBooth token [V] is replaced with the original class name.
  • Full-finetuning FID and related image metrics use 2K randomly sampled training examples and images generated from their prompts; these are not held-out target-domain measurements.
  • GenEval and DPG use original benchmark prompts rather than Prompt-Enhanced variants.

3.2 Main Results

In LoRA customization, D-OPSD achieves the highest CLIP-S, Quality-S, and Aesthetic-S for both base models in Table 1. For Z-Image-Turbo, its CLIP-S is 0.3664 and Quality-S is 3.7965, compared with PSO’s 0.2893 and 3.3422 and Vanilla SFT’s 0.3095 and 2.4236.

  • PSO achieves lower DINO-D and LPIPS-D than D-OPSD on both models, so D-OPSD does not lead every target-similarity metric.
  • For FLUX.2-klein, PSO also has higher VLM-J: 3.3015 versus D-OPSD’s 3.1255. On Z-Image-Turbo, both score 3.3333.
  • Figure 3 illustrates D-OPSD combining learned subjects with requested novel scenes, while the displayed PSO outputs remain closer to training compositions and SFT outputs lose clarity.

D-OPSD leads CLIP-S, Quality-S, and Aesthetic-S for both models, but PSO achieves lower DINO-D and LPIPS-D and higher VLM-J on FLUX.2-klein. The results support a balance of reference learning, new-prompt adherence, and few-step quality rather than superiority on every metric.

LoRA customization comparison across target similarity, concept/style consistency, prompt alignment, quality, and aesthetics.
LoRA customization comparison across target similarity, concept/style consistency, prompt alignment, quality, and aesthetics.

Under full fine-tuning, D-OPSD reduces FID from 48.6858 to 40.4938 for Z-Image-Turbo and from 45.2335 to 38.2932 for FLUX.2-klein. It also improves DINO-D, LPIPS-D, Quality-S, and Aesthetic-S relative to the corresponding base models and tested adaptation baselines.

  • Retention is incomplete: Z-Image-Turbo GenEval changes from 0.7543 to 0.7170 and DPG from 84.7645 to 84.1116.
  • FLUX.2-klein GenEval changes from 0.8155 to 0.7298 and DPG from 85.5624 to 83.8927.
  • Figure 4 shows clearer target-domain and original-domain outputs than SFT or PSO in the selected examples. The authors attribute benchmark declines to a domain-adaptation trade-off, but the experiments do not independently isolate that explanation.

D-OPSD improves target-domain FID and image-quality scores over both base models while keeping GenEval and DPG much closer to base performance than SFT or PSO. Its GenEval and DPG values nevertheless remain below those of the base models, indicating incomplete retention of benchmark capabilities.

Full-finetuning comparison on anime-domain adaptation and general-capability retention for Z-Image-Turbo and FLUX.2-klein.
Full-finetuning comparison on anime-domain adaptation and general-capability retention for Z-Image-Turbo and FLUX.2-klein.

The dog, teapot, rubber duck, and style examples test prompts that differ from the customization images. D-OPSD combines learned appearance with requested new environments in these examples, whereas PSO often retains training-like compositions and SFT produces less distinct outputs. The comparison illustrates why reference similarity alone does not measure instruction-following generalization.

Z-Image-Turbo customization examples comparing training images, the base model, SFT, PSO, and D-OPSD on novel prompts.
Z-Image-Turbo customization examples comparing training images, the base model, SFT, PSO, and D-OPSD on novel prompts.

The left panel illustrates anime-domain adaptation, and the right panel checks original-domain generation. D-OPSD outputs retain clearer structure and detail than the displayed SFT and PSO outputs in both panels, supporting the quantitative quality comparisons without establishing preservation across all prompts.

Target-domain and original-domain generations after full fine-tuning of Z-Image-Turbo.
Target-domain and original-domain generations after full fine-tuning of Z-Image-Turbo.

The user study combines 50 LoRA prompts and 50 full-finetuning prompts, with paired outputs presented in randomized order and generated using identical noise. Figure 5 reports preferences for D-OPSD over PSO, Vanilla SFT, and the base model on quality, aesthetics, and prompt following.

  • Against the base model, D-OPSD receives 62.0% preference for image quality, 60.0% for aesthetics, and 74.0% for prompt following.
  • Reference images or style exemplars are supplied for the customization setting.
  • The report does not specify the number of annotators or provide uncertainty intervals for these preferences.

The preference bars favor D-OPSD across all three evaluated dimensions and all three opponents. Against the base model, the win shares are 62.0% for quality, 60.0% for aesthetics, and 74.0% for prompt following. The study pools customization and full-finetuning prompts rather than reporting the settings separately.

Human pairwise preferences for image quality, aesthetics, and prompt following using outputs generated from identical noise.
Human pairwise preferences for image quality, aesthetics, and prompt following using outputs generated from identical noise.

3.3 Ablation Study

The training-strategy ablation compares SFT on target images, SFT on teacher-generated images, off-policy distillation on a fixed dataset, and D-OPSD on-policy distillation. Figure 6(a), obtained with Z-Image-Turbo LoRA training, shows the on-policy variant increasing DINO similarity faster while maintaining the highest displayed Quality Score.

  • Replacing target images with teacher samples does not reproduce the benefit of distilling on the student’s current trajectory.
  • Off-policy distillation mitigates quality degradation relative to SFT, but converges more slowly and remains below the on-policy variant in the displayed curves.

Teacher construction affects training stability. A frozen base-model teacher works, directly using a student copy collapses, and an EMA teacher with momentum 0.9999 gives the best reported results in Figure 6(b).

  • The figure also includes EMA momentum 0.9, which performs worse than the high-momentum teacher.
  • The authors suggest that slow EMA updates smooth a variable alignment target while tracking student progress; this is an interpretation of the observed curves.

The strategy curves show on-policy distillation increasing DINO similarity faster while maintaining higher image quality than the displayed alternatives. The teacher curves show collapse with a direct student copy, weaker performance with EMA momentum 0.9, and stable improvement with a frozen teacher or EMA momentum 0.9999. Teacher update dynamics therefore matter in this LoRA experiment.

Z-Image-Turbo LoRA ablations of training strategy and teacher construction, measured through DINO similarity and Quality Score over training.
Z-Image-Turbo LoRA ablations of training strategy and teacher construction, measured through DINO similarity and Quality Score over training.

4 Discussion on Limitations and Future Works

D-OPSD requires student rollouts and teacher inference, resulting in roughly 4× FLOPs and 2× training time per iteration relative to Vanilla SFT. Its success also depends on the multimodally conditioned teacher producing useful target-specific supervision.

  • Figure 7 shows a reference toy whose identity is not captured accurately by the teacher; after training, the student follows the teacher’s incorrect appearance.
  • The reported resource advantage concerns avoiding a separate re-distillation stage, not making individual iterations cheaper than SFT.

The multimodal teacher changes the reference toy’s appearance, and the adapted student inherits that mismatch. This example demonstrates that alignment on student-generated states cannot recover reference identity when the teacher supplies inaccurate concept supervision.

A concept-learning failure comparing the target image, teacher sample, and student outputs before and after training.
A concept-learning failure comparing the target image, teacher sample, and student outputs before and after training.

Proposed future directions include richer teacher conditioning from image-editing or video-generation models, additional training constraints, and multi-expert on-policy distillation. These are research proposals, not capabilities demonstrated by the reported experiments.

The paper connects step distillation with on-policy self-distillation. Prior step-distillation methods compress trajectories or match distributions to accelerate inference; language-model OPSD uses richer context to obtain self-generated supervision on current-policy samples.

  • D-OPSD adapts an already distilled image generator rather than introducing a procedure for initially compressing a multi-step model.
  • Its diffusion-specific construction replaces discrete token-distribution alignment with conditional velocity alignment at student-generated states.

6 Conclusion

The reported results support one-stage supervised adaptation of the tested few-step diffusion models through multimodal-context self-distillation, without external reward design or a separate re-distillation stage. D-OPSD preserves few-step image quality and general benchmark performance better than the tested SFT and PSO baselines, while still exhibiting retention losses and teacher-dependent failures.

FLUX.2-klein exhibits the qualitative conditioning behavior used to motivate D-OPSD for Z-Image-Turbo: image–text features better preserve particular object appearances and reference styles than text-only features in the selected examples. This extends the observation to a second tested model, not all LLM/VLM-conditioned generators.

Reference, text-only, and image–text conditioned examples from FLUX.2-klein-4B using 4 steps.
Reference, text-only, and image–text conditioned examples from FLUX.2-klein-4B using 4 steps.

D-OPSD and the listed online RL formulations are marked as on-policy and consistent between training and inference, but only D-OPSD avoids reward requirements among those options. SFT and the listed offline preference methods use ground-truth velocity supervision on target-image-related states. These entries describe formulations rather than quantitative performance.

Training-paradigm comparison by supervision source, on-policy status, reward requirements, and training–inference consistency.
Training-paradigm comparison by supervision source, on-policy status, reward requirements, and training–inference consistency.

Multi-step SFT achieves the lowest DINO-D but requires 100 inference NFEs. DMD restores an 8-NFE budget with substantial additional GPU-hour cost, whereas D-OPSD retains 8 NFEs with higher Quality-S and, in full fine-tuning, higher GenEval than the tested two-stage pipeline. The result is specific to the open-source DMD implementation; some full-finetuning SFT entries also differ from the Z-Image-Turbo entries in Table 2.

One-stage adaptation and two-stage SFT plus DMD compared by similarity, quality, retained capabilities, inference NFEs, and H800 GPU hours.
One-stage adaptation and two-stage SFT plus DMD compared by similarity, quality, retained capabilities, inference NFEs, and H800 GPU hours.

Direct Qwen3-VL conditioning produces excessive detail and visible artifacts in the dog and painting examples. Replacing its language-model weights with Qwen3-4B weights reduces those effects while retaining image-conditioned generation, illustrating the feature-compatibility adjustment used for the teacher encoder.

Z-Image-Turbo multimodal generations using Qwen3-VL-4B directly and with its LLM component replaced by Qwen3-4B weights.
Z-Image-Turbo multimodal generations using Qwen3-VL-4B directly and with its LLM component replaced by Qwen3-4B weights.

Appendix

  • A Investigation of FLUX.2-klein: Figure 8 repeats the text-only versus image–text conditioning comparison with FLUX.2-klein-4B using 4 steps. Reference-conditioned generations preserve aspects of subject appearance or style, extending the qualitative observation beyond Z-Image-Turbo without establishing universality across encoder architectures.
  • B Discussion and Comparison of Different Training Paradigms: Table 3 distinguishes supervision source, on-policy training, reward requirements, and training–inference consistency. D-OPSD uses self-distilled velocities on current student states; the compared online RL methods also use current-policy samples but require rewards, whereas SFT and the discussed offline preference methods use target-image-induced states and ground-truth velocity supervision.
  • C Comparison with Two-stage Training Pipeline: Table 4 compares direct adaptation with Z-Image SFT followed by open-source DMD distillation back to 8 NFEs. In LoRA customization, D-OPSD uses 9 H800 GPU hours versus 5 + 882 for SFT + DMD and achieves Quality-S 3.7965 versus 3.3372; in full fine-tuning, it uses 1918 versus 1066 + 2054 GPU hours and achieves Quality-S 3.8438 versus 3.3652. The comparison does not reproduce the unavailable original foundation-model distillation recipes.
  • D Implementation Details.: Both students retain Qwen3-4B text encoding. For multimodal teacher encoding, directly substituting Qwen3-VL-4B produces high-frequency artifacts and excessive sharpening; the implementation instead replaces its LLM component with Qwen3-4B weights while retaining the ViT and Connector weights, as illustrated in Figure 9.
  • LoRA training on small customized dataset: Both models use rank 64, alpha 128, total batch size 4, EMA momentum 0.9999, and 1K iterations on a single H800 GPU. Learning rates are 4e-5 for Z-Image-Turbo and 1e-5 for FLUX.2-klein; EMA is applied to LoRA weights while sharing the backbone weights.
  • Full finetuning on larger scale dataset: All diffusion-transformer backbone parameters are updated with total batch size 256 and EMA momentum 0.9999. Both models train for 10K iterations on 32 H800 GPUs, with learning rates 3e-5 for Z-Image-Turbo and 8e-6 for FLUX.2-klein.
  • E More Details of Evaluation: LoRA target-similarity metrics use caption paraphrases, while concept/style generalization uses four groups of new prompts and a VLM-J scale from 1 to 4. Full-finetuning image metrics use 2K examples sampled from the training set; the quality and aesthetic evaluator is an internal VLM-based reward model not optimized during training. The human study aggregates preferences across 100 prompts from the two adaptation settings.
  • F More Related Works: The appendix reviews text-to-image diffusion, knowledge and self-distillation, and adaptation techniques including DreamBooth, textual inversion, and LoRA. It distinguishes D-OPSD from standard adaptation methods developed for multi-step models, which do not explicitly target preservation of an already distilled generator’s native few-step behavior.

Brief Thoughts

The fixed-inference-budget comparisons and training-strategy ablation provide the clearest support for the method: improved quality is associated with distillation on student-generated states, not merely with generating replacement training targets. The LoRA results also distinguish reference matching from new-scene instruction following, since PSO leads some similarity metrics while scoring lower on CLIP-S.

The evidence is narrower than a general solution to continual learning: the experiments cover individual customization and anime-domain adaptation, not a sequence of successive tasks. Private training data and quality evaluators, training-set-based domain metrics, missing uncertainty estimates, and the teacher failure case limit reproducibility and broad claims about knowledge retention.