DSAQuant Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

Preserve the structure. Recover the detail. Train quantized video diffusion models in alignment with the distinct role of each denoising stage.

Simplicity

Minimal changes. Integrates quickly with existing training and inference pipelines.

Generality

Across models and scales. Works across Wan and CogVideoX at multiple sizes and resolutions.

Effectiveness

Stable gains. Averaging +2.22 VBench at INT4 and +3.34 at INT3.

02 / Motivation

Quantization errors are stage-dependent

Conventional QAT does not fail uniformly: early-stage perturbations are relatively mild, while late errors destroy detail and texture.

Observation 01 · Conventional QAT

Challenges of conventional QAT

Full-precision video frames used as the visual reference
Full-precision reference.

Uniform QAT retains the scene composition and temporal motion, but fails to recover high-frequency appearance and texture.

PreservedStructure & motion DegradedAppearance & texture
Conventional QAT results with preserved structure but degraded appearance and texture
Failure mode of conventional quantization-aware training.
Observation 02 · Stage-wise analysis

Where quantization happens matters.

The same perturbation has sharply different effects along the denoising trajectory.

Early steps Limited impact

Slight layout changes, while overall visual quality remains high.

Late steps Severe degradation

Visible artifacts emerge and fine detail breaks down.

Stage-wise comparison showing mild impact from early-step quantization and severe degradation from late-step quantization
Quantization errors affect different denoising stages differently.
03 / Method

One principle, applied twice

Quantized video diffusion models should be trained according to the role of each denoising stage.

Training

Denoising-Stage Oriented Supervision

Teacher anchoring stabilizes early structure, then smoothly yields to the native target loss so late stages can reconstruct detail.

Early stepsDenoisingLate steps
Teacher anchoring
Target learning
Inference

Denoising-Stage Gated Guidance

Classifier-free guidance is gated near the end of denoising, preventing conditional mismatch from amplifying quantization noise.

Early stepsDenoisingLate steps
CFG Active CFG Drop
DSAQuant framework Figure 02
Overview of the DSAQuant training and inference framework
Stage-oriented supervision during training and gated guidance during inference.
01 / Results

Low precision, high quality.

Explore eighteen INT4 DSAQuant outputs across three Wan model scales. Each collection cycles automatically between two groups of three samples.

04 / Comparisons

Inspect every comparison

Switch between system-level comparisons and ablations, then choose a model family.

All videos are muted and looped
Acknowledgements

We would like to thank Yichong Lu, Jiahao Wang, Yufeng Yuan, Nan Zhou, Jiahao Shao, and Yudong Jin for their insightful discussions.