Self Gradient Forcing PlusDecoupling Gradient Flows for Autoregressive Video Generation

Zihan Su1,2* Junhao Zhuang2*† Yaowei Li2 Siwen Lu1 Haoran Li2 Lingen Li3 Haoyu Wu2
Weiyang Jin2 Songchun Zhang2 Haoyang Huang2 Chun Yuan1† Zeyue Xue2 Nan Duan2

1Tsinghua University2Joy Future Academy, JD3The Chinese University of Hong Kong

* Equal contribution   † Corresponding authors

SGF+ separates context-writing and denoising parameters to resolve conflicting gradient updates, improving autoregressive video generation and enabling rollouts of up to 24 hours from only 5s training windows.

Abstract

SGF+ paper teaser: long-horizon comparisons of SF, SGF and SGF+, and sampled frames from continuous 24-hour SGF+ generation

Autoregressive video models must both denoise current frames and write context for future predictions. Self Gradient Forcing (SGF) restores context-writing gradients, but these gradients conflict with denoising updates on shared parameters.

Self Gradient Forcing Plus (SGF+) assigns independent parameters to the two roles and trains them jointly with the original objective. It improves visual quality and long-horizon consistency in both framewise and chunkwise generation, supporting continuous generation for up to 24 hours from only 5s training rollouts.

Method

Autoregressive video generation has two roles: writing context for future predictions and denoising the current frames. SF blocks context-writing gradients; SGF restores them on shared parameters. SGF+ assigns independent parameter sets to the two roles, separating their gradient updates while training both jointly with the original objective.

Method lineage: SF blocks context gradients, SGF restores context gradients, and SGF+ separates context-writing and denoising parameters
Method lineage for autoregressive video generation. Restoring context gradients reveals competing optimization demands. SGF+ separates the parameters responsible for context writing and denoising.

Context writing and denoising exhibit distinct gradient distributions in both Attention and FFN. Across 128 prompts and four denoising timesteps, all 512 paired cosine similarities are negative in each module family. These conflicting updates partially cancel on shared weights, motivating SGF+’s role-specific parameters.

Gradient conflict in SGF: distributions, mean gradient directions, and paired cosine similarities across timesteps for Attention and FFN
Conflicting role gradients. Left: angular-distance t-SNE distributions. Middle: mean gradient directions. Right: paired cosine similarities across denoising timesteps.
View layer-wise gradient distributionsHide layer-wise gradient distributions
Layer-wise gradient distributions across Attention blocks, from the paper
Attention. Context-writing and denoising gradient distributions across layers. Click the figure to open the vector PDF.
Layer-wise gradient distributions across FFN blocks, from the paper
FFN. Block 29 has zero context-writing gradients, so its angular distances are undefined and no t-SNE is shown. Click the figure to open the vector PDF.

The two-pass training procedure first records a no-gradient autoregressive rollout. A second pass reconstructs the forward computation from detached context and noisy target latents, allowing future-generation losses to supervise context writing through differentiable KV states. SGF+ routes context-writing and denoising gradients to their respective parameter sets.

Two-pass SGF and SGF+ training pipeline with independent context-writing and denoising parameters
Training pipeline. SGF accumulates both gradient paths on shared parameters. SGF+ uses separate context-writing and denoising parameters, with the same training objective and short rollout horizon.

SEE THE DIFFERENCE

Same prompt. Same timeline.

CHUNKWISE · 240 SECONDS

A train journey

Self Forcing
Self Gradient Forcing
Self Gradient Forcing Plus
0:00 / 4:00
Check long-horizon consistency

A woman in a kimono, seated inside a moving train.

Full generation prompt

QUANTITATIVE RESULTS

Long-horizon generation.

Long-horizon metrics reported in the paper
MethodSubjectBackgroundFlickeringMotionDynamicsAestheticsImaging

All scores are multiplied by 100, higher is better. Bold indicates the highest score in each column. 60s: VBench-Long prompts. 240s: MovieGen-128 prompts.

Scene jumps and deformation can inflate Dynamic Degree, while SGF+ exhibits less error accumulation and greater scene stability.

NATIVE TEMPORAL EXTRAPOLATION

From 5-second training
to 24-hour generation.

Continuous SGF+ rollouts
without long-video fine-tuning.

24 HOURS
Explore the full 24 hours

Citation

Download .bib
@misc{su2026sgfdecouplinggradientflows,
      title={SGF+: Decoupling Gradient Flows for Autoregressive Video Generation},
      author={Zihan Su and Junhao Zhuang and Yaowei Li and Siwen Lu and Haoran Li and Lingen Li and Haoyu Wu and Weiyang Jin and Songchun Zhang and Haoyang Huang and Chun Yuan and Zeyue Xue and Nan Duan},
      year={2026},
      eprint={2610.10429},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.10429},
}