Safe-Sora: Safe Text-to-Video Generation
via Graphical Watermarking
Graphical watermarks embedded during generation.
From adaptive matching to watermark recovery.
Coarse to fine
Adaptive patch matchingSpatiotemporal fusion
3D wavelet-enhanced MambaWatermark extraction
Reconstruction under distortionsCreate. Protect. Recover.
Graphical watermarks, embedded during generation.
Watermark reconstructionSafe-Sora is the first framework that integrates graphical watermarks directly into the video generation process. The following results show the original video, the watermarked video, the difference between them (×5), the original watermark, the recovered watermark, and the difference between them (×5).
Watermark protection,
built into generation.
AI-generated videos need reliable copyright protection. Embedding an invisible graphical watermark requires preserving both video quality and the recoverability of the watermark.
Safe-Sora embeds graphical watermarks directly during video generation. Adaptive patch matching aligns watermark content with video regions, while wavelet-enhanced Mamba supports embedding and recovery across frames.
Read the full abstract ↗
From adaptive matching to recovery


Quality, fidelity, and robustness



Cite this work
@article{su2025safe,
title={Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking},
author={Su, Zihan and Qiu, Xuerui and Xu, Hongbin and Jiang, Tangyu and Zhuang, Junhao and Yuan, Chun and Li, Ming and He, Shengfeng and Yu, Fei Richard},
journal={arXiv preprint arXiv:2505.12667},
year={2025}
}