For the best experience, it's better to use a desktop computer to view this website.
Wan logo Wan Team

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu1,2*, Yawen Shao2,3*, Ziyun Dai2,4*, Zixun Fang2,3*, Kai Zhu2†‡, Siyang Sun2, Haolan Xue2, Chuxin Wang2, Tingyu Weng2, Jingming Luo2, Chen Shi2, Lianghua Huang2, Yufeng Ai2, Yuzheng Wang2, Wenyuan Zhang2, Yu Shang2, Yuxiang Bao2, Zoubin Bi2, Jie Xiao2, Jinbo Xing2, Jiaxing Zhao2, Chongyang Zhong2, Hengjian Chen2, Chenwei Xie2, Akide Liu2, Zhehan Kan5, Yu Liu2, Wei Zhai3‡, Sheng Zhong1, Wei Tong1‡

1Nanjing University 2Wan Team, Alibaba Group 3University of Science and Technology of China 4Fudan University 5Tsinghua University

*Equal contribution †Project leader ‡Corresponding author

Abstract

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction, and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by ~11K blind pairwise assessments. When powering Wan3.0’s video generator, WanPE-397B boosts human preference over raw user prompts by 10.7–18.8 points at 5–15 seconds, and by a dramatic 50.9 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5–15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

Motivation

WanPE motivation

Method

WanPE training pipeline

Results

WanPE results

Preference Scores on WanPEval

(i) 30-second subset against Seedance 2.5, under Wan3.0. (ii) Reverse-constructed vs. forward-based enhancement, under Wan3.0. (iii) Format-adapted WanPE vs. native enhancers on LTX-2.5 and MiniMax-H3 in two separate battles.

Method Action Anim. Speech Ad. Sing & Dance Drama Know. Overall
Exp1: 30-second subset
Seedance 2.5 68.18 46.43 55.00 81.25 60.00 56.00 60.71 59.76
downstream generator: Wan3.0's video generator
Original request 4.179.3815.795.00 22.504.003.579.38
+ WanPE-397B 50.00 81.25 73.68 75.00 47.50 53.85 57.14 60.24
Exp2: Reverse vs. forward enhancement
downstream generator: Wan3.0's video generator
Original Request 23.9721.0129.8519.32 35.0022.606.3823.05
Forward Rewriting 39.5738.4342.1343.57 38.6134.6241.8939.49
Forward-target SFT 30.3644.2332.2837.00 35.3433.7234.2635.17
WanPE-397B-SFT 46.4855.9848.3149.32 45.5151.5451.2349.86
Exp3: Cross-generator transfer (5–15s)
downstream generator: LTX-2.5-Base
LTX-2.5-PE 8.3318.3330.0032.50 7.5026.6725.0021.11
WanPE-397B 25.0038.3353.3327.50 27.5036.6735.0035.56
downstream generator: MiniMax-H3-Base
H3-Context-IR 44.4431.0333.3339.47 37.5029.3140.0035.92
WanPE-397B 29.6337.9340.0044.74 47.5050.0040.0041.09

Video Demos

Play speed

Acknowledgements

We sincerely thank Junjie He, Xinhua Cheng, Zeyinzi Jiang, Xiaowen Li, Wei Wang ~1, Tianyi Gui, Xiaoyi Bao, Wenting Shen, Tianxing Wang, Ang Wang, Weize Duan, Wei Wang ~2, Zhi-Fan Wu, Chaojie Mao, and Lei Shang for their valuable support and contributions to this work.

Citation

@misc{zhu2026wanpecinematicpromptenhancement, title={WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation}, author={Yubo Zhu and Yawen Shao and Ziyun Dai and Zixun Fang and Kai Zhu and Siyang Sun and Haolan Xue and Chuxin Wang and Tingyu Weng and Jingming Luo and Chen Shi and Lianghua Huang and Yufeng Ai and Yuzheng Wang and Wenyuan Zhang and Yu Shang and Yuxiang Bao and Zoubin Bi and Jie Xiao and Jinbo Xing and Jiaxing Zhao and Chongyang Zhong and Hengjian Chen and Chenwei Xie and Akide Liu and Zhehan Kan and Yu Liu and Wei Zhai and Sheng Zhong and Wei Tong}, year={2026}, eprint={2609.30221}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2609.30221}, }