SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Published in Neural Information Processing Systems (NeurIPS) , 2026 [CCF A]

Kai Zeng, Moran Li, Zhengwei Wang, Yingchen Yu, Yiheng Lin, Ruichuan An, Ming Lu, Qi She, Wentao Zhang

Project Page · Paper · Code · Model · Benchmark

SteerVTE is a unified framework for precise video text editing that preserves the original text style while maintaining temporal coherence. It augments a frozen video diffusion transformer with a lightweight Text Context Adapter that combines mask-guided localization, native-resolution style encoding, and dual-granularity glyph control.

The work introduces the Glyph-Aware Spatial-Focal (GLAS) Loss, a progressive image-to-video training curriculum, the SteerVTE-1M training dataset, and VTE-Bench for evaluating text accuracy, style consistency, temporal coherence, and background preservation.