multimodal
fact
bullish
A pretrained image animation model that naturally decouples appearance from motion can be repurposed as a base model for instruction-based human-centric video editing
we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset
Computer Vision30 Aug 2026