multimodalfactbullishAlign4D provides a flexible framework that translates any-modal input into coherent video-3D pairs for 4D generationComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02516v1