multimodalfactbullishVideo re-shooting can be achieved without explicit 3D priors or paired training data by using text-driven semantic viewpoint specification and self-supervised learning of camera dynamicsComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28261v1