multimodal
critique
bearish
Existing video diffusion model methods for image-to-scene generation rely on incomplete conditioning signals, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency
Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency.
Computer Vision30 Aug 2026