multimodal
fact
bearish
Current video generation models fail to achieve probabilistic alignment - they cannot reproduce the correct distribution of possible behaviors under the same initial observation and action
Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors.
Computer Vision30 Aug 2026