robotics
fact
neutral
State-of-the-art action-conditioned video models are restricted to single robot embodiments and cannot leverage heterogeneous video data
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics.
Computer Vision30 Aug 2026