robotics
fact
bullish
CLAP is capable of being trained on diverse, internet-scale videos across human and robotic agents for cross-embodiment action-conditioned video generation
we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents.
Computer Vision30 Aug 2026