multimodal
fact
neutral
Modern vision foundation models are trained almost exclusively on RGB images despite many embodied systems having access to explicit depth sensing that provides geometric information monocular inputs cannot recover
While modern vision foundation models are trained almost exclusively on RGB images, many embodied systems have access to explicit depth sensing, which provides geometric information that monocular inputs cannot recover.
Computer Vision30 Aug 2026