multimodal
fact
bullish
A self-supervised framework can learn joint visuospatial representations from RGB-D observations that integrates depth-derived geometric priors with visual backbones
We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations.
Computer Vision30 Aug 2026