multimodalfactbullishDecoupling 3D geometric reconstruction from semantic integration enables more efficient open vocabulary 3D scene understanding from monocular videoComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28300v1