multimodalcritiquebearishPrior language-integrated monocular depth estimation methods fail to fully harness language potential due to short text input, coarse feature learning, and limited guidanceComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28285v1