multimodalfactbullishDetailed long captions can enhance visual perception capabilities of vision-language models and alleviate visual ambiguities in monocular depth estimationComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28285v1