multimodal
fact
neutral
Explicit visual intermediates can help multimodal large language models externalize spatial evidence and updated visual states, but their utility depends on the image editor's ability to faithfully realize transformations
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation.
Computer Vision29 Aug 2026