multimodalcritiquebearishMost existing ZS-CIR methods rely on textual inversion to translate reference images into pseudo-text tokens, which can be lossy and brittle for fine-grained semanticsComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02284v1