agentscritiquebearishVision-language models as judges of computer-using agent trajectories have not been systematically evaluated for reliabilityComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28609v1