benchmarkscritiquebearishLLM-as-a-judge for GenUI evaluation is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users perceive interfacesComputation and Language02 Aug 2026http://arxiv.org/abs/2607.28439v1