benchmarkscritiquebearishExisting art-focused evaluations rely on synthetic questions and rarely report item-level propertiesComputation and Language27 Jul 2026http://arxiv.org/abs/2607.02007v1