benchmarkscritiquebearishExact match is too brittle, text similarity ignores structure, and LLM judges are expensive, opaque, and non-deterministic for evaluating JSON outputsComputation and Language27 Jul 2026http://arxiv.org/abs/2607.01972v1