benchmarks
fact
bearish
Double-difference designs used to audit LLM judge bias are not properly identified on bounded rating scales, confounding differential preference with differential attenuation
We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation
Computation and Language30 Aug 2026