benchmarksfactneutralPosition bias in multiple-choice LLM evaluation is only statistically detectable within a 60-95% base-accuracy rangeComputation and Language28 Jul 2026http://arxiv.org/abs/2607.20864v1