safetyfactneutralDifferent open instruction-tuned models show distinct policies under skeptical pressure on consensus science topics, with Llama showing reactive assertion rather than false balanceComputation and Language27 Jul 2026http://arxiv.org/abs/2607.01951v1