benchmarksfactneutralClinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroupsMachine Learning02 Aug 2026http://arxiv.org/abs/2607.28608v1