safetyfactbearishTask-specific fine-tuning of LLMs is notoriously plagued by overconfidence, severely hindering trustworthy deploymentMachine Learning27 Jul 2026http://arxiv.org/abs/2607.02182v1