benchmarksfactbearishA general-purpose helpfulness rubric may not distinguish direct answer-giving from pedagogical guidance in LLM tutoringComputation and Language01 Aug 2026http://arxiv.org/abs/2607.28128v1