agentscritiquebearishFinal verifier success metrics are too coarse for evaluating agentic skill-use because agents may succeed through trial-and-error while making errors in skill selection and compositionComputation and Language27 Jul 2026http://arxiv.org/abs/2607.01874v1