benchmarks
fact
neutral
Human evaluation for non-verifiable tasks is reliable but expensive, while automatic metrics are scalable but often biased
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased.
Machine Learning (Statistics)29 Aug 2026