rlhf
critique
bearish
Most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text-to-image generation and instruction-based image editing, where different inputs require different evaluation dimensions.
Computer Vision28 Aug 2026