agents
fact
bullish
Learned preference functions in multi-objective environments exhibit contextual priority switching, graded trade-offs, and temporal persistence, outperforming fixed-preference and handcrafted-preference strategies
Experiments in self-constructed multi-objective exploration environments show that the learned preference function exhibits contextual priority switching, graded trade-offs, and temporal persistence, and outperforms the evaluated fixed-preference and handcrafted-preference strategies.
Machine Learning30 Aug 2026