rlhf
opinion
neutral
Use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts; and MOPD when preserving domain-specific gains matters more than surpassing teachers
use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost
Computation and Language30 Aug 2026