PoEM: Predicting RL Outcomes from Existing Policies
Updated
Updated · arxiv.org · Sep 24
PoEM: Predicting RL Outcomes from Existing Policies
1 articles · Updated · arxiv.org · Sep 24
Summary
Researchers have introduced PoEM, a framework that predicts reinforcement learning (RL) outcomes for new rewards without running additional RL training.
PoEM uses models already post-trained on other rewards to approximate the target RL policy by composing existing single-reward adapters at inference time.
This approach could significantly reduce computational costs and enable rapid evaluation of new reward functions, especially when RL training is expensive or unstable.