TY - GEN
T1 - PREFINE
T2 - 25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026
AU - Verma, Richa
AU - Kulur, Bavish
AU - Chawla, Sanjay
AU - Ravindran, Balaraman
N1 - Publisher Copyright:
© 2026 International Foundation for Autonomous Agents and Multiagent Systems.
PY - 2026/5/24
Y1 - 2026/5/24
N2 - We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it from scratch. While costs could be numerically encoded, we assume a more general setting is when costs are provided as preferences. Given a reward-optimized policy and a small dataset of preferred (low-cost) and dispreferred (high-cost) trajectories, our goal is to fine-tune the policy to generate low-cost behaviors while retaining high rewards. Unlike standard RLHF in language models, where preferences are defined over responses to the same prompt, our setting involves trajectory-level preferences in continuous control environments. We introduce PREFINE: Preference-based Implicit Reward and Cost Fine-Tuning for Safety Alignment which is a preference-based fine-tuning method that adapts Direct Preference Optimization (DPO), which is now widely used for LLM fine-tuning, to the sequential decision making setting. PREFINE constructs policy-sampled counterfactual trajectories to establish meaningful preference contrasts and jointly optimizes for reward retention and safety alignment. Empirically, PREFINE reduces constraint violations and catastrophic failures by over 60% while maintaining original reward behavior. PREFINE produces policies that achieve low-cost, high-reward performance with significantly improved data and computational efficiency compared to full offline RL or imitation learning, bridging preference alignment and safe policy adaptation in continuous domains.
AB - We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it from scratch. While costs could be numerically encoded, we assume a more general setting is when costs are provided as preferences. Given a reward-optimized policy and a small dataset of preferred (low-cost) and dispreferred (high-cost) trajectories, our goal is to fine-tune the policy to generate low-cost behaviors while retaining high rewards. Unlike standard RLHF in language models, where preferences are defined over responses to the same prompt, our setting involves trajectory-level preferences in continuous control environments. We introduce PREFINE: Preference-based Implicit Reward and Cost Fine-Tuning for Safety Alignment which is a preference-based fine-tuning method that adapts Direct Preference Optimization (DPO), which is now widely used for LLM fine-tuning, to the sequential decision making setting. PREFINE constructs policy-sampled counterfactual trajectories to establish meaningful preference contrasts and jointly optimizes for reward retention and safety alignment. Empirically, PREFINE reduces constraint violations and catastrophic failures by over 60% while maintaining original reward behavior. PREFINE produces policies that achieve low-cost, high-reward performance with significantly improved data and computational efficiency compared to full offline RL or imitation learning, bridging preference alignment and safe policy adaptation in continuous domains.
KW - Policy adaptation
KW - Preference-based Optimization
KW - Safe Reinforcement Learning
UR - https://www.scopus.com/pages/publications/105041466129
U2 - 10.65109/SDRB4374
DO - 10.65109/SDRB4374
M3 - Conference contribution
AN - SCOPUS:105041466129
T3 - AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
SP - 1004
EP - 1012
BT - AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
PB - Association for Computing Machinery, Inc
Y2 - 25 May 2026 through 29 May 2026
ER -