Skip to main navigation Skip to search Skip to main content

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

  • Richa Verma*
  • , Bavish Kulur
  • , Sanjay Chawla
  • , Balaraman Ravindran
  • *Corresponding author for this work
  • Indian Institute of Technology Madras
  • University of Alberta

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it from scratch. While costs could be numerically encoded, we assume a more general setting is when costs are provided as preferences. Given a reward-optimized policy and a small dataset of preferred (low-cost) and dispreferred (high-cost) trajectories, our goal is to fine-tune the policy to generate low-cost behaviors while retaining high rewards. Unlike standard RLHF in language models, where preferences are defined over responses to the same prompt, our setting involves trajectory-level preferences in continuous control environments. We introduce PREFINE: Preference-based Implicit Reward and Cost Fine-Tuning for Safety Alignment which is a preference-based fine-tuning method that adapts Direct Preference Optimization (DPO), which is now widely used for LLM fine-tuning, to the sequential decision making setting. PREFINE constructs policy-sampled counterfactual trajectories to establish meaningful preference contrasts and jointly optimizes for reward retention and safety alignment. Empirically, PREFINE reduces constraint violations and catastrophic failures by over 60% while maintaining original reward behavior. PREFINE produces policies that achieve low-cost, high-reward performance with significantly improved data and computational efficiency compared to full offline RL or imitation learning, bridging preference alignment and safe policy adaptation in continuous domains.

Original languageEnglish
Title of host publicationAAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
PublisherAssociation for Computing Machinery, Inc
Pages1004-1012
Number of pages9
ISBN (Electronic)9798400723179
DOIs
Publication statusPublished - 24 May 2026
Event25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026 - Paphos, Cyprus
Duration: 25 May 202629 May 2026

Publication series

NameAAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems

Conference

Conference25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026
Country/TerritoryCyprus
CityPaphos
Period25/05/2629/05/26

Keywords

  • Policy adaptation
  • Preference-based Optimization
  • Safe Reinforcement Learning

Fingerprint

Dive into the research topics of 'PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment'. Together they form a unique fingerprint.

Cite this