EcoSta 2026: Start Registration
View Submission - EcoSta2026
A1704
Title: Balanced policy switching in reinforcement learning Authors:  Tao Ma - University College London (United Kingdom) [presenting]
Abstract: Reinforcement learning (RL) aims to learn optimal policies that maximize long-term cumulative reward and has achieved widespread success across domains. In many real-world applications, however, updating or switching policies incurs non-negligible costs, making frequent changes undesirable. This challenge arises both in offline settings, where decisions rely solely on historical data, and in online scenarios where policy updates must balance improvement against switching costs. A unified framework for policy switching in both offline and online RL is introduced, leveraging ideas from optimal transport to propose a principled formulation that captures the trade-off between performance gains and switching costs. Key theoretical properties are established and a Net Actor-Critic algorithm is developed for the resulting problem. Experiments on robot control tasks and traffic signal management demonstrate the effectiveness of the approach.