Offline Reinforcement Learning (RL) addresses settings where online interaction is impractical, costly, or unsafe, enabling applications from healthcare to robotics. Learning from offline data is challenging due to distributional shift, which causes extrapolation errors that cannot be corrected without further exploration. Model-free RL methods regularize policies to stay close to the behavior policy but often generalize poorly due to sample complexity. Model-based approaches, which first learn an empirical MDP from offline data and then optimize policies via simulated interactions, improve sample efficiency. Hence, the ongoing research is focused on building more reliable world models for policy search.
Existing approaches often rely on accurate uncertainty estimation of the dynamics model and incorporate the uncertainty by penalizing the reward or value function, which can make the learned policy unnecessarily conservative. This project will investigate a different use of dynamics uncertainty for the state transitions itself (see [1]) to determine whether an imagined transition should be trusted, corrected, or halted. This will involve investigating whether controlling the transition process itself can reduce compounding model error while retaining useful model-generated experience outside the observed data distribution.
1. Guo, K., Shao, Y., & Geng, Y. (2022). Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics Belief. NeurIPS 2022.
Maryam Tavakol