Offline reinforcement learning relies on learning policies from previously collected data without further interaction with the environment. A key challenge is accurately estimating uncertainty in learned dynamics models, particularly for state-action pairs that are poorly represented in the dataset. Recent work in our group developed a theoretically motivated uncertainty estimator based on Neural Tangent Kernels (NTKs), Gaussian Processes, and RKHS-based error bounds. While the approach provides promising theoretical foundations, its empirical performance revealed an important trade-off between theoretically justified uncertainty estimates and predictive accuracy.
This MSc project will build on this work by investigating how the proposed uncertainty estimation framework can be improved and made more practically effective. Possible directions include improving the tightness and calibration of the error bounds, enhancing assumptions required by the current theoretical framework, developing alternative approaches to incorporating prior uncertainty, and evaluating the resulting methods in model-based offline reinforcement learning algorithms.
Maryam Tavakol