Distributionally Robust Average Reward Reinforcement Learning
Project Description
Distributionally robust average-reward reinforcement learning studies sequential decision-making problems in which the objective is to maximize long-run average performance under uncertainty about the underlying transition dynamics or reward distributions. While distributionally robust reinforcement learning has been studied primarily in discounted and finite-horizon settings, its theoretical foundations and computational methods under the average-reward criterion remain relatively underdeveloped. Nevertheless, this formulation is particularly relevant for continuing systems that operate indefinitely and face model misspecification or distributional shifts, including queueing systems, inventory control, supply chains, and online platforms. In this project, students will learn the necessary background in average-reward Markov decision processes and distributionally robust optimization, investigate the statistical and structural properties of robust average-reward policies, and develop scalable algorithms supported by numerical experiments. Due to the technical complexity of the project, advanced proficiency in real analysis and Markov chains is required.
Supervisor
SI, Nian
Quota
1
Course type
UROP1100
Applicant's Roles
Read relevant literature and become familiar with the theoretical background of the topic.
Investigate the statistical complexity of the proposed methods.
Implement and evaluate algorithms through numerical experiments.
Investigate the statistical complexity of the proposed methods.
Implement and evaluate algorithms through numerical experiments.
Applicant's Learning Objectives
Understand the fundamentals of distributionally robust average-reward reinforcement learning and its distinction from discounted and finite-horizon settings.
Develop a solid grasp of the theoretical foundations, including convergence and statistical properties.
Gain hands-on experience in implementing and evaluating scalable RL algorithms.
Develop a solid grasp of the theoretical foundations, including convergence and statistical properties.
Gain hands-on experience in implementing and evaluating scalable RL algorithms.
Complexity of the project
Challenging