Distributionally Robust Constrained Reinforcement Learning with Function Approximation
Implementing Organization
Indian Institute Of Technology Kanpur
Principal Investigator
Prof. Washim Uddin Mondal
Indian Institute Of Technology Kanpur
wmondal@iitk.ac.in
Project Overview
Reinforcement Learning (RL) is a widely adopted learning framework where an agent learns to navigate an unknown environment via repeated interaction and receiving feedback in the form of reward/cost. Most works in this setup assume that the environment that trains the policy is identical to the environment where the agent is intended to be deployed. This is not true for many practical cases. For example, in safety-critical applications, the agents are not allowed to interact with the real environment during the training phase, and, therefore, the training must be done solely using an artificial simulator, which may not be an accurate replica of the reality. For a concrete example, consider a car-driving scenario. If the driving agent is allowed to drive on the road during training, catastrophic accidents might occur. How does the policy trained on an inaccurate simulator perform in its intended reality? Distributionally robust RL (DRRL) attempts to answer such questions. Although some recent papers provide sample complexity-based convergence guarantees of DRRL, most only apply to tabular setups. Many real-world applications, however, require a large or infinite state space to be accurately described. How well DRRL algorithms perform with infinite state space is currently underexplored in the literature. On the other hand, one of the potential beneficiaries of DRRL is safety-critical applications. Traditionally, such applications are modeled via constrained RL. Can we incorporate constraints in the DRRL setup and still show its efficacy? Such answers are also unknown in the literature. Our proposal precisely tackles these two questions in a unified distributional robust constrained RL (DRCRL) framework. One way to incorporate infinite state space is via function approximation (also known as general parameterization), where policies can be represented by neural networks. Recent works show that policy-gradient-type results are valid even for robust MDPs. Another recent work shows that Bellman-type updates are valid for DRRL setup. Combining these two insights, we suggest exploring the idea of function approximation-based actor-critic algorithms for DRRL (which is compatible with infinite state space) in this proposal. To tackle the constraints, we suggest exploring the primal-dual-based algorithms that showed tremendous success in PI Mondal's recent NeurIPS papers. We hope that all of these insights together will produce provably convergent algorithms for DRCRL. We will empirically validate the efficacy and robustness of our algorithm by (a) training a driving agent in a commercial traffic simulator such as SUMO and (b) training a wireless transmission strategy in an integrated sensing and communication (ISAC)-enabled test bench.
Disclaimer:
Information available on this portal is sourced from various organizations and is provided for informational purposes only. Users are advised to verify details from the respective official sources.
Please enter your details
Please provide your name and email to continue. Your details are saved in this browser for future use.
Latest Updates
Loading…
⚠️
You are leaving this website
You are about to be redirected to an external website that is not operated by
India Science, Technology & Innovation (ISTI) Portal.