Improving in-context generalization in reinforcement learning through asymmetric model-based methods using adequate representation and exploration techniques
€238K
01 Apr 2026 → 30 Jun 2028
2
organizations
Objective
Reinforcement learning (RL) is an appealing framework for solving decision-making problems, notably because it makes few assumptions about the problem at hand. In its purest form, the promise of an RL algorithm is to learn an optimal behavior from interaction with an unknown environment. There has been a plethora of empirical successes in real-world applications ranging from games to robotics. However, most of these achievements have required a dedicated training, and the learned behaviors have demonstrated limited generalization abilities. Compared to the generalization capabilities recently obtained in generative modeling with large pretrained models, notably in language generation, it appears clearly that RL has much room for improvement when it comes to generalization. We define generalization as the ability to maintain performance by adapting to environment changes (perception, dynamics, or rewards) based on the observable context only. Interestingly, in-context generalization is known to be equivalent to the problem of optimally controlling a partially observable environment. Coincidentally, the RL field has been bursting with discoveries over the last few years, with notable progresses in several domains that are closely related to generalization and partial observability: representation learning, model-based RL, asymmetric RL, and exploration. These advances inspire enthusiasm about the future of RL, and in particular about the idea of developing RL algorithms able to learn behaviors that generalize well. It motivates this research project that will improve generalization by (i) developing new world model architectures for model-based RL with effective imagination, (ii) developing asymmetric representation learning objectives for better convergence and sample-efficiency, (iii) designing suitable exploration strategies relying on the aforementioned representations, and (iv) benchmarking generalization on a real-world application, tertiary voltage control.
Click “Summarize” to get an AI-powered analysis of this project.
Call Topics
Consortium(2 organizations)
| Organization | Country | Type | SME | Website |
|---|---|---|---|---|
UNIVERSITE DE LIEGE ULIEGE | BE | HES | — | |
ROYAL INSTITUTION FOR THE ADVANCEMENT OF LEARNING MCGILL UNIVERSITY McGill University | CA | HES | — |