Reinforcement Learning / Fall 2026
Updates
- New Assignment released: [Assignment #2 - Tabular RL]
- New Lecture is up: Lecture 04: Generalized Policy Iteration via Monte-Carlo and Temporal Difference
- New Lecture is up: Lecture 03: Policy and Value Iteration
- New Assignment released: [Assignment #1 - Basics of RL]
- New Lecture is up: Lecture 2: MDPs and Bellman
- New Lecture is up: Lecture 1: RL Framework
- New Lecture is up: Lecture 0: Course Overview and Logistics
For the Quercus page of the course please click here
Course Description
This course develops fundamental understanding and hands-on skills in reinforcement learning. The course is designed in three major parts: Part I gives the students a warm welcome by taking them through the basic definitions and fundamental concepts. Part II explains fundamental reinforcement learning methods by touching the key model-based and model-free techniques and providing deep understanding of these methods. Part III explores deep reinforcement learning, where deep neural networks are employed to efficiently approximate the developed techniques in Part II. In this part, we take a look into several algorithms, such as deep Q-learning, policy gradient methods, e.g., trust-region and proximal policy optimization algorithms, and actor-critic methods.
Time and Place
Lectures
Lectures start on September 10, 2025.
| Day | Time | Place |
| Thursdays | 1 PM - 4 PM | BA-1170 - Bahen Centre for Information Technology |
Tutorials
Tutorials sessions start on September 18, 2025.
| Day | Time | Place |
| Fridays | 5 PM - 6 PM | BA-1130 - Bahen Centre for Information Technology |
Course Office Hours
| Day | Time |
| Thursdays | 4 PM - 5 PM |
Previous Offerings
Course Description
This course provides a concrete understanding of reinforcement learning and its applications. The ultimate goal of the course is to develop hands-on skills in deep reinforcement learning, for which fundamentals of reinforcement learning are first discussed and then deep reinforcement learning algorithms are studied. The course is designed in three major parts: Part I introduces basic definitions and fundamental concepts; Part II covers fundamental reinforcement learning methods (model-based and model-free) with a deep understanding of each; Part III explores deep reinforcement learning, employing deep neural networks to approximate the techniques in Part II (e.g., function approximation, deep Q-learning, policy gradient methods, and proximal policy optimization).
Part I: First Things in Reinforcement Learning
- General framework of reinforcement learning
- The multi-armed bandit problem
- Components: Agent, Environment, State, Action, Reward, Policy
- Comparison to supervised learning
- Value function and policy design
- Problem of Credit Assignment
- Exploration versus Exploitation
- Revisiting the multi-armed bandit problem
- Trade-off between exploration and exploitation
- Introduction to Gymnasium library
- Generating an environment in Gymnasium
- Our first try: a simple game
Part II: Fundamentals of Reinforcement Learning
- Model-based reinforcement learning
- Markov Decision Processes (MDPs)
- Value and policy with MDPs
- Dynamic programming and Bellman equation
- Value iteration and Policy iteration algorithms
- Model-free reinforcement learning
- On-policy versus off-policy approaches for model-free reinforcement learning
- Difference and properties of on-policy and off-policy methods
- On-policy approach 1: Monte-Carlo (MC) learning
- On-policy approach 2: Temporal Difference (TD) learning
- From value function to Q-function
- On-policy approach 3: State-Action-Reward-State-Action (SARSA)
- Off-policy approach: Q-learning
- Revisiting our simple game
- Implementing value and policy iteration in Gymnasium
Part III: Deep Reinforcement Learning
- Reviewing main concepts in deep learning
- Universal approximation theorem
- Deep neural networks
- Training a neural net via gradient descent
- Reviewing neural network implementation in PyTorch
- Preliminaries of Deep Reinforcement Learning
- Function approximation
- Space reduction via function approximation
- Simple function approximator
- Deep neural networks as function approximators
- Looking into a new example
- Deep off-policy reinforcement learning
- Value networks: value function approximation via deep neural networks
- Deep Q-learning and Deep Q-networks (DQNs)
- Properties of deep Q-learning: sample efficiency and instability
- Visiting our new example
- Deep on-policy methods
- Policy networks: policy approximation via deep neural networks
- Policy gradient methods
- Direct policy update
- Properties of deep policy networks: sample inefficiency versus stability
- Trust Region Policy Optimization (TRPO)
- Constraining policy update via Kullback-Leibler divergence
- Idea of surrogate objective function
- Proximal Policy Optimization (PPO)
- Clipping
- Complexity of PPO
- Revisiting our new example
- Actor-Critic Methods
- Advantage Actor Critic (A2C)
- TRPO and PPO with value network
- Deterministic Policy Gradient
- Deep Deterministic Policy Gradient (DDPG)
- Soft Actor Critic (SAC)
- Extensions and modification
- Applications and advancements of deep reinforcement learning
- Looking into some successful examples: Alpha-Go, Alpha-Zero, Pluribus, and OpenAI Five
- Sample applications of deep reinforcement learning and project poster session
- Recent advancements in deep reinforcement learning
Course Evaluation
The learning procedure consists of three components:
| Component | Grade % | |
| Assignments | 40% | 4 assignment sets |
| Exam | 30% | 2 in-term exams |
| Project | 30% | open-ended projects |
Assignment
There will be 4 sets of assignments. Roughly speaking, the first one goes through fundamentals of reinforcement learning. The second assignment gets more serious in terms of implementation and develops your knowledge on tabular reinforcement learning methods. The last two assignments gives you a chance to implement some mini-projects on Deep Reinforcement Learning. The assignments will count for 40% of your final mark.
Exam
We have two in-term exams in this semester. The exams are on Friday, October 23 and Friday, December 4. Each exam will ask two questions similar to what you have seen in assignments, tutorials, and practice questions shared with you. The exam evaluates the understanding of fundamental concepts.
Course Project
An open-ended project is to be completed through semester. Students work in groups. The final projects will be presented in the second week of December.
