Online Learning, Bandit, and RL (2026 Summer)

Course Information

This course aims to provide an overview of online learning, bandits, and RL. This is a part of “人工智能的优化方法前沿讲习班” (2026 年国家自然科学基金数学天元基金项目)。

Course agenda

IndexDateTime
Topic
Slides
107.1409:30-11:30Online Learning/OptimizationLecture 1
207.1509:30-11:30Online Mirror DescentLecture 2
307.1514:30-16:30Multi-Armed BanditsLecture 3
407.1609:30-11:30Upper Confidence BoundLecture 4
507.1614:30-16:30Online MDPs and One-Pass BanditsLecture 5

Prerequisites

Familiar with calculus, probability, and linear algebra. Basic knowledge in convex optimization and machine learning.

This mini-course is a crash short version of the Advanced Optimization (AOPT) course developed for graduate students at the School of Artificial Intelligence in Nanjing University. You can use the links below to access lecture slides and related materials from my AOPT courses in the last four years.While the overall structure remains consistent each year, I continually refine the content by adding new topics and improving the logical flow based on the feedback and my latest research understanding.

Reading

Unfortunately, we don't have a specific textbook for this course. In addition to the course slides, the following books are very good materials for extra readings.



Last modified: 2026-07-15 by Peng Zhao