This course aims to provide an overview of online learning, bandits, and RL. This is a part of “人工智能的优化方法前沿讲习班” (2026 年国家自然科学基金数学天元基金项目)。
Instructor: Peng Zhao ([email protected])
Location: 教二327,南京邮电大学(仙林校区)
| Index | Date | Time | Topic | Slides |
|---|---|---|---|---|
| 1 | 07.14 | 09:30-11:30 | Online Learning/Optimization | Lecture 1 |
| 2 | 07.15 | 09:30-11:30 | Online Mirror Descent | Lecture 2 |
| 3 | 07.15 | 14:30-16:30 | Multi-Armed Bandits | Lecture 3 |
| 4 | 07.16 | 09:30-11:30 | Upper Confidence Bound | Lecture 4 |
| 5 | 07.16 | 14:30-16:30 | Online MDPs and One-Pass Bandits | Lecture 5 |
Familiar with calculus, probability, and linear algebra. Basic knowledge in convex optimization and machine learning.
This mini-course is a crash short version of the Advanced Optimization (AOPT) course developed for graduate students at the School of Artificial Intelligence in Nanjing University. You can use the links below to access lecture slides and related materials from my AOPT courses in the last four years.While the overall structure remains consistent each year, I continually refine the content by adding new topics and improving the logical flow based on the feedback and my latest research understanding.
Advanced Optimization (For Undergraduate and Graduate Students, 2023 Fall)
Advanced Optimization (For Undergraduate and Graduate Students, 2022 Fall)
Unfortunately, we don't have a specific textbook for this course. In addition to the course slides, the following books are very good materials for extra readings.
Francesco Orabona. Online Learning: A Modern Introduction Using Convex Optimization. Cambridge University Press, 2026.
Tor Lattimore and Csaba Szepesvári. Bandit Algorithms. Cambridge University Press, 2021.
Amir Beck. First-Order Methods in Optimization. MOS-SIAM Series on Optimization, 2017.
Elad Hazan. Introduction to Online Convex Optimization (second edition). MIT Press, 2022.
Last modified: 2026-07-15 by Peng Zhao