Hi all,
Today's Learning Theory Circle session:
Speaker: M.Mahdi Mostafaei (Shahid Beheshti University)
Paper: Why Does SGD Prefer Flat Minima? Through the Lens of Dynamical Systems
Paper link:
https://openreview.net/forum?id=Zffca0v-i5SWhen: Sunday, August 16 · 7:00 PM Tehran (UTC+3:30)
Where:
https://meet.google.com/csa-sbkw-ctmAbstract
The paper shows that stochastic gradient descent (SGD) escapes from sharp minima exponentially fast even before SGD reaches a stationary distribution.
Everyone is welcome — no preparation needed, and questions are encouraged
throughout.
Full schedule and past talks:
https://learning-theory-circle.github.io/Would you like to present a future session? Just reply to this email.
See you,
Erfan