Lectures

Here, you can find the recordings of the lecture videos.

  • Lecture 0: Course Overview and Logistics
    Overview: This lecture goes through the course logistics. Unfortunately, it did not get recorded!
    [link]

    Lecture Notes:

  • Lecture 1: Fundamentals of DL
    DL Components: This lecture goes over preliminaries. We study key components of ML problems, work with the example of classification, and formulate the training of a model via empirical risk minimization. Unfortunately, the lecture was not recorded!
    [link]

    Lecture Notes:

    Further Reads:

    • Motivation: Chapter 1 - Section 1.1 of [BB]
    • Review on Linear Algebra: Chapter 2 of [GYC]
    • ML Components: Chapter 1 - Sections 1.2.1 to 1.2.4 of [BB]
    • Binary Classification: Chapter 5 - Sections 5.1 and 5.2 of [BB]
    • McCulloch-Pitts Model: Paper A logical calculus of the ideas immanent in nervous activity published in the Bulletin of Mathematical Biophysics by Warren McCulloch and Walter Pitts in 1943, proposing a computational model for neuron. This paper is treated as the pioneer study leading to the idea of artificial neuron
    • Overview on Risk Minimization: Paper An overview of statistical learning theory published as an overview of his life-going developments in ML in the IEEE Transactions on Neural Networks by Vladimir N. Vapnik in 1999
  • Lecture 2: Deep NNs and Gradient Descent
    NNs: We talk about Universal Approximation by NNs, this motivates us to use Deep NNs for ML. We then formally define Deep NNs and get some basics on MLPs. Finally, we study the gradient descent algorithm which is used for training of Deep NNs.
    [link]

    Lecture Notes:

    Further Reads:

  • Lecture 08: Multiclass Classification and Backpropagation
    Forward Pass: We review neural classification using an MLP. We see that cross-entropy is a better loss helping us make differentiable risk. We can deploy that by interpreting the output of a neural network to be probability of each label. At the second part of the course, we discuss computation graphs. We see that the flow of information in a NN is similar to a computation graph. We could use this fact to build an algorithmic approach based on chain-rule to compute the gradient of the loss with respect to each weight in the network, this algorithm is called Backpropagation. We develop this algorithm for an MLP and see that it's dual to the forward pass.
    [link]

    Lecture Notes:

    Further Reads:

    • Deep FNNs: Chapter 6 - Sections 6.3 and 6.4 of [GYC]
    • Backpropagation: Chapter 8 of [BB]
    • Backpropagation of Error Paper Learning representations by back-propagating errors published in Nature by D. Rumelhart, G. Hinton and R. Williams in 1986 advocating the idea of systematic gradient computation of a computation graph
  • Lecture 4: Optimizers and Generalization
    Optimizer & Overfitting: We study the SGD algorithm. We see that using mini-batches we can control the tradeoff between variance and complexity of gradient computation. We then take a look at extensions of SGD, namely momentum method and Rprop, RMSprop, and Adam. In the second part, we study the overfitting, its main sources, and practical means to handle it, i.e., cross-validation, data augmentation, and regularization.
    [link]

    Lecture Notes:

    Further Reads:

    • SGD: Chapter 5 - Section 5.9 of [GYC]
    • Regularization: Chapter 7 of [GYC]
    • Learning Rate Scheduling Paper Cyclical Learning Rates for Training Neural Networks published in Winter Conference on Applications of Computer Vision (WACV) by Leslie N. Smith in 2017 discussing learning rate scheduling
    • Rprop Paper A direct adaptive method for faster backpropagation learning: the RPROP algorithm published in IEEE International Conference on Neural Networks by M. Riedmiller and H. Braun in 1993 proposing Rprop algorithm
    • Dropout 1 Paper Improving neural networks by preventing co-adaptation of feature detectors published in 2012 by G. Hinton et al. proposing Dropout
    • Dropout 2 Paper Dropout: A Simple Way to Prevent Neural Networks from Overfitting published in 2014 by N. Srivastava et al. providing some analysis and illustrations on Dropout