📚 Study Notes / Home / Neural Nets / Exam Notes
Neural Nets · Exam Notes

Concise exam notes — straight from the slides

Distilled from the actual lecture decks: only what you need to pass — definitions, formulas, the key intuition, and likely exam questions. No fluff. Each note also links its original slide PDF.

📝 14 topics (Weeks 1–7) ⏱ ~80 min total 📚 Full detailed sessions 🌐 Source slides
W1S1

Why Deep Learning?

⏱ 5 min read · ⚡ concise

Hand-crafted features vs learned features, universal approximation, representation learning.

Revise →
W2S1

Perceptron to MLP

⏱ 5 min read · ⚡ concise

Neuron, activation functions, multi-layer perceptron, forward pass, the XOR problem.

Revise →
W2S2

Backpropagation & Computational Graphs

⏱ 7 min read · ⚡ concise

Chain rule, computational graphs, local gradients, the backprop algorithm, gradient flow.

Revise →
W2S3

Loss Functions & Optimization

⏱ 6 min read · ⚡ concise

MSE, cross-entropy, SGD, momentum, Adam, learning-rate schedules.

Revise →
W3S1

Regularization

⏱ 5 min read · ⚡ concise

Overfitting, L1/L2 weight decay, dropout, early stopping, augmentation, batch norm.

Revise →
W3S2

Initialization, Normalization & Debugging

⏱ 7 min read · ⚡ concise

Xavier/He init, vanishing/exploding gradients, BatchNorm/LayerNorm, gradient checking.

Revise →
W4S1

Convolutions

⏱ 5 min read · ⚡ concise

Local connectivity, filters, padding, stride, receptive field, pooling, output-size formula.

Revise →
W4S2

CNN Architectures

⏱ 7 min read · ⚡ concise

LeNet, AlexNet, VGG, Inception, ResNet + skip connections — and each one's key idea.

Revise →
W5S1

Transfer Learning & CNN Applications

⏱ 5 min read · ⚡ concise

Feature extraction vs fine-tuning, when to use which, detection, segmentation.

Revise →
W5S2

Sequence Modeling: RNNs

⏱ 5 min read · ⚡ concise

Hidden-state recurrence, BPTT, vanishing/exploding gradients, sequence task types.

Revise →
W6S1

LSTMs, GRUs & Gating

⏱ 5 min read · ⚡ concise

Cell state, forget/input/output gates, how gating fixes vanishing gradients, GRUs, BiRNNs.

Revise →
W6S2

Seq2Seq & Attention

⏱ 5 min read · ⚡ concise

Encoder-decoder, the information bottleneck, attention, alignment scores, context vector.

Revise →
W7S1

The Transformer & Self-Attention

⏱ 6 min read · ⚡ concise

Self-attention Q/K/V, scaled dot-product, multi-head, positional encoding, why it beats RNNs.

Revise →
W7S2

Transformers in Practice — BERT & GPT

⏱ 7 min read · ⚡ concise

BERT (masked LM, encoder-only), GPT (autoregressive, decoder-only), pre-training, BPE/WordPiece.

Revise →