CS7.501 Advanced NLP

Monsoon 2025

Course Logo

Course Overview ↑ Top

Designed for those seeking an advanced understanding, this course explores in-depth topics in Natural Language Processing (NLP). It assumes foundational knowledge and may not suitable for newcomers to the field. While specific prerequisites are not strictly enforced, prior experience is highly encouraged.

Instructor

Teaching Assistants

Logistics ↑ Top

All course related annoucements will be made on Moodle.

Teaching Assistant Office Hours

TA Name Office Hours
Aparajitha Monday: 13:30 - 14:30 (MT NLP Lab)
Saumitra Thursday: 10:30 - 11:30 (MT NLP Lab)
Maitreya Friday: 9:00 - 10:00 (Appointments must be requested via email by Thursday, 9:00 PM)
Ketaki Monday: 14:00 - 15:00
Vivek Monday: 10:00 - 11:00
Raveesh Thursday: 14:00 - 15:00

Course Schedule ↑ Top

# Date Topic Section Resources
1 31-Jul-25 Course Introduction Logistics, Introduction Post-Reading: Neural Machine Translation by Jointly Learning to Align and Translate
Attention is all you need
2 4-Aug-25 Recap Transformers Post-Reading: Illustrated Transformer
Self Attention, Transformers
Intro to Transformers
Attention in Transformers
Dolma
Training Compute-Optimal Large Language Models
3 7-Aug-25 Transformers Attention Pre-Reading: Self Attention from Scratch

Post-Reading: Longformer
MQA
GQA
DeepSeekV2 - MHLA
Flash Attention
Ring Attention
Log Linear Attention
4 11-Aug-25 Feed Forward Layers Pre-Reading:Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps

Post-Reading: Hands-on: Mixture of Experts with Transformers
5 14-Aug-25 Positional Embedding Pre-Reading:What do position embeddings learn
Rethinking positional encoding in language pretraining

Post-Reading: Rotary Position Embedding
6 18-Aug-25 Normalization Pre-Reading: Backpropogation
Residual Network
Backpropagation
Yes you should understand backprop

Post-Reading: Layer Norm
Batch Norm
Weight Norm
RMS Norm
Pre Norm - On Layer Normalization in the Transformer Architecture
Normalization in Deep Learning
7 21-Aug-25 Optimizers Pre-Reading: Gradient Descent

Post-Reading: Optimizers
Adam
AdaGrad
RMSProp
Muon
Muon is Scalable for LLM Training
Practical Efficiency of Muon for Pretraining
8 25-Aug-25 Decoding & Generation Introduction to Decoding Pre-Reading: Autoregressive Models

Post-Reading: Hierarchical Neural Story Generation
Beam Search Strategies
Thorough Examination of Decoding Strategies in LLM Era
9 28-Aug-25 Speculative Decoding Post-Reading:
Speculative Decoding for Seq2Seq Generation
Speculative Decoding
Speculative Sampling
A Survey on Speculative Decoding
Looking back at speculative decoding
10 1-Sep-25 Recent Decoding Strategies Post-Reading:
Min-p Sampling
Lookahead Decoding
11 8-Sep-25 Parallel Decoding Post-Reading:
Accelerating Transformer Inference for Translation via Parallel Decoding
Lookahead Decoding
Medusa
12 11-Sep-25 Accelerated Inference Inference Optimizations Post-Reading:
Inference-time Algorithms
Large Language Monkeys
13 15-Sep-25 Evaluation Evaluation Metrics Post-Reading:
BLEU
ROUGE
BERTScore
COMET
MEE4
14 18-Sep-25 Benchmarking Post-Reading:
GLUE: A Multi-Task Benchmark and Analysis Platform for NLU
GEM Benchmark
Gptscore: Evaluate as you desire
M3t: A new benchmark dataset for multi-modal document-level machine translation
G-eval: Nlg evaluation using gpt-4 with better human alignment
A Comprehensive Survey on Sentence-Level Translation Evaluation
15 29-Sep-25 Quantization Quantization Post-Reading: Visual Guide to Quantization
Smooth Quant
Optimal Brain Compression (OBQ)
LoRA: Low Rank Adaptation
QLoRA: Efficient Finetuning of Quantized LLMs
GPTQ: Accurate Post-Training Quantization
AWQ: Activation-aware Weight Quantization
Zero Quant
LLM.int8
16 6-Oct-25 Distributed Training Distributed Pretraining Post-Reading: Switch Transformer
Distributed training at Scale
Switch Transformer
Megatron-LM
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Large Scale Distributed Training
17 9-Oct-25 Sharding Post-Reading: ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
ZeRO
PyTorch Distributed
NCCL
GPipe
Picotron
18 13-Oct-25 Mixture of Experts Introduction - Conceptual, Architecture Pre-Reading: Adaptive Mixture of Local Experts
Hierarchical Mixture of Experts

Post-Reading: Mamba DeepSeek
DeepSpeed MoE
DeepSeek MoE
Mixture of Experts
A Visual Guide to Mixture of Experts
MoE Explained
19 16-Oct-25 Multimodality Vision Contrastive Learning, Modality Fusion Post-Reading: Learning Transferable Visual Models From Natural Language Supervision
Chameleon
20 23-Oct-25 VLMs - Hybrid/Combined, ViT, LLaVA, PaliGemma Post-Reading: ViT
LLaVA
PaliGemma
Molmo
21 27-Oct-25 Advanced Topics Multimodality - Speech Post-Reading: Whisper
Google USM
F5 TTS
22 30-Oct-25 Graphs Post-Reading: GCN
GAT
GraphSAGE
23 3-Nov-25 RL in NLP, RLHF Post-Reading: TBA
24 6-Nov-25 Guest Lecture Dr. Balaji Srinivasan
25 10-Nov-25 DPO Post-Reading: DPO
26 13-Nov-25 Guest Lecture Dr. Vamshi Ambati
27 17-Nov-25 Project Evaluations
28 20-Nov-25 Reasoning Post-Reading: Training models to follow instructions with human feedback
Scaling Instruction tuned Model
AlpacaFarm

Grading and Weightage ↑ Top

Your final grade in this course will be determined by the following components:

Component Weightage
Assignments * 3 10 + 10 + 10 (30%)
Quiz * 2 10 + 10 (20%)
Project 50%
Total 100%

Due Dates ↑ Top

Important dates for the course:

Component Due Date
Team Details Submission 7-Aug
Project Proposals (Interim) 10-Aug
Proposal (Interim) Rejection by TAs 12-Aug
Project Finalization, Mentor Assignment 15-Aug
A1 Release 17-Aug
Proposal (Final) 26-Aug
Proposal Grades 5-Sep
A1 Due 10-Sep
A2 Release 12-Sep
A1 Grades 21-Sep
Project Mid Submission 2-Oct
A2 Due 10-Oct
Project Mid Grades 12-Oct
A3 Release 12-Oct
A2 Grades 3-Nov
Final Project Due 7-Nov
A3 Due 10-Nov
Q1 Release 12-Nov
Q1 Due 14-Nov
Q2 Release 17-Nov
Q2 Due 19-Nov
A3 Grades 23-Nov
Final Grade Assignment 5-Dec

Course Projects ↑ Top

The course project is a significant component of this course (50%), providing an opportunity to apply the concepts and techniques learned to a real-world NLP problem. Students will work in teams of 3 to propose, develop, and present a project. Projects are expected to be at an advanced level. Submission to conferences/workshops is highly encouraged. Read more on research opportunities.

We encourage students to think creatively and explore areas that genuinely interest them.


Component Weightage
Project Outline 5%
Project Mid 15%
Project Final 30%
Total 50%

Project Areas

  • Efficient/Low-Resource Methods
  • Ethics, Fairness, Bias
  • Generation/Language Modeling
  • Interpretability/Explainability
  • Multilingual/Cross-lingual
  • Multimodal
  • NLP Applications
  • Evaluation/Benchmarking
  • Retrieval & Extraction
  • Document Understanding
  • Graph + LLMs
  • Entertainment + LLM

Assignments and Project Policies ↑ Top

Late Submission Policy

Assignment and project deadlines are strictly enforced. Late submissions will not be allowed and no extensions will be provided.

Collaboration Policy

Collaboration on assignments is encouraged for discussion of concepts and general approaches, but all submitted answers must be your individual work. For projects, group collaboration is expected, and the contribution of each member should be clearly documented. Any specific collaboration guidelines for individual assignments or projects will be provided with the assignment description.

Consultation and Research Opportunities

Students are highly encouraged to consult with the instructor, teaching assistants and mentors during office hours or by appointment for guidance on assignments, projects, or course topics. If a project developed in this course shows significant potential for a research paper, students may, at their discretion, invite a TA, a mentor or the instructor whose advice substantially benefited the work to be a co-author. This is not a requirement, but an opportunity to acknowledge significant contributions and foster academic collaboration.

Other Resources ↑ Top