Course Materials
# Module 1
Getting Started: Google Colab and GitHub
Learning objectives: Open, run, and save a notebook in Google Colab; navigate a GitHub repository and open a notebook from it; describe the course workflow for accessing, editing, and retaining copies of code.
Before class
- Sign in to Google ColabOptional
During class
- Practice opening, running, copying, and saving a course notebook.Required
# Module 2
Tokens and Embeddings
Learning objectives: Tokenize short texts and interpret token IDs; use sentence embeddings to compare semantic similarity; distinguish representation models from text-generation models; explain how next-token probabilities produce generated text.
Before class
- Read Chapter 1, “An Introduction to Large Language Models,” and Chapter 2, “Tokens and Embeddings.”Optional
During class
# Module 3
Text Classification I: Encoder Models
Learning objectives: Distinguish binary, multiclass, and multilabel classification problems; apply task-specific and embedding-based classifiers; evaluate predictions using suitable metrics and structured error analysis.
Before class
- Read Chapter 3, “Looking Inside Large Language Models,” and the Chapter 4 sections on task-specific and embedding-based text classification.Optional
During class
After class
# Module 4
Text Classification II: Zero-Shot and Few-Shot Models
Learning objectives: Compare NLI-based zero-shot classification with prompted generative classification; use label descriptions, few-shot examples, and constrained outputs; test how label wording and examples affect classification performance.
Before class
- Read Chapter 4, “Text Classification,” from “What If We Do Not Have Labeled Data?” through “Text Classification with Generative Models,” including the Flan-T5 and ChatGPT classification examples.Optional
During class
# Module 5
Text Classification III: Evaluation and Error Analysis
Learning objectives: Compare classifiers on the same held-out examples using aggregate and per-class results; categorize errors and disagreements with reference to the original text; recommend an approach while distinguishing development decisions from final evaluation.
Before class
- Chapter 4, “Text Classification”: metrics in “Using a Task-Specific Model”Optional
During class
After class
# Module 6
Text Clustering and Topic Modeling
Learning objectives: Create document embeddings for an unlabeled text collection; cluster and visualize semantically related documents; interpret topic representations while identifying instability, labeling choices, and other limitations.
Before class
- Chapter 5, “Text Clustering and Topic Modeling”: through the BERTopic discussion before “Adding a Special Lego Block”Optional
During class
# Module 7
RAG I: Dense Retrieval and Grounded Generation
Learning objectives: Trace how a question moves through chunking, embeddings, retrieval, and answer generation; use a scaffolded dense-retrieval workflow; construct a grounded-answer prompt and check whether cited passages support the answer.
Before class
- Chapter 8: semantic search overview, opening dense-retrieval example, and “From Search to RAG”Optional
During class
After class
# Module 8
RAG II: Evaluation and Improvement
Learning objectives: Distinguish retrieval failures from generation failures; assess evidence relevance, answer correctness, groundedness, and unanswerable questions; compare one retrieval or prompt change using fixed examples and report its tradeoffs.
Before class
- Review two question–passage–answer examples from Module 7 and bring one proposed improvement.Required
- Chapter 8, “Retrieval Evaluation Metrics”: queries and relevance judgments; calculations optionalOptional
# Module 9
Context Design and Model Benchmarks
Learning objectives: Select and organize instructions, examples, and source evidence for a task; interpret a model benchmark in terms of its tasks, metrics, and testing conditions; explain what additional application-specific evidence is needed before choosing a model.
Before class
- Chapter 6, “Prompt Engineering”: “The Basic Ingredients of a Prompt” and “Instruction-Based Prompting”Optional
During class
- Context-design exercise: revise a RAG prompt to separate instructions, evidence, and output; propose a fair comparison.Required
# Module 10
LLM Workflows and Tool Use
Learning objectives: Distinguish a fixed LLM workflow from model-directed tool use; trace a tool request, execution, observation, and final response; diagnose a failed task and propose a concrete safeguard or stopping condition.
Before class
- Chapter 7: introductions to chains and agents through ReAct; skip the LangChain implementationOptional
# Module 11
Model Adaptation
Learning objectives: Distinguish changes to prompts and retrieved evidence from updates to model weights; explain the purpose of parameter-efficient fine-tuning; assess a before-and-after adaptation comparison and decide whether adaptation is justified for a project.
Before class
- Bring one project result, one failure example, and one question for the project clinic.Required
- Chapter 12, “Fine-Tuning Generation Models”: SFT, full fine-tuning, and PEFT through LoRA; stop before QLoRAOptional
During class
After class
- Chapter 10, “What Is Contrastive Learning?” or Chapter 11, “SetFit: Efficient Fine-Tuning with Few Training Examples”Optional
# Module 12
Project Presentations
Learning objectives: Present the project question, data, workflow, evaluation, and results as a coherent analytical story; use clear visuals or examples to explain model behavior; discuss limitations and responsible use and respond thoughtfully to questions and peer feedback.
Before class
- Check that your GitHub repository and README are ready for questions.Required
- Submit presentation materials through Canvas by the published deadline.Required