Introduction to VLMs & Core
Foundations
Module 1:
Understanding Vision-Language Models
- What are VLMs?
- How VLMs use images, text, video
- Real-world applications
(captioning, Q&A, document understanding)
Module 2:
Modern Open-Source VLMs
- Overview of Qwen 3 VL
- Overview of Gemma 3 Vision
- Overview of Kimi-VL
- Capabilities comparison
Module 3:
Architecture
- Vision encoder (“eyes”)
- Language model (“brain”)
- Fusion mechanism (“how eyes talk to
brain”)
- Why modern VLMs support video &
long-context
Module 4:
Data Fundamentals
- What is an image–text pair?
- What is a video–text pair?
- Basics of captions and annotations
- Introduction to publicly available
datasets (COCO, WebVid, small open sets)
Module 5:
Simple Dataset Preparation
- Preparing a small image dataset
- Writing simple captions
- Basic image preprocessing (resize,
normalize)
- Basic video preprocessing (extract
frames, understanding frame sampling)
Module 6:
Running Your First VLM (Hands-on)
- Loading Qwen 3 VL / Gemma 3 Vision
/ Kimi-VL from Hugging Face
- Performing:
- image captioning
- image Q&A
- simple video understanding
- Understanding outputs
Fine-Tuning
(Images + Simple Video)
Module 7:
Introduction to Fine-Tuning
- What is fine-tuning?
- Why beginners use small datasets
- Full fine-tuning vs LoRA (simple
explanation)
- Why LoRA is preferred for beginners
Module 8:
Fine-Tuning on Images
- Creating a small training set
- Using PEFT/LoRA with open-source
VLMs
- Training loop (simple walkthrough)
- Observing model improvements
visually
- Avoiding overfitting on tiny
datasets (basic tips)
Module 9:
Evaluating Fine-Tuned Models
- Understanding good vs bad captions
- Simple evaluation:
- manual inspection
- intuition behind BLEU, CIDEr (no
math)
- Saving and loading fine-tuned
models
Module 10:
Intro to Video Fine-Tuning
- Why video = “many images”
- Extracting short clips
- Simple frame sampling
- Preparing video–text pairs
Module 11:
Fine-Tuning Kimi-VL or Qwen 3 VL on Small Video Tasks
- Using 3–10 video clips for
demonstration
- Simple LoRA-based video fine-tuning
- Testing video Q&A or video
captioning
- Discussing common mistakes
Evaluation, Deployment &
Responsible AI
Module 12:
Model Evaluation
- What is a benchmark?
- Why we evaluate models
- High-level introductions:
- Comparing outputs from Qwen, Gemma,
Kimi
Module 13:
Error Analysis & Debugging Basics
- Identifying hallucinations
- Identifying misinterpretations in
images
- Identifying temporal errors in
video
- Dataset quality issues
Module 14:
Deployment Basics
- Running models locally (CPU/GPU)
- Running models in Google Colab
- Exporting models
- Creating simple demos with Hugging
Face Spaces
- Gradio UI for:
- image captioning
- image Q&A
- short video Q&A
Module 15:
Building a Simple End-to-End Demo
- Upload an image or short video
- Run inference using a VLM
- Display outputs in a web UI
- Optional: Add a text box for user
questions
Module 16:
Responsible AI Essentials
- Image bias & fairness issues
- Video privacy considerations
- Avoiding misleading outputs
- Best practices for safe usage of
VLMs
Module 17:
Wrap-Up & Learning Path