Course Name
Course Code : UJN77
Venue Details
Postal Code : Z05H9K3
Session Dates
Duration: 2 days (14 hours)
This course provides a practical and theoretical understanding of how to fine-tune Vision-Language Models (VLMs) such as CLIP, BLIP, Flamingo, LLaVA, and Kosmos for domain-specific tasks. Participants will learn transfer learning, prompt-tuning, parameter-efficient fine-tuning, and evaluation techniques for multimodal tasks like image captioning, visual question answering (VQA), multimodal retrieval, and visual grounding.
Foundations of VLMs and Fine-Tuning
Introduction to Vision-Language Models
Multimodal pretraining paradigms
Fine-tuning strategies: Full vs. Parameter-efficient methods
Dataset preparation: Image-text pairs, preprocessing pipelines
Fine-Tuning Techniques
Full model fine-tuning (when and why)
LoRA and Adapter tuning for efficiency
Prompt-tuning and instruction tuning for VLMs
Case studies: Fine-tuning for Image Captioning & VQA
Evaluation and Deployment
Metrics for multimodal tasks
Error analysis and model debugging
Deployment approaches (TorchServe, ONNX, Hugging Face Spaces, APIs)
Responsible AI: Bias, fairness, and ethical concerns in multimodal AI
Hands-on Exercises
Summary and Conclusion
Mode of Delivery : The event can be attended both online and at nearby ProgNXT classroom by Individual Professionals and Corporate Employees as per the seat availability. Please Contact Us at [email protected] for checking the seat availability
Audience : We have a global audience that logs in to using their own computers to work hand in hand with our world-class instructors.
Assessment : Each training course will have ProgNXT Assessment at the end.
Certification : After successful passing of ProgNXT Assessment, ProgNXT Certification will be provided, which has got acceptance in 55+ Countries.
| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |