Modern VLMs: Qwen, Gemma & Kimi Training in Australia
Modern VLMs: Qwen, Gemma & Kimi Training in Australia
This program introduces participants to Vision-Language Models (VLMs) using fully open-source, modern models like Qwen 3 VL, Gemma 3 Vision, and Kimi-VL.
Modern VLMs: Qwen, Gemma & Kimi Training is a professional training program delivered by ProgNXT, a globally recognized corporate training provider. ProgNXT's Modern VLMs: Qwen, Gemma & Kimi Training course in Australia equips professionals with industry-relevant skills through hands-on, instructor-led sessions. This program introduces participants to Vision-Language Models (VLMs) using fully open-source, modern models like Qwen 3 VL, Gemma 3...
Expert Panel
Designed by the ProgNXT AI & Data Science Expert Panel, specializing in Generative AI, Machine Learning, and ChatGPT applications
ProgNXT AI & Data Science Expert PanelCourse Overview
Course Code: SZW49
21 Hrs
- Course Rating 5/5
Last Updated:
Overview
This program introduces participants to Vision-Language Models
(VLMs) using fully open-source, modern models like Qwen 3 VL, Gemma 3
Vision, and Kimi-VL.
Learners will understand how multimodal AI works, prepare simple image/video
datasets, run open-source VLMs, perform fine-tuning using LoRA,
and deploy multimodal applications.
The course focuses on intuitive explanations, hands-on guided exercises, and easy-to-use tools (Google Colab, Hugging Face Spaces, Gradio).
By the end, learners will be able to build and deploy their own simple VLM applications.
Welcome to the official Modern VLMs: Qwen, Gemma & Kimi Training certification program. This comprehensive training is designed to elevate your professional skills and provide you with practical, industry-relevant knowledge in in Australia. As a globally recognized corporate training provider operating in 55+ countries, ProgNXT ensures that our curriculum meets the highest standards of excellence.
Whether you are looking to upskill your team or advance your personal career, our expert-led sessions will guide you through the core concepts of this domain. Upon successful completion of the 21 Hrs program, participants will receive a globally accepted certification, demonstrating their proficiency and readiness to tackle complex challenges in the field.
Pre-Requisites
- Basic Python syntax (functions, lists, loops)
- Very basic understanding of machine learning concepts (optional but helpful)
- Ability to use Google Colab notebooks
What Skills It Will Add
Skillset Achieved After the Course
Vision-Language Understanding
- Understanding how AI interprets images, text, and video together
- Working with modern open-source VLMs
Data Handling
- Preparing image–text datasets
- Preparing small video–text datasets
- Basic image and frame preprocessing
Model Usage & Fine-Tuning
- Running Qwen 3 VL, Gemma 3, and Kimi-VL
- Performing simple LoRA fine-tuning on images
- Introductory video fine-tuning workflows
Evaluation & Debugging
- Identifying good vs bad model outputs
- Recognizing hallucinations and common model errors
Deployment
- Building a simple multimodal web app using Gradio
- Deploying on Hugging Face Spaces
- Running models on Google Colab (free)
Responsible AI Awareness
- Understanding bias, fairness, and privacy concerns in multimodal AI
Course Outcomes
Course Outcomes
By the end of the training, participants will be able to:
1. Understand key VLM concepts
- Explain what VLMs are
- Describe how vision encoders and language models work together
- Explain differences between Qwen 3 VL, Gemma 3, and Kimi-VL
2. Work with multimodal data
- Prepare image/text datasets
- Extract video frames and create small video caption datasets
3. Use and fine-tune open-source VLMs
- Load and experiment with VLMs on Colab
- Perform beginner-friendly LoRA-based fine-tuning
- Run inference on images and short videos
4. Deploy multimodal applications
- Build a simple UI that accepts images/videos and answers questions
- Deploy it publicly on Hugging Face Spaces
5. Understand responsible AI considerations
- Know where VLMs might hallucinate
- Recognize privacy and fairness concerns in vision/video AI
Modern VLMs: Qwen, Gemma & Kimi Training Events in Other Locations
Online Events| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |