Course Name
Course Code : PRP06
Venue Details
Postal Code : 30-644
Session Dates
Duration: 3 days (21 hours)
GPU Programming with CUDA training provides participants with the knowledge to harness the computational power of GPUs for high-performance parallel computing. The course covers fundamental CUDA concepts, including threads, blocks, and grids, enabling participants to write efficient parallel programs. Attendees will learn to optimize memory usage, handle data transfers between host and device, and debug GPU applications effectively. Advanced topics include leveraging libraries like cuBLAS and cuDNN for accelerated mathematical operations and integrating CUDA with existing workflows. By the end of the training, participants will be skilled in developing and optimizing GPU-accelerated applications for a wide range of computational tasks.
Introduction to GPU Programming Introduction to GPU Computing Evolution of GPU computing. Comparison of CPU vs. GPU architectures. Applications of GPU programming in various domains. CUDA Fundamentals Overview of the CUDA architecture and toolkit. Setting up the CUDA environment. Writing and executing the first CUDA program. CUDA Programming Basics Thread hierarchy: grids, blocks, and threads. Kernel functions and their execution. Practical exercise Memory Model in CUDA Types of memory: global, shared, constant, and texture. Memory access patterns and optimization techniques. Practical exercise Intermediate CUDA Programming Thread Synchronization and Communication Synchronization techniques using shared memory. Atomic operations and their usage. Practical exercise Performance Optimization Techniques Optimizing memory access and minimizing latency. Efficient utilization of registers and shared memory. Profiling tools for performance analysis (Nsight, nvprof). Practical exercise CUDA Libraries and APIs Introduction to Thrust, cuBLAS, and cuDNN libraries. Using CUDA libraries for common computational tasks. Practical exercise Advanced Topics and Applications Advanced CUDA Concepts Dynamic parallelism and unified memory. Streams and asynchronous execution. Practical exercise Multi-GPU Programming Basics of multi-GPU programming. Using CUDA-aware MPI for distributed GPU computation. Practical exercise Real-World Applications Case studies: GPU-accelerated applications in AI, scientific computing, and image processing. Hands-on project Debugging and Error Handling Common debugging tools and techniques. Best practices for error handling in CUDA applications. Practical exercise Course Wrap-Up Summary of key concepts and tools. Open Q&A session and guidance on further learning. Feedback and certification distribution.
Mode of Delivery : The event can be attended both online and at nearby ProgNXT classroom by Individual Professionals and Corporate Employees as per the seat availability. Please Contact Us at [email protected] for checking the seat availability
Audience : We have a global audience that logs in to using their own computers to work hand in hand with our world-class instructors.
Assessment : Each training course will have ProgNXT Assessment at the end.
Certification : After successful passing of ProgNXT Assessment, ProgNXT Certification will be provided, which has got acceptance in 55+ Countries.
| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |