Course Name
Course Code : PRQ99
Venue Details
Postal Code : 11001
Session Dates
Duration: 3 days (21 hours)
PySpark training introduces learners to the power of Apache Spark for processing large-scale data using Python. It covers core concepts such as Resilient Distributed Datasets (RDDs), DataFrames, and SparkSQL, providing the foundation for data manipulation and analysis. Participants gain hands-on experience in working with PySpark for both batch and real-time data processing, leveraging its scalability and speed. The training also explores advanced topics like machine learning with MLlib, graph processing with GraphX, and performance optimization techniques. By completing this course, professionals can apply PySpark to handle complex big data tasks and build efficient data pipelines for analytics.
Introduction to PySpark and Basic Operations Introduction to Big Data and Apache Spark: What is Big Data? Overview of Apache Spark and its ecosystem. Components of Spark: SparkCore, Spark SQL, Spark Streaming, MLlib, and GraphX. Benefits of using Spark for big data processing. Hands-on Exercise Understanding RDDs (Resilient Distributed Datasets): What is an RDD? Creating RDDs from existing data sources. Basic RDD operations: map(), filter(), reduce(), and collect(). Transformations and actions in RDDs. Hands-on Exercise Introduction to DataFrames and Spark SQL: What is a DataFrame? Differences between RDD and DataFrame. Loading data into Spark DataFrames. Performing basic operations on DataFrames (select, filter, groupBy). Hands-on Exercise Data Transformation and Analysis Data Processing with PySpark DataFrames: Using DataFrame APIs for data manipulation. Data cleaning and preprocessing using PySpark. Handling missing values, filtering, and data aggregation. Hands-on Exercise Advanced DataFrame Operations: Joins and aggregations with DataFrames. Working with nested data and complex data structures (arrays, maps). Window functions and ranking operations. Hands-on Exercise Working with PySpark SQL: Writing SQL queries in PySpark using the spark.sql() method. Understanding and working with temporary views and permanent tables. Using SQL functions in PySpark. Hands-on Exercise Spark Jobs, Performance Optimization, and Advanced Topics Running Spark Jobs: Spark job execution model: stages, tasks, and DAG (Directed Acyclic Graph). Running Spark jobs locally vs on a cluster. Submitting jobs to a Spark cluster using Spark-submit. Hands-on Exercise Performance Tuning and Optimization in PySpark: Understanding Spark job execution and performance bottlenecks. Optimizing Spark jobs using caching, partitioning, and parallelism. Tuning Spark configurations (e.g., memory, shuffle partitions). Hands-on Exercise Introduction to Spark Streaming: What is Spark Streaming? Processing real-time data streams with PySpark. Working with DStreams and window operations. Hands-on Exercise Integrating PySpark with Other Big Data Tools: Using PySpark with Hadoop and HDFS. Integration with Hive and other big data tools. Hands-on Exercise
Mode of Delivery : The event can be attended both online and at nearby ProgNXT classroom by Individual Professionals and Corporate Employees as per the seat availability. Please Contact Us at [email protected] for checking the seat availability
Audience : We have a global audience that logs in to using their own computers to work hand in hand with our world-class instructors.
Assessment : Each training course will have ProgNXT Assessment at the end.
Certification : After successful passing of ProgNXT Assessment, ProgNXT Certification will be provided, which has got acceptance in 55+ Countries.
| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |