Course Name
Course Code : PRQ96
Venue Details
Postal Code : 28014
Session Dates
Duration: 3 days (21 hours)
Apache Spark training provides a deep understanding of big data processing and analytics using the Spark framework. It covers essential concepts such as RDDs, DataFrames, and SparkSQL for efficient data manipulation and querying. Participants gain hands-on experience with Spark's machine learning library (MLlib) and real-time stream processing capabilities through Spark Streaming. The training emphasizes optimizing Spark jobs, managing clusters, and integrating Spark with other big data tools like Hadoop. By completing the course, professionals can harness the full power of Spark to process large-scale data and perform advanced analytics in distributed environments.
Introduction to Apache Spark and Core Concepts Introduction to Apache Spark: What is Apache Spark and its advantages over Hadoop. Spark components: Spark Core, Spark SQL, Spark Streaming, MLlib, Understanding Spark’s architecture: Driver, Executors, Cluster Manager. Hands-on Exercise Spark RDDs (Resilient Distributed Datasets): Understanding RDDs: Creation, transformations, and actions. Working with RDDs in Spark (e.g., map(), filter(), reduce(), flatMap()). Hands-on Exercise Spark DataFrames and Datasets: Introduction to DataFrames and Datasets in Spark. Benefits of DataFrames over RDDs. Creating DataFrames from RDDs, JSON, and CSV files. Hands-on Exercise Spark SQL and Data Processing Spark SQL: Introduction to Spark SQL and its integration with DataFrames. Running SQL queries on Spark DataFrames. Creating and querying temporary views in Spark SQL. Hands-on Exercise DataFrame Transformations and Actions: Using various DataFrame transformations: select(), groupBy(), join(), and agg(). Common actions on DataFrames: collect(), show(), count(), save(). Hands-on Exercise Data Sources and Formats in Spark: Reading from and writing to various data formats: JSON, CSV, Parquet, JDBC, HDFS. Connecting to different data sources (e.g., relational databases, HDFS). Hands-on Exercise Spark Streaming and Optimization Techniques Spark Streaming: Introduction to real-time stream processing with Spark Streaming. Working with DStreams and Spark Streaming transformations. Handling time windows and sliding windows in streaming data. Hands-on Exercise Spark Machine Learning (MLlib): Introduction to MLlib for building machine learning models in Spark. Basic machine learning algorithms: classification, regression, clustering. Hands-on Exercise Spark Optimization and Performance Tuning: Understanding Spark’s execution plan and how to optimize jobs. Techniques for performance tuning: partitioning, caching, broadcasting, and serialization. Hands-on Exercise Final Project and Review: Using all the concepts learned to solve a real-world big data problem. Reviewing key topics and answering participant questions. Hands-on Exercise
Mode of Delivery : The event can be attended both online and at nearby ProgNXT classroom by Individual Professionals and Corporate Employees as per the seat availability. Please Contact Us at [email protected] for checking the seat availability
Audience : We have a global audience that logs in to using their own computers to work hand in hand with our world-class instructors.
Assessment : Each training course will have ProgNXT Assessment at the end.
Certification : After successful passing of ProgNXT Assessment, ProgNXT Certification will be provided, which has got acceptance in 55+ Countries.
| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |