Showing 31 total results
Learn to process, analyze, and derive insights from massive datasets using advanced tools and technologies.
The course covers essential concepts like lifecycle phases, repositories, and artifact management, ensuring efficient and reproducible builds.
It covers key concepts such as Kafka architecture, producers, consumers, topics, partitions, and brokers.
It covers essential topics such as data ingestion, transformation, routing, and delivery across diverse systems.
It covers essential topics like connecting to data sources, designing charts, and building custom dashboards for data exploration.
It covers essential concepts such as RDDs, DataFrames, and SparkSQL for efficient data manipulation and querying.
It covers core concepts such as stream processing, stateful computations, event time processing, and fault tolerance.
It covers core concepts such as Resilient Distributed Datasets (RDDs), DataFrames, and SparkSQL, providing the foundation for data manipulation and analysis.
It covers key components such as HDFS (Hadoop Distributed File System), YARN (Yet Another Resource Negotiator), and MapReduce for parallel data processing.
Participants will gain hands-on experience in leveraging Spark for real-world data analytics and machine learning applications.
It covers various technologies, including Hadoop, Apache Spark, and NoSQL databases.
This course provides system administrators with the skills necessary to set up, configure, manage, and troubleshoot Apache Spark clusters.
The training covers the concepts of hubs, links, and satellites to design a robust data architecture.
The course covers essential topics like setting up virtual hosts, managing server modules, and securing the server environment.
The course covers topics such as load balancing, reverse proxy setup, and configuring server blocks to handle multiple websites.
The training focuses on building proficiency in using Talend Studio for data integration, transformation, and managing data workflows.
Participants will explore foundational principles, key components, and best practices for creating scalable, efficient, and secure data architectures.
This course is suitable for data engineers, data scientists, and cloud professionals aiming to leverage Spark's powerful distributed computing capabilities in cloud environments.
Participants will learn the core Kafka architecture, installation and configuration, cluster administration, topic management, monitoring, and best practices for reliability and security.
Apache Superset basics and with deeper understanding of advanced configuration, custom visualizations, data security, performance optimization, and embedding Superset in enterprise applications.
This training provides a comprehensive understanding of Apache ZooKeeper, a centralized service for maintaining configuration information, naming, synchronization, and group services in distributed systems.
Apache Camel, a powerful open-source integration framework that enables seamless communication between different systems using Enterprise Integration Patterns (EIPs).
Participants will learn how to leverage data analytics to improve governance, enhance citizen services, and enable evidence-based decision-making.
This course provides administrators with in-depth knowledge of Apache Superset, an open-source data exploration and visualization platform.
Participants will learn how to set up, configure, manage, and optimize ActiveMQ for reliable asynchronous communication in distributed systems.
Participants will learn how to install, configure, and administer Kylin, build and optimize data models, and perform interactive analytics on large datasets using SQL on Hadoop.
It focuses on core administration skills, including deployment, configuration, monitoring, and troubleshooting.
The course focuses on writing and optimizing MapReduce programs, understanding Hadoop’s data flow, and integrating with other ecosystem tools such as Hive, Pig, and HBase
This training program is designed to help developers and testers leverage Apache Maven for managing and executing automated tests.
This training provides a practical and project-based introduction to Spark NLP, a production-grade natural language processing library built on Apache Spark.
This course provides hands-on training in developing and deploying modular OSGi applications using Apache Karaf.
This course introduces Apache Iceberg, an open table format for managing large analytic datasets on data lakes.