Course Name
Course Code : NOB56
Venue Details
Postal Code : 10330
Session Dates
Duration: 2 days (14 hours)
This foundation-level course introduces the principles and practices of Site Reliability Engineering (SRE). It emphasizes reliability, observability, service-level objectives, incident response, and the cultural alignment between development and operations.
Introduction to SRE Principles and Practices
What is SRE? History and evolution from Google
Key differences: SRE vs. DevOps vs. traditional operations
Principles of SRE: Error Budgets, Toil Reduction, Automation
Reliability Metrics: SLI, SLO, SLA explained
Service Level Objectives (SLO) Design and Measurement
Toil identification and automation strategies
Introduction to monitoring, logging, and alerting
Incident Response, Observability, and SRE Culture
Incident Management and Response (IMR)
Postmortems and Blameless Culture
Observability: logs, metrics, traces
Alert fatigue and alert tuning
Capacity Planning and Load Testing basics
Building SRE culture: collaboration, communication, and reliability ownership
Introduction to SRE tools: Prometheus, Grafana, Alertmanager, PagerDuty, etc.
Maturity models and SRE implementation roadmap
Hands-On Activities
Mode of Delivery : The event can be attended both online and at nearby ProgNXT classroom by Individual Professionals and Corporate Employees as per the seat availability. Please Contact Us at [email protected] for checking the seat availability
Audience : We have a global audience that logs in to using their own computers to work hand in hand with our world-class instructors.
Assessment : Each training course will have ProgNXT Assessment at the end.
Certification : After successful passing of ProgNXT Assessment, ProgNXT Certification will be provided, which has got acceptance in 55+ Countries.
| Global Region | Location | Start Date | End Date | Action |
|---|---|---|---|---|
| | | | | |
| | | | | |
| | | | | |
| | | | | |
| | | | | |