Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open-Source Solutions <\/p>
- Understanding AIOps concepts and their advantages <\/li>
- The role of Prometheus and Grafana in the observability stack <\/li>
-
The place of ML in AIOps: predictive versus reactive analytics
<\/li>
<\/ul>
Configuring Prometheus and Grafana <\/p>
- Setting up and configuring Prometheus for time series data collection <\/li>
- Designing Grafana dashboards with real-time metrics <\/li>
-
Investigating exporters, relabeling, and service discovery mechanisms
<\/li>
<\/ul>
Data Preparation for Machine Learning <\/p>
- Extracting and processing Prometheus metrics <\/li>
- Structuring datasets for anomaly detection and forecasting tasks <\/li>
-
Leveraging Grafana’s transformation features or Python pipelines
<\/li>
<\/ul>
Utilizing Machine Learning for Anomaly Detection <\/p>
- Applying fundamental ML models for outlier identification (e.g., Isolation Forest, One-Class SVM) <\/li>
- Training and assessing models against time series data <\/li>
-
Displaying detected anomalies within Grafana dashboards
<\/li>
<\/ul>
Metric Forecasting via ML <\/p>
- Developing basic forecasting models (ARIMA, Prophet, introductory LSTM) <\/li>
- Anticipating system load or resource consumption <\/li>
-
Utilizing predictions for proactive alerting and scaling actions
<\/li>
<\/ul>
Merging ML with Alerting and Automation <\/p>
- Creating alert rules based on ML outputs or defined thresholds <\/li>
- Implementing Alertmanager and notification routing strategies <\/li>
-
Initiating scripts or automation workflows upon anomaly detection
<\/li>
<\/ul>
Scaling and Implementing AIOps Operations <\/p>
- Connecting external observability tools (e.g., ELK stack, Moogsoft, Dynatrace) <\/li>
- Integrating ML models into observability workflows <\/li>
-
Best practices for large-scale AIOps deployment
<\/li>
<\/ul>
Recap and Future Directions <\/p>
Requirements
- A solid grasp of system monitoring and observability principles <\/li>
- Practical experience with Grafana or Prometheus <\/li>
-
Knowledge of Python and fundamental machine learning concepts
<\/li>
<\/ul>
Target Audience <\/p>
- Observability engineers <\/li>
- Infrastructure and DevOps teams <\/li>
- Monitoring platform architects and site reliability engineers (SREs) <\/li> <\/ul>