Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- Evolution of IT automation: moving from static runbooks to reasoning agents
- Anatomy of an agent: reasoning loops, tool utilization, memory management, and planning
- Determining when to automate versus retaining human involvement in the loop
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm patterns
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent implementations
- Building your first operational agent: querying monitoring systems, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Infrastructure tool usage: executing kubectl, Terraform, and Ansible commands via agent actions
- Designing safe tool interfaces using parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and intelligent routing
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation: executing restart, scaling, rollback, and failover actions
- Developing an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Action classification: categorizing read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical operations
- Guardrail patterns: implementing action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance adherence
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Evaluation
- Tracing agent reasoning chains to facilitate debugging and audits
- Evaluating decision quality: assessing precision, recall, and time-to-resolution
- Implementing feedback loops: learning from operator overrides and operational outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services via APIs, webhooks, and scheduled jobs
- Gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
- Runbooks for agent failures: protocols for when the agent itself encounters issues
- Constructing the business case and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST API integration.
- A foundational understanding of LLM capabilities and prompt engineering techniques.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leads evaluating agentic AI solutions for incident management.
14 Hours