Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Concepts of Agentic Systems in Live Environments
- Architectural components: loops, tooling, memory structures, and orchestration layers
- Agent lifecycle: phases of development, deployment, and continuous operation
- Key challenges in managing agents at production scale
Infrastructure and Deployment Frameworks
- Implementing agents within containerized and cloud-native settings
- Scaling strategies: horizontal vs. vertical expansion, concurrency management, and throttling
- Coordinating multi-agent interactions and workload distribution
Monitoring and Observability Practices
- Essential metrics: latency, success rates, memory consumption, and agent invocation depth
- Tracking agent activities and visualizing call graphs
- Enhancing observability with tools like Prometheus, OpenTelemetry, and Grafana
Logging, Auditing, and Compliance Management
- Aggregating logs and collecting structured event data
- Ensuring compliance and auditability within agentic workflows
- Creating audit trails and replay capabilities for effective debugging
Performance Optimization and Resource Efficiency
- Minimizing inference overhead and refining agent orchestration cycles
- Utilizing model caching and lightweight embeddings to accelerate retrieval
- Conducting load and stress testing for AI pipelines
Cost Governance and Strategic Control
- Analyzing cost drivers: API usage, memory, compute resources, and third-party integrations
- Monitoring individual agent costs and applying chargeback models
- Enforcing automation policies to curb agent sprawl and eliminate idle resource usage
CI/CD and Agent Rollout Methodologies
- Embedding agent workflows into CI/CD pipelines
- Strategies for testing, versioning, and rolling back iterative agent updates
- Implementing progressive rollouts and secure deployment mechanisms
Resilience Engineering and Failure Recovery
- Building for fault tolerance and managing graceful degradation
- Applying retry, timeout, and circuit breaker patterns to ensure agent stability
- Establishing incident response and post-mortem processes for AI operations
Capstone Project
- Constructing and launching an agentic AI system with comprehensive monitoring and cost tracking
- Simulating load conditions, assessing performance, and refining resource consumption
- Presenting the final architecture and monitoring dashboards to colleagues
Conclusion and Future Directions
Requirements
- A solid grasp of MLOps and production-grade machine learning systems
- Hands-on experience with containerized environments (Docker/Kubernetes)
- Knowledge of cloud cost management and observability platforms
Intended Audience
- MLOps Engineers
- Site Reliability Engineers (SREs)
- Engineering Managers responsible for AI infrastructure
21 Hours
Testimonials (3)
The trainer is patient and very helpful. He knows the topic well.
CLIFFORD TABARES - Universal Leaf Philippines, Inc.
Course - Agentic AI for Business Automation: Use Cases & Integration
Good mixvof knowledge and practice
Ion Mironescu - Facultatea S.A.I.A.P.M.
Course - Agentic AI for Enterprise Applications
The mix of theory and practice and of high level and low level perspectives