AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Traditional observability relies on dashboards, threshold alerts, and manual log diving. AI-driven observability transforms this with natural language querying of telemetry data, LLM-powered root cause analysis, anomaly detection using foundation models, and automated incident summaries that understand context.
This instructor-led, live training (online or onsite) is aimed at observability and SRE engineers who want to integrate LLMs and AI into their monitoring, alerting, and incident analysis workflows.
By the end of this training, participants will be able to:
- Build natural language interfaces for querying Prometheus, Elasticsearch, and SQL-based observability stores.
- Implement LLM-powered log analysis and anomaly detection pipelines.
- Generate automated incident summaries and postmortem drafts from raw telemetry.
- Design AI-assisted root cause analysis workflows with evidence chaining.
- Integrate foundation models for time-series anomaly detection and forecasting.
- Deploy an AI-augmented on-call experience with smart alert enrichment.
Format of the Course
- Interactive lecture and discussion.
- Lots of exercises and practice.
- Hands-on implementation in a live-lab environment.
Course Customization Options
- To request a customized training, please contact us to arrange.
Course Outline
The AI Observability Landscape
- From dashboards to conversations: the shift toward AI-augmented observability
- LLM capabilities relevant to observability: summarization, reasoning, pattern matching
- Architecture patterns: embedding AI into existing observability stacks
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries
- NL querying for Elasticsearch, OpenSearch, and Loki log stores
- SQL generation from natural language for structured telemetry
- Building a query assistant agent with tool use and context awareness
LLM-Powered Log Analysis
- Automated log parsing and structuring with LLMs
- Anomaly detection in log streams using embedding similarity
- Log clustering and pattern discovery at scale
- Generating human-readable explanations from raw log sequences
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding
- Automated incident context gathering from runbooks, past incidents, and docs
- Smart alert routing based on content understanding and team expertise
- Reducing alert fatigue with AI-driven noise reduction
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation
- Evidence chaining: connecting symptoms across metrics, logs, and traces
- Guided troubleshooting with interactive AI diagnosis sessions
- Building a root cause analysis agent with progressive investigation
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry
- Automated postmortem drafting with timeline reconstruction
- Stakeholder communication tailored to technical and executive audiences
- Runbook suggestion and automated remediation recommendations
ML for Observability
- Time-series forecasting for capacity planning and anomaly prediction
- Foundation models for zero-shot anomaly detection on metrics
- Embedding-based service dependency mapping and topology discovery
- Training and deploying lightweight ML models alongside observability pipelines
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability
- Data privacy: ensuring LLMs do not leak sensitive telemetry
- Human oversight: when AI diagnosis needs operator validation
- Measuring impact: MTTD, MTTR, and on-call experience metrics
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic Python scripting for data processing.
Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers building next-generation monitoring pipelines.
- DevOps leads evaluating LLM integration into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Provisional Upcoming Courses (Require 5+ participants)
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity serves as an agentic development platform specifically engineered to create autonomous agents that leverage Gemini 3's multimodal capabilities for planning, reasoning, coding, and execution.
This live, instructor-led training, available online or on-site, is tailored for advanced technical professionals aiming to design, construct, and deploy autonomous agents utilizing Gemini 3 within the Antigravity environment.
By the end of this program, participants will be equipped to:
- Construct autonomous workflows that leverage Gemini 3 for logical reasoning, strategic planning, and task execution.
- Engineer agents within Antigravity capable of analyzing tasks, generating code, and interacting with various tools.
- Seamlessly integrate Gemini-driven agents into enterprise systems and API ecosystems.
- Enhance agent performance, safety, and reliability within complex operational environments.
Course Format
- Expert-led demonstrations paired with interactive peer discussions.
- Practical, hands-on experimentation in autonomous agent development.
- Real-world implementation utilizing Antigravity, Gemini 3, and complementary cloud services.
Customization Options
- For teams requiring specialized domain behaviors or bespoke integrations, please reach out to us to tailor the program to your specific needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework designed for exploring long-term agent lifecycles and the emergence of dynamic interactive behaviors.
Delivered as an instructor-led live session—available either online or on-site—this training is tailored for senior professionals seeking to engineer, assess, and refine agents that possess memory retention, feedback-driven improvement capabilities, and long-term evolutionary potential.
By the end of this course, participants will have acquired the proficiency to:
- Architect robust long-term memory structures that ensure agent persistence.
- Deploy efficient feedback mechanisms to actively steer agent conduct.
- Analyze learning paths and monitor model drift.
- Embed memory functions within intricate multi-agent frameworks.
Delivery Methodology
- Expert-facilitated dialogue complemented by technical walkthroughs.
- Practical application via curated design exercises.
- Implementation of theoretical concepts within simulated agent environments.
Bespoke Training Solutions
- Should your enterprise require specific content adjustments or industry-relevant case studies, please reach out to tailor this training to your needs.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra is a framework designed to facilitate deep integration between AI agents, APIs, enterprise applications, and external data systems.
This instructor-led live training (available online or onsite) targets intermediate-level engineers looking to build reliable, secure, and scalable integrations between Mastra agents and the wider enterprise ecosystem.
Upon completing this training, participants will be equipped to:
- Implement API-driven integrations connecting Mastra agents with external services.
- Link enterprise data systems and tools to automated agent workflows.
- Apply best practices for secure data exchange and authentication.
- Design integration layers that are scalable, maintainable, and ready for production use.
Course Format
- Interactive lectures and discussions.
- Hands-on engineering exercises involving integrations and APIs.
- Live laboratory implementation using real-world enterprise scenarios.
Course Customization Options
- Custom API scenarios, enterprise system mappings, or data-integration workshops can be arranged upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore empowers AI agents to create dynamic, interactive, and context-aware experiences by offering memory persistence, a secure code interpreter, and a browser tool.
This live, instructor-led training (available online or onsite) is designed for technical practitioners at intermediate to advanced levels who aim to build and deploy AI agents that retain long-term context, perform on-the-fly calculations, and interact directly with web interfaces.
Upon completion, participants will be equipped to:
- Implement AgentCore memory for workflows that are stateful and context-aware.
- Utilize the secure code interpreter for performing dynamic calculations and data transformations.
- Integrate the browser tool to facilitate real-time data retrieval and UI interactions.
- Design interactive agents for applications in analytics, customer support, and research.
Course Format
- Engaging lectures and open discussions.
- Practical lab exercises utilizing AgentCore memory and tools.
- Case studies focusing on analytics, automation, and customer support scenarios.
Course Customization
- Contact us to arrange a customized training session for this course.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursAgentCore Runtime & Gateway serves as a paired AWS service designed to streamline the packaging, deployment, and secure exposure of AI agents, facilitating seamless integration with external systems.
This instructor-led training session, available online or onsite, targets intermediate-level engineering teams aiming to transition their agent prototypes into production. The curriculum focuses on mastering the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration.
Upon completion of this course, participants will be capable of:
- Establishing AgentCore Runtime environments and packaging agents for deployment.
- Exposing agents via the Gateway using authenticated endpoints with rate limiting.
- Integrating external tools and APIs into agent workflows through stable contracts.
- Implementing observability, logging, and usage monitoring essential for production operations.
Course Format
- Interactive lectures and discussions.
- Practical labs featuring Runtime deployments and Gateway integrations.
- Targeted exercises emphasizing reliability, security, and deployment strategies.
Customization Options
- To arrange customized training for this course, please contact us.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity serves as a specialized development platform engineered for the creation of AI-powered, agent-first applications.
This instructor-led training, available both online and onsite, is tailored for intermediate developers aiming to build practical applications utilizing autonomous AI agents within the Antigravity ecosystem.
Upon completion, participants will possess the ability to:
- Engineer applications that leverage autonomous and coordinated AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for comprehensive development cycles.
- Orchestrate multi-agent workflows utilizing the Agent Manager.
- Embed agent capabilities into robust, production-grade software systems.
Training Structure
- Integrated presentations paired with detailed technical demonstrations.
- Substantial hands-on exercises and guided practical sessions.
- Real-world implementation tasks performed directly within the Antigravity live environment.
Customization Opportunities
- To align content specifically with your development stack, please reach out to us to discuss a tailored version of this training.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity serves as an agent-first development environment, engineered to optimize engineering workflows by leveraging intelligent automation.
This live, instructor-led training session, available both online and onsite, is tailored for entry-level practitioners eager to explore the core principles of Antigravity and comprehend how agent-driven coding environments can boost productivity.
By the end of this training, participants will have the capability to:
- Install and properly configure Google Antigravity.
- Navigate and grasp the functionalities of both the Editor View and Manager View.
- Collaborate effectively with agents to automate routine development tasks.
- Utilize Antigravity for generating, refining, and managing project files.
Course Format
- Instructor-led explanations complemented by real-time live demonstrations.
- Structured exercises emphasizing the practical application of agents.
- Hands-on exploration of essential Antigravity features within a controlled lab setting.
Customization Options
- Should you need a customized version of this training, please reach out to arrange a tailored program.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity serves as a robust platform for developing agents that interact seamlessly with web applications, browser environments, and complex multi-surface workflows.
Designed for intermediate-level professionals, this instructor-led live training (available online or onsite) focuses on building, automating, and testing browser-based workflows using Google Antigravity.
Upon completion, participants will be equipped to:
- Develop agents that effectively interact with web applications within browser surfaces.
- Streamline end-to-end workflows across various browser contexts.
- Validate and resolve agent behavior issues in UI-driven environments.
- Deploy cross-surface automation strategies leveraging Antigravity.
Course Format
- Structured guidance backed by practical demonstrations.
- Hands-on activities and scenario-based exercises for real-world application.
- Implementation of agent workflows within an interactive lab setting.
Customization Options
- Contact us to tailor the curriculum to your specific training objectives and requirements.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the development, optimization, and monitoring of fully managed AI agents by offering a comprehensive suite of services designed for large-scale deployment.
This instructor-led live training, available online or on-site, is designed for practitioners ranging from beginners to intermediates who seek practical experience in building production-ready AI agents using AgentCore.
Upon completing this training, participants will be equipped to:
- Grasp the fundamental capabilities of AgentCore in AI agent development.
- Architect and set up simple AI agents utilizing managed services.
- Incorporate workflows to augment agent functionality.
- Release and oversee AI agents within production environments.
Course Format
- Engaging lectures and interactive discussions.
- Practical labs focused on AgentCore services.
- Structured exercises guiding the journey from concept to deployment.
Customization Options
- For tailored training requests regarding this course, please reach out to us to coordinate the arrangement.
AI Agent Development with Mastra
14 HoursThis live, instructor-led training, offered online or onsite, is designed for intermediate-level developers and engineering teams looking to build scalable, observable AI systems using Mastra.
By the end of the session, learners will be able to:
- Understand Mastra's architecture and how it connects with LLMs and external APIs.
- Design and build AI agents and workflows in TypeScript.
- Utilize Mastra's observability and memory tools to track and optimize agent performance.
- Deploy production-ready AI applications using Mastra's framework features.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra is a framework that offers structured tools to evaluate, debug, and ensure the reliability of AI agents operating within complex workflows.
This instructor-led, live training (available online or onsite) targets intermediate-level practitioners who want to rigorously test agent behavior, enhance reliability, and implement measurable evaluation processes.
By the end of this training, participants will be able to confidently:
- Apply debugging techniques to identify and resolve agent behavior issues.
- Evaluate agents using structured metrics, benchmarks, and quality scores.
- Implement tooling and workflows that track reliability, drift, and hallucinations.
- Design QA strategies that ensure consistent and predictable agent performance.
Course Format
- Interactive lectures and discussions.
- Hands-on debugging and evaluation exercises.
- Live-lab analysis of agent behaviors using observability tools.
Course Customization Options
- Customized reliability testing scenarios and industry-specific QA methods can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra is an operational framework designed to streamline the deployment, scaling, and lifecycle management of AI agents in production environments.
This instructor-led, live training (online or onsite) is aimed at intermediate-level to advanced-level technical professionals who need to operationalize AI agents reliably and efficiently across production systems.
Upon completion of this training, attendees will be equipped to:
- Deploy Mastra-based AI agents into controlled, production-grade environments.
- Scale agents horizontally and vertically using platform-native primitives.
- Implement observability pipelines to track agent behaviour and performance.
- Optimize runtime configurations to reduce latency, costs, and operational risks.
Format of the Course
- Interactive lecture and discussion.
- Hands-on exercises focused on real deployment scenarios.
- Live-lab implementation using containerized and orchestrated environments.
Course Customization Options
- Customization of topics, hands-on labs, or industry-specific scenarios is available upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra is a framework that empowers sophisticated workflow automation and coordination across multiple AI agents operating within distributed systems. <\/p>
This instructor-led training session (available online or onsite) targets intermediate-level practitioners who aim to design, orchestrate, and manage multi-agent workflows at scale. <\/p>
Upon completion of this training, participants will acquire the ability to: <\/p>
- Design complex workflows utilizing Mastra’s orchestration capabilities. <\/li>
- Coordinate multiple agents executing parallel or dependent tasks. <\/li>
- Deploy monitoring and debugging tools for efficient workflow execution. <\/li>
-
Optimize orchestration logic to enhance reliability, throughput, and automation efficiency.
<\/li>
<\/ul>
Course Format <\/p>
- Interactive lectures and discussions. <\/li>
- Hands-on exercises focused on workflow design and automation. <\/li>
-
Practical implementation within a containerized live-lab environment.
<\/li>
<\/ul>
Customization Options <\/p>
- Custom automation scenarios, enterprise integrations, or workflow patterns can be tailored upon request. <\/li> <\/ul>
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity serves as an agent-centric platform designed to coordinate, supervise, and orchestrate automation and coding workflows powered by AI.
This live, instructor-led session, available either online or on-site, is tailored for intermediate professionals seeking to manage, design, and refine multi-agent workflows within the Google Antigravity environment.
By the end of this program, participants will be equipped to:
- Set up agent responsibilities and orchestration pipelines using the Manager interface.
- Create and analyze Antigravity artifacts, such as logs, browser recordings, plans, and task lists.
- Apply verification methods that keep agent actions clear and auditable.
- Enhance multi-agent collaboration for demanding operational and development tasks.
Training Format
- Guided demonstrations and practical walkthroughs.
- Scenario-driven exercises addressing real-world workflow challenges.
- Interactive experimentation in a live Antigravity workspace.
Customization
- For a customized version of this course, please reach out to discuss available options.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity is a framework designed to support advanced agent-driven development workflows.
This live, instructor-led training, available online or onsite, is tailored for intermediate to advanced professionals aiming to verify, validate, and secure the output generated by AI agents within Antigravity-driven environments.
By the end of this program, participants will be equipped to:
- Evaluate the accuracy and safety of code artifacts created by agents.
- Employ structured methods to verify tasks executed by agents.
- Analyze browser recordings and trace agent activity with precision.
- Implement QA and security principles to guarantee the reliability of agent workflows.
Course Format
- Instructor-led technical briefings and interactive discussions.
- Practical exercises centered on verifying real-world agent workflows.
- Hands-on testing and validation conducted in a controlled lab environment.
Customization Options
- Scenarios, workflows, and testing examples can be adapted to specific needs upon request.