Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Fundamentals of Performance and Key Metrics
- Analyzing latency, throughput, power consumption, and resource utilization
- Distinguishing between system-level and model-level bottlenecks
- Approaches to profiling for inference versus training tasks
Profiling on Huawei Ascend
- Leveraging CANN Profiler and MindInsight
- Conducting diagnostics at the kernel and operator levels
- Analyzing offload patterns and memory mapping strategies
Profiling on Biren GPUs
- Utilizing performance monitoring features within the Biren SDK
- Managing kernel fusion, memory alignment, and execution queues
- Implementing power and temperature-aware profiling techniques
Profiling on Cambricon MLU
- Employing BANGPy and Neuware performance utilities
- Gaining visibility into kernel operations and interpreting logs
- Integrating the MLU profiler with deployment frameworks
Optimization at the Graph and Model Levels
- Strategies for graph pruning and quantization
- Techniques for operator fusion and computational graph restructuring
- Standardizing input sizes and fine-tuning batch configurations
Memory and Kernel Optimization Techniques
- Enhancing memory layout efficiency and reuse strategies
- Managing buffers effectively across different chipsets
- Applying platform-specific tuning techniques for kernels
Cross-Platform Best Practices
- Achieving performance portability through abstraction strategies
- Developing shared tuning pipelines suitable for multi-chip environments
- Case Study: Optimizing an object detection model across Ascend, Biren, and MLU architectures
Summary and Future Steps
Requirements
- Prior experience in AI model training or deployment pipelines
- Solid understanding of GPU/MLU compute principles and model optimization techniques
- Familiarity with basic performance profiling tools and metrics
Target Audience
- Performance engineers
- Teams responsible for machine learning infrastructure
- AI system architects
21 Hours