Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Overview of Chinese AI GPU Ecosystem
- Comparison of Huawei Ascend, Biren, and Cambricon MLU architectures
- CUDA versus CANN, Biren SDK, and BANGPy models
- Industry trends and vendor ecosystems
Preparing for Migration
- Assessing your existing CUDA codebase
- Identifying target platforms and required SDK versions
- Toolchain installation and environment setup
Code Translation Techniques
- Porting CUDA memory access patterns and kernel logic
- Mapping compute grid/thread models
- Evaluating automated versus manual translation options
Platform-Specific Implementations
- Utilizing Huawei CANN operators and custom kernels
- Navigating the Biren SDK conversion pipeline
- Rebuilding models using BANGPy (Cambricon)
Cross-Platform Testing and Optimization
- Profiling execution on each target platform
- Memory tuning and comparing parallel execution methods
- Performance tracking and iterative refinement
Managing Mixed GPU Environments
- Hybrid deployments involving multiple architectures
- Fallback strategies and device detection mechanisms
- Implementing abstraction layers for code maintainability
Case Studies and Best Practices
- Porting vision and NLP models to Ascend or Cambricon platforms
- Retrofitting inference pipelines on Biren clusters
- Handling version mismatches and API gaps
Summary and Next Steps
Requirements
- Experience programming with CUDA or GPU-based applications
- Understanding of GPU memory models and compute kernels
- Familiarity with AI model deployment or acceleration workflows
Audience
- GPU programmers
- System architects
- Porting specialists
21 Hours