Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Introduction:
- Apache Spark within the Hadoop Ecosystem
- Brief introduction to Python and Scala
Core Concepts (Theory):
- Architecture
- RDD (Resilient Distributed Datasets)
- Transformations and Actions
- Stages, Tasks, and Dependencies
Exploring Fundamentals via Databricks (Hands-on Workshop):
- Practical exercises using the RDD API
- Basic action and transformation functions
- PairRDD
- Join operations
- Caching strategies
- Practical exercises using the DataFrame API
- SparkSQL
- DataFrame operations: select, filter, group, and sort
- UDF (User Defined Functions)
- An overview of the DataSet API
- Structured Streaming
Deployment Strategies via AWS (Hands-on Workshop):
- Fundamentals of AWS Glue
- Key differences between AWS EMR and AWS Glue
- Example job implementations in both environments
- Analysis of advantages and limitations
Additional Topics:
- Introduction to orchestration with Apache Airflow
Requirements
Programming skills (Python and Scala preferred)
Fundamental knowledge of SQL
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift