Get in Touch
 Duration 21 hours

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief introduction to Python and Scala

Core Concepts (Theory):

  • Architecture
  • RDD (Resilient Distributed Datasets)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring Fundamentals via Databricks (Hands-on Workshop):

  • Practical exercises using the RDD API
  • Basic action and transformation functions
  • PairRDD
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDF (User Defined Functions)
  • An overview of the DataSet API
  • Structured Streaming

Deployment Strategies via AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Key differences between AWS EMR and AWS Glue
  • Example job implementations in both environments
  • Analysis of advantages and limitations

Additional Topics:

  • Introduction to orchestration with Apache Airflow

Requirements

Programming skills (Python and Scala preferred)

Fundamental knowledge of SQL

Number of participants


Price per participant

Testimonials (3)

Provisional Upcoming Courses (Require 5+ participants)

Related Categories