Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Spark Streaming - Stream Processing in Lakehouse - PySpark
Highest Rated
Rating: 4.8 out of 5(2,221 ratings)
21,539 students

Spark Streaming - Stream Processing in Lakehouse - PySpark

Master Spark Structured Streaming using Python (PySpark) on Azure Databricks Cloud with a end-to-end Project
Last updated 8/2024
English
German [Auto],English [Auto],

What you'll learn

  • Real-time Stream Processing Concepts
  • Spark Structured Streaming APIs and Architecture
  • Working with Streaming Sources and Sinks
  • Kafka for Data Engineers
  • Working With Kafka Source and Integrating Spark with Kafka
  • State-less and State-full Streaming Transformations
  • Windowing Aggregates using Spark Stream
  • Watermarking and State Cleanup
  • Streaming Joins and Aggregation
  • Handling Memory Problems with Streaming Joins
  • Working with Azure Databricks
  • Capstone Project - Streaming application in Lakehouse

Course content

9 sections108 lectures22h 23m total length
  • About the Course5:57

    Master Spark Structured Streaming and Kafka integration in a Lakehouse on the Databricks platform, designing unified batch and streaming applications with testing and CI/CD practices.

  • Course Prerequisite2:27

    Meet the prerequisites: strong Python programming and Spark proficiency, including PySpark, Spark DataFrame API, and Spark SQL, to learn Spark Structured Streaming, Kafka basics, and Databricks Lakehouse concepts.

  • Source Code and Other Resources0:13
  • Note for Students - Before Start2:05

    Encourage learners to leave reviews and five-star ratings when benefiting from the course, to motivate updates, new content, and the latest technologies, with a 30-day refund policy.

Requirements

  • Spark Fundamentals and exposure to Spark Dataframe APIs
  • Programming Knowledge Using Python Programming Language

Description

About the Course

I am creating Apache Spark and Databricks - Stream Processing in Lakehouse using the Python Language and PySpark API. This course will help you understand Real-time Stream processing using Apache Spark and Databricks Cloud and apply that knowledge to build real-time stream processing solutions. This course is example-driven and follows a working session-like approach. We will take a live coding approach and explain all the needed concepts.

Capstone Project

This course also includes an End-To-End Capstone project. The project will help you understand the real-life project design, coding, implementation, testing, and CI/CD approach.

Who should take this Course?

I designed this course for software engineers willing to develop a Real-time Stream Processing Pipeline and application using Apache Spark. I am also creating this course for data architects and data engineers who are responsible for designing and building the organization’s data-centric infrastructure. Another group of people is the managers and architects who do not directly work with Spark implementation. Still, they work with those implementing Apache Spark at the ground level.

Spark Version used in the Course.

This Course is using the Apache Spark 3.5. I have tested all the source code and examples used in this Course on Azure Databricks Cloud using Databricks Runtime 14.1.


Who this course is for:

  • Software Engineers and Architects who are willing to design and develop a Bigdata Engineering Projects using Apache Spark and Databricks Cloud
  • Programmers and developers who are aspiring to grow and learn Data Engineering using Apache Spark and Databricks Cloud