Get Free Shipping on orders over $79
Real-Time Stream Processing Using Apache Spark 3 for Python Developers - ScholarNest

Real-Time Stream Processing Using Apache Spark 3 for Python Developers

By: ScholarNest

eText | 25 February 2022 | Edition Number 1

At a Glance

eText


$251.89

or 4 interest-free payments of $62.97 with

 or 

Instant online reading in your Booktopia eTextbook Library *

Why choose an eTextbook?

Instant Access *

Purchase and read your book immediately

Read Aloud

Listen and follow along as Bookshelf reads to you

Study Tools

Built-in study tools like highlights and more

* eTextbooks are not downloadable to your eReader or an app and can be accessed via web browsers only. You must be connected to the internet and have no technical issues with your device or browser that could prevent the eTextbook from operating.

Build your own real-time stream processing applications using Apache Spark 3.x and PySpark

Key Features

  • Learn real-time stream processing concepts
  • Understand Spark structured streaming APIs and architecture
  • Work with file streams, Kafka source, and integrating Spark with Kafka

Book Description

Take your first steps towards discovering, learning, and using Apache Spark 3.0. We will be taking a live coding approach in this carefully structured course and explaining all the core concepts needed along the way.

In this course, we will understand the real-time stream processing concepts, Spark structured streaming APIs, and architecture.

We will work with file streams, Kafka source, and integrating Spark with Kafka. Next, we will learn about state-less and state-full streaming transformations. Then cover windowing aggregates using Spark stream. Next, we will cover watermarking and state cleanup. After that, we will cover streaming joins and aggregation, handling memory problems with streaming joins. Finally, learn to create arbitrary streaming sinks.

By the end of this course, you will be able to create real-time stream processing applications using Apache Spark.

All the resources for the course are available at https://github.com/PacktPublishing/Real-time-stream-processing-using-Apache-Spark-3-for-Python-developers

What you will learn

  • Explore state-less and state-full streaming transformations
  • Windowing aggregates using Spark stream
  • Learn Watermarking and state cleanup
  • Implement streaming joins and aggregations
  • Handling memory problems with streaming joins
  • Learn to create arbitrary streaming sinks

Who this book is for

This course is designed for software engineers and architects who are willing to design and develop big data engineering projects using Apache Spark. It is also designed for programmers and developers who are aspiring to grow and learn data engineering using Apache Spark.

For this course, you need to know Spark fundamentals and should be exposed to Spark Dataframe APIs. Also, you should know Kafka fundamentals and have a working knowledge of Apache Kafka. One should also have programming knowledge of Python programming.
on
Desktop
Tablet
Mobile

More in Data Capture & Analysis

China's Megatrends : The 8 Pillars of a New Society - John Naisbitt

eBOOK

AI Model Evaluation - Leemay Nassery

eBOOK

Learn AI Data Engineering - David Melillo

eBOOK

Nominalization : Exploring Aspect and Countability - University of Pardubice

eBOOK