Get Free Shipping on orders over $79
Apache Spark 3 for Data Engineering and Analytics with Python - David Mngadi

Apache Spark 3 for Data Engineering and Analytics with Python

By: David Mngadi

eText | 30 August 2021 | Edition Number 1

At a Glance

eText


$112.19

or 4 interest-free payments of $28.05 with

 or 

Instant online reading in your Booktopia eTextbook Library *

Why choose an eTextbook?

Instant Access *

Purchase and read your book immediately

Read Aloud

Listen and follow along as Bookshelf reads to you

Study Tools

Built-in study tools like highlights and more

* eTextbooks are not downloadable to your eReader or an app and can be accessed via web browsers only. You must be connected to the internet and have no technical issues with your device or browser that could prevent the eTextbook from operating.

Master Python and PySpark 3.0.1 for Data Engineering / Analytics (Databricks)

Key Features

  • Apply PySpark and SQL concepts to analyze data
  • Understand the Databricks interface and use Spark on Databricks
  • Learn Spark transformations and actions using the RDD (Resilient Distributed Datasets) API

Book Description

Apache Spark 3 is an open-source distributed engine for querying and processing data. This course will provide you with a detailed understanding of PySpark and its stack. This course is carefully developed and designed to guide you through the process of data analytics using Python Spark. The author uses an interactive approach in explaining keys concepts of PySpark such as the Spark architecture, Spark execution, transformations and actions using the structured API, and much more. You will be able to leverage the power of Python, Java, and SQL and put it to use in the Spark ecosystem.

You will start by getting a firm understanding of the Apache Spark architecture and how to set up a Python environment for Spark. Followed by the techniques for collecting, cleaning, and visualizing data by creating dashboards in Databricks. You will learn how to use SQL to interact with DataFrames. The author provides an in-depth review of RDDs and contrasts them with DataFrames.

There are multiple problem challenges provided at intervals in the course so that you get a firm grasp of the concepts taught in the course.

The code bundle for this course is available here: https://github.com/PacktPublishing/Apache-Spark-3-for-Data-Engineering-and-Analytics-with-Python-

What you will learn

  • Learn Spark architecture, transformations, and actions using the structured API
  • Learn to set up your own local PySpark environment
  • Learn to interpret DAG (Directed Acyclic Graph) for Spark execution
  • Learn to interpret the Spark web UI
  • Learn the RDD (Resilient Distributed Datasets) API
  • Learn to visualize (graphs and dashboards) data on Databricks

Who this book is for

This course is designed for Python developers who wish to learn how to use the language for data engineering and analytics with PySpark. Any aspiring data engineering and analytics professionals. Data scientists/analysts who wish to learn an analytical processing strategy that can be deployed over a big data cluster. Data managers who want to gain a deeper understanding of managing data over a cluster.
on
Desktop
Tablet
Mobile

More in Programming & Scripting Languages

Grokking Statistics - Thomas Nield

eBOOK

Learn Calculus with Python - Nick McIntyre

eBOOK

RRP $61.44

$49.16

20%
OFF