- Course
Manage Data and Delta Tables with Apache Spark on Databricks
Learn to organize, read, write, and manage Delta tables in Databricks using Unity Catalog and Apache Spark. This course will teach you the hands-on data engineering skills needed to build reliable, production-ready pipelines on the Lakehouse.
- Course
Manage Data and Delta Tables with Apache Spark on Databricks
Learn to organize, read, write, and manage Delta tables in Databricks using Unity Catalog and Apache Spark. This course will teach you the hands-on data engineering skills needed to build reliable, production-ready pipelines on the Lakehouse.
Get started today
Access this course and other top-rated tech content with one of our business plans.
Try this course for free
Access this course and other top-rated tech content with one of our individual plans.
This course is included in the libraries shown below:
- Data
What you'll learn
Data engineers working on Databricks often struggle with the same set of problems: data that lands with no structure, schemas that break pipelines overnight, and tables that are painful to keep current without rebuilding from scratch. In this course, Manage Data and Delta Tables with Apache Spark on Databricks, you'll gain the ability to organize, manage, and operate Delta tables in a way that holds up in production. First, you'll explore how Unity Catalog structures data access using catalogs, schemas, and tables, and how to read and write Delta tables using PySpark and Spark SQL. Next, you'll discover how Delta Lake enforces schemas, how to handle schema evolution safely using mergeSchema, and how to identify the risks that come with changing schemas on live tables. Finally, you'll learn how to use MERGE for upserts, work with semi-structured data, and use time travel to audit and recover from unintended changes. When you're finished with this course, you'll have the skills and knowledge needed to build and maintain a reliable Delta table foundation that your team can depend on.