Koulutus
Overview
In this intermediate course, you will learn to design, build, and optimize robust batch data pipelines on Google Cloud. Moving beyond fundamental data handling, you will explore large-scale data transformations and efficient workflow orchestration, essential for timely business intelligence and critical reporting.
Get hands-on practice using Dataflow for Apache Beam and Serverless for Apache Spark (Dataproc Serverless) for implementation, and tackle crucial considerations for data quality, monitoring, and alerting to ensure pipeline
reliability and operational excellence. A basic knowledge of data warehousing, ETL/ELT, SQL, Python, and Google Cloud concepts is recommended.
Prerequisites
Participants should have:
- Basic proficiency with Data Warehousing and ETL/ELT concepts
- Basic proficiency in SQL
- Basic programming knowledge (Python recommended)
- Familiarity with gcloud CLI and the Google Cloud console
- Familiarity with core Google Cloud concepts and services
Target audience
This course is designed for:
- Data Engineers
- Data Analysts
Objectives
By the end of this course, learners will be able to:
- Determine whether batch data pipelines are the correct choice for your business use case.
- Design and build scalable batch data pipelines for high-volume ingestion and transformation.
- Implement data quality controls within batch pipelines to ensure data integrity.
- Orchestrate, manage, and monitor batch data pipeline workflows, implementing error handling and observability using logging and monitoring tools.
Outline
Module 1 When to choose batch data pipelines
Topics
- You will learn the critical role of a data engineer in developing and maintaining batch data pipelines, understand their core components and lifecycle, and analyze common challenges in batch data processing. You'll also identify key Google Cloud services that address these challenges.
Objectives
- Explain the critical role of a data engineer in developing and maintaining batch data pipelines.
- Describe the core components and typical lifecycle of batch data pipelines from ingestion to downstream consumption.
- Analyze common challenges in batch data processing, such as data volume, quality, complexity, and reliability, and identify key Google Cloud services that can address them.
Module 2 Design and build batch data pipelines
Topics
- You will design scalable batch data pipelines for high-volume data ingestion and transformation. You'll also optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques.
Objectives
- Design scalable batch data pipelines for high-volume data ingestion and transformation.
- Optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques.
Module 3 Control data quality in batch data pipelines
Topics
- You will develop data validation rules and cleansing logic to ensure data quality within batch pipelines. You'll also implement strategies for managing schema evolution and performing data deduplication in large datasets.
Objectives
- Develop data validation rules and cleansing logic to ensure data quality within batch pipelines.
- Implement strategies for managing schema evolution and performing data deduplication in large datasets.
Module 4 Orchestrate and monitor batch data pipelines
Topics
- You will orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking. You'll also implement robust error handling, monitoring, and observability for batch data pipelines.
Objectives
- Orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking
- Implement robust error handling, monitoring, and observability for batch data pipelines
Exams and assessments
There is no specific certification related to this course.
Hands-on learning
There are four practical labs in this course.
Osta liput
QA’s online-courses from Tieturi
Questions about QA courses?
Find out how QA’s live online courses work, what you need to participate, and what to expect before booking your training.
Accreditation and trademark notice
ITIL® and PRINCE2® courses are provided by QA Ltd, an ATO of People Cert.
ITIL®, PRINCE2® are registered trademarks of the PeopleCert group. Used under licence from PeopleCert. All rights reserved.
TOGAF® is a registered trademark of The Open Group.