Setting up a scalable data pipeline does not require expensive cloud infrastructure. With the right local tools, you can build, test, and run powerful data engineering workflows directly on your machine. This text-based course guides you through setting up a modern local data stack to ingest, clean, and merge complex datasets.
You will start with foundational data engineering concepts, understanding how DuckDB acts as a high-performance analytical engine and how dbt manages SQL transformations. From there, you will learn how to integrate Python to resolve duplicate records and merge disparate data sources into a single, clean source of truth.
What you'll learn:
- Understand the core architecture of a modern local data stack using DuckDB and dbt
- Configure a local development environment with virtual environments and modern dependency management
- Write efficient SQL transformations and materializations within dbt
- Apply entity resolution techniques in Python to identify and merge duplicate records
- Establish robust data quality checks and schema testing to ensure pipeline reliability
- Build an end-to-end local data pipeline that transforms raw data into clean, analytical models
Through clear explanations and structured code walk-throughs, you will progress from basic database setup to advanced data transformation and entity matching. This course is designed for beginner data analysts, aspiring data engineers, and Python developers who want to master local data modeling without cloud overhead. No prior experience with DuckDB or dbt is required. Dive in to start building your local data pipeline today.
สิ่งที่คุณจะได้รับ
📜ใบประกาศนียบัตร เพิ่มในโปรไฟล์ LinkedIn ของคุณ
💬ติวเตอร์ AI ส่วนตัว ติดขัดในบทเรียน? ถามติวเตอร์ในตัวของคุณได้ทุกอย่าง ทุกเวลา