Log in to access more pages.
Create an account or log in to continue reading more pages.
Log in
Table of contents
Data Engineering on AWS
From fundamentals to production-grade pipelines, analytics platforms, PySpark, SQL, and observability with Datadog
Read each section in order. Every title can be opened as a TheoryTrace document.
- Cover1
- Copyright2
- How to read this book3
- Introduction4
- Chapter 1: The Data Engineering Mindset5
- Chapter 2: Data Fundamentals: Files, Tables, Events, and Schemas6
- Chapter 3: SQL for Data Engineering7
- Chapter 4: Python for Data Pipelines8
- Chapter 5: Linux, Networking, and Cloud Foundations9
- Chapter 6: AWS Core Services for Data Engineers10
- Chapter 7: Amazon S3 as a Data Lake Foundation11
- Chapter 8: Data Modeling for Analytics12
- Chapter 9: Batch Ingestion Patterns13
- Chapter 10: Streaming and Event-Driven Data Engineering14
- Chapter 11: Apache Spark and PySpark Fundamentals15
- Chapter 12: Production PySpark Transformations16
- Chapter 13: AWS Glue and EMR for Distributed Processing17
- Chapter 14: Querying the Lake with Athena and Redshift18
- Chapter 15: Orchestration and Workflow Design19
- Chapter 16: Data Quality, Testing, and Reliability20
- Chapter 17: Observability for Data Systems with Datadog21
- Chapter 18: Security, Governance, and Compliance22
- Chapter 19: CI/CD and Infrastructure as Code for Data Engineering23
- Chapter 20: Cost, Performance, and Scalability24
- Chapter 21: Building an End-to-End AWS Data Platform25
- Chapter 22: Career-Ready Data Engineering Practice26
- Conclusion27