Evaluate Storage for Data Warehousing Success

External: Coursera Courses ↗ · Coursera

Open Course on External: Coursera

Free to audit · Opens on External: Coursera

Evaluate Storage for Data Warehousing Success

Coursera · Advanced ·🔄 Data Engineering ·3mo ago

Key Takeaways

Evaluates storage for data warehousing success using columnar versus row-oriented storage formats

Original Description

Master the critical decision-making skills for optimizing data warehouse storage architecture. This course equips data professionals with the analytical expertise to evaluate columnar versus row-oriented storage formats based on workload characteristics, query patterns, and performance requirements. You'll learn to analyze compression ratios, assess ingestion performance implications, and conduct systematic benchmarking of formats like Parquet, ORC, and Avro. Transform your ability to make informed storage architecture decisions that directly impact analytical performance and cost-effectiveness in enterprise data warehousing environments. This course is unique because it combines theoretical understanding with hands-on benchmarking practice, giving you real-world experience in evaluating storage formats using actual performance metrics. To be successful in this project, you should have basic understanding of data warehousing concepts and familiarity with SQL queries.
Watch on External: Coursera ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Learn how to build a production-ready ETL pipeline with Python, Docker, PostgreSQL, and Kestra by thinking like a data engineer
Towards Data Science
📰
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Learn how to efficiently transfer large volumes of data using JuiceFS Sync, which offers resumable sync, encryption, and bandwidth control, ideal for PB-scale data transfers.
Dev.to AI
📰
How Airflow is using AI to make data engineering more resilient, not more complex
Airflow uses AI to make data engineering more resilient by detecting data drift, resuming failed pipelines, and fixing issues automatically, reducing complexity and improving reliability.
Medium · AI
📰
What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
Learn how to overcome memory bottlenecks in data engineering using Pandas chunking, Dask, and Polars, and why it matters for processing large datasets
Towards Data Science
Up next
A Moment Frozen in Time | Arnav Iyengar | TEDxJenks Youth
TEDx Talks
Watch →