✕ Clear all filters
141 articles
▶ Videos →

Articles

141 articles · Updated every 3 hours · View all reads

All Articles 185,344Blog Posts 168,177Tech Tutorials 49,634Research Papers 36,197News 22,822 ⚡ AI Lessons
BUCKETING VS. TIME-PARTITIONING IN ICEBERG
Medium · Data Science 🔄 Data Engineering 3w ago
BUCKETING VS. TIME-PARTITIONING IN ICEBERG
How a single Python list controls the storage layout of ~180 CDC tables and why the busiest tables get the fewest buckets Continue reading on Medium »
Databricks Lakehouse Architecture: The Modern Way to Store, Manage, and Access Data
Medium · Data Science 🔄 Data Engineering 3w ago
Databricks Lakehouse Architecture: The Modern Way to Store, Manage, and Access Data
Initially, organizations relied heavily on data warehouses, storing structured data in traditional databases like MySQL and building… Continue reading on Medium
How AI Agents Fix Data Pipelines Failures Before You Wake Up.
Medium · Data Science 🔄 Data Engineering 3w ago
How AI Agents Fix Data Pipelines Failures Before You Wake Up.
Today’s pipelines send alerts. Tomorrow’s pipelines will investigate, validate, and recover before engineers even wake up. Continue reading on Towards Data Engi
Data Lakes, Data Warehouses & More (Without the Jargon)
Medium · Data Science 🔄 Data Engineering 3w ago
Data Lakes, Data Warehouses & More (Without the Jargon)
If you’ve spent any time around a data team, you’ve probably heard people throw around words like “data lake,” “warehouse,” “lakehouse,”… Continue reading on Me
Terminate PySpark Spark Sessions | Apache Iceberg Guide
Medium · Data Science 🔄 Data Engineering 3w ago
Terminate PySpark Spark Sessions | Apache Iceberg Guide
Efficiently manage and terminate Spark sessions in PySpark using Apache Iceberg for optimized data workflows and resource utilization Continue reading on Medium
Medium · Python 🔄 Data Engineering 4w ago
An Introduction to Lakeflow Declarative Pipelines: From Definition to S3 Storage
Databricks has been moving data engineering toward a declarative model. Lakeflow Declarative Pipelines is the current framework for… Continue reading on Medium
Databricks FILE: A Step Toward Making Unstructured Data a First-Class Citizen in the Lakehouse
Medium · RAG 🔄 Data Engineering 4w ago
Databricks FILE: A Step Toward Making Unstructured Data a First-Class Citizen in the Lakehouse
Why PDFs, images and documents are becoming part of the modern lakehouse — and what this means for AI, RAG and data engineering Continue reading on Medium »
Databricks vs Snowflake: Who Will Own the Enterprise AI Entry Point?
Hackernoon 🔄 Data Engineering 4w ago
Databricks vs Snowflake: Who Will Own the Enterprise AI Entry Point?
Enterprise AI is moving beyond models. The real competition is about data, context, governance, and task ownership.
Data Engineering for RAG: Building Reliable AI with Better Data Pipelines
Medium · LLM 🔄 Data Engineering 1mo ago
Data Engineering for RAG: Building Reliable AI with Better Data Pipelines
Introduction Continue reading on Towards AI »
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Medium · Machine Learning 🔄 Data Engineering 1mo ago
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Databricks was founded in 2013 by the original creators of Apache Spark at UC Berkeley’s AMPLab. The founders open-sourced Spark and then… Continue reading on M
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Medium · Data Science 🔄 Data Engineering 1mo ago
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Databricks was founded in 2013 by the original creators of Apache Spark at UC Berkeley’s AMPLab. The founders open-sourced Spark and then… Continue reading on M
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Medium · Python 🔄 Data Engineering 1mo ago
Databricks Job Seeker Tips 25 Interview Questions Answer You Should Need to Know || Full Course…
Databricks was founded in 2013 by the original creators of Apache Spark at UC Berkeley’s AMPLab. The founders open-sourced Spark and then… Continue reading on M
Understanding Atlan MCP: Metadata Control for AI and Data
Medium · Data Science 🔄 Data Engineering 1mo ago
Understanding Atlan MCP: Metadata Control for AI and Data
Atlan MCP transforms metadata into real-time control for Snowflake, Databricks, and AI agents Continue reading on Medium »
LLM Evaluation for Data Pipelines: LangSmith, TruLens, Ragas & Snowflake Cortex Search Ops
Medium · Python 🔄 Data Engineering 1mo ago
LLM Evaluation for Data Pipelines: LangSmith, TruLens, Ragas & Snowflake Cortex Search Ops
llm evaluation for data pipelines is the load-bearing correctness discipline of the 2026 data stack — the difference between a RAG chatbot… Continue reading on
Stop Debugging Your Data Pipeline at 3 AM: A Practical Validation Framework with Airflow, Python…
Medium · Python 🔄 Data Engineering 1mo ago
Stop Debugging Your Data Pipeline at 3 AM: A Practical Validation Framework with Airflow, Python…
Subtitle: From “hope and pray” to “validate early, validate often” — how to catch bad data before it reaches your CEO’s dashboard. Continue reading on Medium »
AWS Machine Learning 🔄 Data Engineering 1mo ago
Agentic Data Operations Platform (ADOP): Data engineering into hours
The Agentic Data Operations Platform (ADOP) is a reference architecture on Amazon Bedrock that uses specialized AI agents to automate the full Bronze-to-Silver-
A Reference Architecture for AI-Driven Healthcare Data Engineering
Hackernoon 🔄 Data Engineering 1mo ago
A Reference Architecture for AI-Driven Healthcare Data Engineering
Healthcare data platforms are evolving beyond ETL, using AI for anomaly detection, entity matching, forecasting, compliance, and data quality.
Refactoring data pipelines with LLMs: notes from a SSIS to dbt migration
Medium · LLM 🔄 Data Engineering 1mo ago
Refactoring data pipelines with LLMs: notes from a SSIS to dbt migration
Migrating ~1000 stored procedures and ~100 SSIS jobs to dbt with Claude Code: equivalence testing, decomposition, context engineering. Continue reading on Mediu
The Silent Fracture of Modern Data Engineering
Medium · Python 🔄 Data Engineering 1mo ago
The Silent Fracture of Modern Data Engineering
Most production data pipelines do not fail because their underlying mathematical algorithms are flawed they implode because human domain… Continue reading on IL
Top 50 Data Architecture Patterns Interview Questions and Answers
Medium · Data Science 🔄 Data Engineering 1mo ago
Top 50 Data Architecture Patterns Interview Questions and Answers
Data Engineering, Lakehouse, Medallion, Streaming, Data Mesh, Data Vault & Modern Platform Architecture Continue reading on Medium »