
The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI How Airflow orchestration decisions impact Spark performance
6 snips
Aug 13, 2026 Meni Shmueli, Co-founder and CEO of DataFlint and former data engineer focused on Spark performance. He explores how Airflow orchestration choices can make Spark jobs faster or far more costly. Short stories show parallelism pitfalls and serial-run waste. He explains DataFlint’s Airflow/Astro integrations and how AI enables holistic scheduling, cost-aware optimization, and huge real-world savings.
AI Snips
Chapters
Transcript
Episode notes
Airflow Orchestrates While Spark Does The Heavy Lifting
- Airflow schedules and Spark execute complementary roles: Airflow decides when and how pipelines run while Spark performs heavy distributed compute.
- Meni Shmueli says most customers use Airflow to order pipelines and Spark as the underlying engine moving data at scale.
Parallelizing Tasks Made Spark Slower And Costlier
- Parallelizing independent Spark tasks in Airflow led one customer to slower, unstable, and far more expensive runs because tasks contended for the same cluster resources.
- Meni explains contention caused spills to disk and crashes; sequential runs used memory and ran faster.
Benchmark Pipelines By Duration And Underlying Spark Cost
- Do benchmark pipelines with both Airflow and Spark contexts; measure duration alongside actual Spark resource usage and cost.
- Meni warns a faster DAG runtime can hide a jump from 100 to 1,000 cluster machines and huge cost increases.
