The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI

How Airflow orchestration decisions impact Spark performance

6 snips
Aug 13, 2026
Meni Shmueli, Co-founder and CEO of DataFlint and former data engineer focused on Spark performance. He explores how Airflow orchestration choices can make Spark jobs faster or far more costly. Short stories show parallelism pitfalls and serial-run waste. He explains DataFlint’s Airflow/Astro integrations and how AI enables holistic scheduling, cost-aware optimization, and huge real-world savings.
Ask episode
AI Snips
Chapters
Transcript
Episode notes
INSIGHT

Airflow Orchestrates While Spark Does The Heavy Lifting

  • Airflow schedules and Spark execute complementary roles: Airflow decides when and how pipelines run while Spark performs heavy distributed compute.
  • Meni Shmueli says most customers use Airflow to order pipelines and Spark as the underlying engine moving data at scale.
ANECDOTE

Parallelizing Tasks Made Spark Slower And Costlier

  • Parallelizing independent Spark tasks in Airflow led one customer to slower, unstable, and far more expensive runs because tasks contended for the same cluster resources.
  • Meni explains contention caused spills to disk and crashes; sequential runs used memory and ran faster.
ADVICE

Benchmark Pipelines By Duration And Underlying Spark Cost

  • Do benchmark pipelines with both Airflow and Spark contexts; measure duration alongside actual Spark resource usage and cost.
  • Meni warns a faster DAG runtime can hide a jump from 100 to 1,000 cluster machines and huge cost increases.
Get the Snipd Podcast app to discover more snips from this episode
Get the app