Many reliability challenges in workflow orchestration aren’t unique to Airflow — they’re the same failure-isolation, retry, and dependency problems that show up across large-scale distributed systems.
In this lightning talk, I’ll distill a handful of battle-tested patterns from high-throughput production systems — bounding blast radius so one failure doesn’t cascade, designing retries and backoff that don’t overwhelm downstream services, and observability that surfaces issues early — and connect each to how data teams can think about resilience in their Airflow deployments. The goal is a tight set of transferable principles attendees can apply immediately.
Tejas Pravinbhai Patel
IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect