Check out the full program for Airflow Summit.

If you prefer, you can also see this as sessionize layout or list of sessions.

All times in Coordinated Universal Time (UTC).

Wednesday, November 04, 2026

15:00
15:30
16:00
16:30
17:00
20:00
20:30
21:00
21:30
15:00 - 15:30.
By Ashir Alam
Track: Data & AI Applications
Room: Online
11/04/2026 3:00 PM 11/04/2026 3:30 PM America/Chicago AS26: The Alert Is Firing. Now What? LLM-Assisted Triage for Data Quality DAGs

At Airflow Summit 2025 we showed how we generate data quality DAGs from YAML using DAGFactory, putting check authoring in the hands of 30+ engineers, analysts, and BI users on Cloud Composer. It worked and it created a new problem. Check coverage grew faster than our ability to respond to it, and on-call was spending a median of one to two hours per alert just deciding whether it was real.

Almost none of that time was judgment. It was gathering: querying historical values for the failing metric, checking upstream task state, looking for schema drift, correlating recent deploys, comparing sibling checks on the same table. All of it was context Airflow already had.

This talk covers how we turned that gathering loop into a triage DAG. A failing check triggers parallel context-assembly tasks, hands the result to an LLM constrained to a strict output schema, and posts a structured verdict to Slack: true positive, false positive, or inconclusive, with confidence, reasoning, and cited evidence. The engineer approves a rerun, suppresses, or escalates with one click - Airflow stays in control of every action taken. Median time to a triage decision dropped from one to two hours to under ten minutes.

We’ll walk through the DAG design, the output schema, how the Slack approval loop doubles as a labeled evaluation set, and what we measure to know the model isn’t quietly wrong, including false-negative rate and how often engineers override the verdict. We’ll also cover what broke along the way: hallucinated root causes, context bloat, cost per triage, and the open question of whether humans keep reviewing carefully once the verdicts start getting good.

Online

At Airflow Summit 2025 we showed how we generate data quality DAGs from YAML using DAGFactory, putting check authoring in the hands of 30+ engineers, analysts, and BI users on Cloud Composer. It worked and it created a new problem. Check coverage grew faster than our ability to respond to it, and on-call was spending a median of one to two hours per alert just deciding whether it was real.

15:30 - 16:00.
By Rafael Pierre
Track: Data & AI Applications
Room: Online
11/04/2026 3:30 PM 11/04/2026 4:00 PM America/Chicago AS26: Who Judges the Agents? Orchestrating AI Evals with Airflow

Agentic AI systems are difficult to ship safely: their behavior is non-deterministic, traditional assertions only cover part of the problem, and manual evaluation quickly becomes a release bottleneck.

In this talk, I’ll share how we built Pegasus, an automated evaluation platform for agentic applications, and used Apache Airflow as its orchestration backbone. The platform combines deterministic tests with LLM-as-a-Judge and Agents-as-Judges evaluations. Airflow coordinates the evaluation harness across scheduled and on-demand runs. a DAG checks the commit currently deployed across different environments, determines whether that version has already been evaluated, and runs the required test suites only when needed. The same workflow can also be triggered asynchronously during development and release processes.

This approach automated roughly 80% of tests that previously required manual execution and helped move our production release cadence from about once a month to twice or more a week.

Attendees will leave with a practical architecture for orchestrating AI evaluations with Airflow, patterns for combining scheduled and on-demand evaluation workflows, and lessons for making agentic AI testing repeatable, observable, and useful as part of a production release process.

Online

Agentic AI systems are difficult to ship safely: their behavior is non-deterministic, traditional assertions only cover part of the problem, and manual evaluation quickly becomes a release bottleneck.

In this talk, I’ll share how we built Pegasus, an automated evaluation platform for agentic applications, and used Apache Airflow as its orchestration backbone. The platform combines deterministic tests with LLM-as-a-Judge and Agents-as-Judges evaluations. Airflow coordinates the evaluation harness across scheduled and on-demand runs. a DAG checks the commit currently deployed across different environments, determines whether that version has already been evaluated, and runs the required test suites only when needed. The same workflow can also be triggered asynchronously during development and release processes.

16:00 - 16:30.
By Vikram Koka
Track: Data Strategy
Room: Online
11/04/2026 4:00 PM 11/04/2026 4:30 PM America/Chicago AS26: The State of Airflow

Airflow 3 has been out for a year, and the systems we are being asked to build are changing quickly. In this keynote, we ask a broader question: what does orchestration need to become as production workloads get more intelligent, dynamic, and distributed?

Airflow has always been about coordinating work, managing dependencies, recovering from failure, and running reliably at scale. Those problems do not go away in the AI era. In many ways, they become more important.

We will look at how Airflow is evolving in response: supporting richer execution models, working across a much broader set of models, frameworks, infrastructure, and data systems, and giving teams the controls they need to run these workloads confidently in production.

We close with what this means for Data Engineers. AI is changing the abstractions we work with and the systems we are responsible for. The opportunity for Data Engineering is to take these new capabilities and make them dependable enough to run in production.

Twenty-five minutes. A lot of ground to cover

Online

Airflow 3 has been out for a year, and the systems we are being asked to build are changing quickly. In this keynote, we ask a broader question: what does orchestration need to become as production workloads get more intelligent, dynamic, and distributed?

Airflow has always been about coordinating work, managing dependencies, recovering from failure, and running reliably at scale. Those problems do not go away in the AI era. In many ways, they become more important.

16:30 - 17:00.
Track: Sponsored
Room: Online
11/04/2026 4:30 PM 11/04/2026 5:00 PM America/Chicago AS26: Astronomer Sponsored

TBD

Online
17:00 - 17:30.
Track: Sponsored
Room: Online
11/04/2026 5:00 PM 11/04/2026 5:30 PM America/Chicago AS26: Google Sponsored

TBD

Online
20:00 - 20:30.
By Saket Milind Karve
Track: Data & AI Applications
Room: Online
11/04/2026 8:00 PM 11/04/2026 8:30 PM America/Chicago AS26: Compositional DAGs in Airflow: Fixing the Hidden Data Bottleneck in Physical AI Training

Robot foundation model training is data-bound, not compute-bound. GPUs sit idle while heterogeneous trajectory data, synchronized camera streams, IMU, joint-state, and force/torque signals, waits to be transformed into training-ready tensors. Production teams generate tens of terabytes daily, and generic data-loader patterns don’t compose the way this data needs.

This talk shares a compositional DAG architecture built on Airflow for exactly this problem. Each ingestion step declares its own CPU/GPU/memory profile, ships as a versioned container image, fans out across Kubernetes workers using a distribute-wait pattern, and is selectively gated to skip re-execution of unaffected steps when downstream logic changes. Business logic runs off-orchestrator so the scheduler’s worker pool never saturates under compute load.

In production, this eliminated over 30% of training slowdown while coordinating an 88 TB cache across 16 H100 nodes.

Attendees will leave with concrete DAG design patterns for per-step resource specification, selective re-execution, and orchestrator/compute separation, applicable to any team running heterogeneous, resource-heavy pipelines on Airflow, not just robotics.

Online

Robot foundation model training is data-bound, not compute-bound. GPUs sit idle while heterogeneous trajectory data, synchronized camera streams, IMU, joint-state, and force/torque signals, waits to be transformed into training-ready tensors. Production teams generate tens of terabytes daily, and generic data-loader patterns don’t compose the way this data needs.

This talk shares a compositional DAG architecture built on Airflow for exactly this problem. Each ingestion step declares its own CPU/GPU/memory profile, ships as a versioned container image, fans out across Kubernetes workers using a distribute-wait pattern, and is selectively gated to skip re-execution of unaffected steps when downstream logic changes. Business logic runs off-orchestrator so the scheduler’s worker pool never saturates under compute load.

20:30 - 21:00.
Track: Sponsored
Room: Online
11/04/2026 8:30 PM 11/04/2026 9:00 PM America/Chicago AS26: Broadcom Sponsored

TBD

Online
21:00 - 21:30.
By Serjesh Sharma & John MacMillan
Track: Data & AI Applications
Room: Online
11/04/2026 9:00 PM 11/04/2026 9:30 PM America/Chicago AS26: Architecting Multi-Tenant ML Governance & Privacy in Airflow

In enterprise SaaS, orchestrating ML models on sensitive workforce data requires zero compromise on data privacy and multi-tenant isolation. Workday’s Scheduling & Labor Optimization product uses Airflow to power predictive workforce forecasts across thousands of customer tenants—where each customer strictly owns their data and model artifacts. This session demonstrates how Workday leverages Airflow for multi-tenant ML pipeline orchestration. We cover how Airflow dynamically materializes isolated ML DAGs per tenant, enforces strict execution boundaries via TENANT_KEY metadata, and automates lifecycle operations for customer opt-in and opt-out workflows, including mandatory 14-day data/model purges.
Key Problems Solved:

  1. Strict Data & Model Privacy: Isolating tenant data and training models exclusively on tenant-owned datasets.
  2. Lifecycle Compliance: Triggering DAG creation on opt-in and managing grace periods with enforced 14-day purges on opt-out. 3. Operational Efficiency: Scaling thousands of dynamic DAGs using VPA/Scaleops autoscaling without scheduler overhead.

Session Outline :

  1. Overview & Challenge
  2. Multi-tenant ML privacy in SaaS.
  3. Dynamic DAG Architecture
  4. Event-driven materialization and TENANT_KEY isolation.
  5. Lifecycle Compliance
  6. Dynamic enrollment and automated 14-day deletion pipelines. Scaling & Infrastructure
  7. Resource optimization with ScaleOps/VPA

Takeaways:

  1. Design patterns for dynamic per-tenant DAG generation.
  2. Enforcing tenant isolation, security boundaries, and model governance in Airflow.
  3. Building compliance workflows for dynamic opt-in/opt-out retention policies.
Online

In enterprise SaaS, orchestrating ML models on sensitive workforce data requires zero compromise on data privacy and multi-tenant isolation. Workday’s Scheduling & Labor Optimization product uses Airflow to power predictive workforce forecasts across thousands of customer tenants—where each customer strictly owns their data and model artifacts. This session demonstrates how Workday leverages Airflow for multi-tenant ML pipeline orchestration. We cover how Airflow dynamically materializes isolated ML DAGs per tenant, enforces strict execution boundaries via TENANT_KEY metadata, and automates lifecycle operations for customer opt-in and opt-out workflows, including mandatory 14-day data/model purges.
Key Problems Solved:

21:30 - 22:00.
Track: Sponsored
Room: Online
11/04/2026 9:30 PM 11/04/2026 10:00 PM America/Chicago AS26: BMC Sponsored

TBD

Online
15:00 - 15:30. Online
By Ashir Alam
Track: Data & AI Applications
11/04/2026 3:00 PM 11/04/2026 3:30 PM America/Chicago AS26: The Alert Is Firing. Now What? LLM-Assisted Triage for Data Quality DAGs

At Airflow Summit 2025 we showed how we generate data quality DAGs from YAML using DAGFactory, putting check authoring in the hands of 30+ engineers, analysts, and BI users on Cloud Composer. It worked and it created a new problem. Check coverage grew faster than our ability to respond to it, and on-call was spending a median of one to two hours per alert just deciding whether it was real.

Almost none of that time was judgment. It was gathering: querying historical values for the failing metric, checking upstream task state, looking for schema drift, correlating recent deploys, comparing sibling checks on the same table. All of it was context Airflow already had.

This talk covers how we turned that gathering loop into a triage DAG. A failing check triggers parallel context-assembly tasks, hands the result to an LLM constrained to a strict output schema, and posts a structured verdict to Slack: true positive, false positive, or inconclusive, with confidence, reasoning, and cited evidence. The engineer approves a rerun, suppresses, or escalates with one click - Airflow stays in control of every action taken. Median time to a triage decision dropped from one to two hours to under ten minutes.

We’ll walk through the DAG design, the output schema, how the Slack approval loop doubles as a labeled evaluation set, and what we measure to know the model isn’t quietly wrong, including false-negative rate and how often engineers override the verdict. We’ll also cover what broke along the way: hallucinated root causes, context bloat, cost per triage, and the open question of whether humans keep reviewing carefully once the verdicts start getting good.

Online

At Airflow Summit 2025 we showed how we generate data quality DAGs from YAML using DAGFactory, putting check authoring in the hands of 30+ engineers, analysts, and BI users on Cloud Composer. It worked and it created a new problem. Check coverage grew faster than our ability to respond to it, and on-call was spending a median of one to two hours per alert just deciding whether it was real.

15:30 - 16:00. Online
By Rafael Pierre
Track: Data & AI Applications
11/04/2026 3:30 PM 11/04/2026 4:00 PM America/Chicago AS26: Who Judges the Agents? Orchestrating AI Evals with Airflow

Agentic AI systems are difficult to ship safely: their behavior is non-deterministic, traditional assertions only cover part of the problem, and manual evaluation quickly becomes a release bottleneck.

In this talk, I’ll share how we built Pegasus, an automated evaluation platform for agentic applications, and used Apache Airflow as its orchestration backbone. The platform combines deterministic tests with LLM-as-a-Judge and Agents-as-Judges evaluations. Airflow coordinates the evaluation harness across scheduled and on-demand runs. a DAG checks the commit currently deployed across different environments, determines whether that version has already been evaluated, and runs the required test suites only when needed. The same workflow can also be triggered asynchronously during development and release processes.

This approach automated roughly 80% of tests that previously required manual execution and helped move our production release cadence from about once a month to twice or more a week.

Attendees will leave with a practical architecture for orchestrating AI evaluations with Airflow, patterns for combining scheduled and on-demand evaluation workflows, and lessons for making agentic AI testing repeatable, observable, and useful as part of a production release process.

Online

Agentic AI systems are difficult to ship safely: their behavior is non-deterministic, traditional assertions only cover part of the problem, and manual evaluation quickly becomes a release bottleneck.

In this talk, I’ll share how we built Pegasus, an automated evaluation platform for agentic applications, and used Apache Airflow as its orchestration backbone. The platform combines deterministic tests with LLM-as-a-Judge and Agents-as-Judges evaluations. Airflow coordinates the evaluation harness across scheduled and on-demand runs. a DAG checks the commit currently deployed across different environments, determines whether that version has already been evaluated, and runs the required test suites only when needed. The same workflow can also be triggered asynchronously during development and release processes.

16:00 - 16:30. Online
By Vikram Koka
Track: Data Strategy
11/04/2026 4:00 PM 11/04/2026 4:30 PM America/Chicago AS26: The State of Airflow

Airflow 3 has been out for a year, and the systems we are being asked to build are changing quickly. In this keynote, we ask a broader question: what does orchestration need to become as production workloads get more intelligent, dynamic, and distributed?

Airflow has always been about coordinating work, managing dependencies, recovering from failure, and running reliably at scale. Those problems do not go away in the AI era. In many ways, they become more important.

We will look at how Airflow is evolving in response: supporting richer execution models, working across a much broader set of models, frameworks, infrastructure, and data systems, and giving teams the controls they need to run these workloads confidently in production.

We close with what this means for Data Engineers. AI is changing the abstractions we work with and the systems we are responsible for. The opportunity for Data Engineering is to take these new capabilities and make them dependable enough to run in production.

Twenty-five minutes. A lot of ground to cover

Online

Airflow 3 has been out for a year, and the systems we are being asked to build are changing quickly. In this keynote, we ask a broader question: what does orchestration need to become as production workloads get more intelligent, dynamic, and distributed?

Airflow has always been about coordinating work, managing dependencies, recovering from failure, and running reliably at scale. Those problems do not go away in the AI era. In many ways, they become more important.

16:30 - 17:00. Online
Track: Sponsored
11/04/2026 4:30 PM 11/04/2026 5:00 PM America/Chicago AS26: Astronomer Sponsored

TBD

Online
17:00 - 17:30. Online
Track: Sponsored
11/04/2026 5:00 PM 11/04/2026 5:30 PM America/Chicago AS26: Google Sponsored

TBD

Online
20:00 - 20:30. Online
By Saket Milind Karve
Track: Data & AI Applications
11/04/2026 8:00 PM 11/04/2026 8:30 PM America/Chicago AS26: Compositional DAGs in Airflow: Fixing the Hidden Data Bottleneck in Physical AI Training

Robot foundation model training is data-bound, not compute-bound. GPUs sit idle while heterogeneous trajectory data, synchronized camera streams, IMU, joint-state, and force/torque signals, waits to be transformed into training-ready tensors. Production teams generate tens of terabytes daily, and generic data-loader patterns don’t compose the way this data needs.

This talk shares a compositional DAG architecture built on Airflow for exactly this problem. Each ingestion step declares its own CPU/GPU/memory profile, ships as a versioned container image, fans out across Kubernetes workers using a distribute-wait pattern, and is selectively gated to skip re-execution of unaffected steps when downstream logic changes. Business logic runs off-orchestrator so the scheduler’s worker pool never saturates under compute load.

In production, this eliminated over 30% of training slowdown while coordinating an 88 TB cache across 16 H100 nodes.

Attendees will leave with concrete DAG design patterns for per-step resource specification, selective re-execution, and orchestrator/compute separation, applicable to any team running heterogeneous, resource-heavy pipelines on Airflow, not just robotics.

Online

Robot foundation model training is data-bound, not compute-bound. GPUs sit idle while heterogeneous trajectory data, synchronized camera streams, IMU, joint-state, and force/torque signals, waits to be transformed into training-ready tensors. Production teams generate tens of terabytes daily, and generic data-loader patterns don’t compose the way this data needs.

This talk shares a compositional DAG architecture built on Airflow for exactly this problem. Each ingestion step declares its own CPU/GPU/memory profile, ships as a versioned container image, fans out across Kubernetes workers using a distribute-wait pattern, and is selectively gated to skip re-execution of unaffected steps when downstream logic changes. Business logic runs off-orchestrator so the scheduler’s worker pool never saturates under compute load.

20:30 - 21:00. Online
Track: Sponsored
11/04/2026 8:30 PM 11/04/2026 9:00 PM America/Chicago AS26: Broadcom Sponsored

TBD

Online
21:00 - 21:30. Online
By Serjesh Sharma & John MacMillan
Track: Data & AI Applications
11/04/2026 9:00 PM 11/04/2026 9:30 PM America/Chicago AS26: Architecting Multi-Tenant ML Governance & Privacy in Airflow

In enterprise SaaS, orchestrating ML models on sensitive workforce data requires zero compromise on data privacy and multi-tenant isolation. Workday’s Scheduling & Labor Optimization product uses Airflow to power predictive workforce forecasts across thousands of customer tenants—where each customer strictly owns their data and model artifacts. This session demonstrates how Workday leverages Airflow for multi-tenant ML pipeline orchestration. We cover how Airflow dynamically materializes isolated ML DAGs per tenant, enforces strict execution boundaries via TENANT_KEY metadata, and automates lifecycle operations for customer opt-in and opt-out workflows, including mandatory 14-day data/model purges.
Key Problems Solved:

  1. Strict Data & Model Privacy: Isolating tenant data and training models exclusively on tenant-owned datasets.
  2. Lifecycle Compliance: Triggering DAG creation on opt-in and managing grace periods with enforced 14-day purges on opt-out. 3. Operational Efficiency: Scaling thousands of dynamic DAGs using VPA/Scaleops autoscaling without scheduler overhead.

Session Outline :

  1. Overview & Challenge
  2. Multi-tenant ML privacy in SaaS.
  3. Dynamic DAG Architecture
  4. Event-driven materialization and TENANT_KEY isolation.
  5. Lifecycle Compliance
  6. Dynamic enrollment and automated 14-day deletion pipelines. Scaling & Infrastructure
  7. Resource optimization with ScaleOps/VPA

Takeaways:

  1. Design patterns for dynamic per-tenant DAG generation.
  2. Enforcing tenant isolation, security boundaries, and model governance in Airflow.
  3. Building compliance workflows for dynamic opt-in/opt-out retention policies.
Online

In enterprise SaaS, orchestrating ML models on sensitive workforce data requires zero compromise on data privacy and multi-tenant isolation. Workday’s Scheduling & Labor Optimization product uses Airflow to power predictive workforce forecasts across thousands of customer tenants—where each customer strictly owns their data and model artifacts. This session demonstrates how Workday leverages Airflow for multi-tenant ML pipeline orchestration. We cover how Airflow dynamically materializes isolated ML DAGs per tenant, enforces strict execution boundaries via TENANT_KEY metadata, and automates lifecycle operations for customer opt-in and opt-out workflows, including mandatory 14-day data/model purges.
Key Problems Solved:

21:30 - 22:00. Online
Track: Sponsored
11/04/2026 9:30 PM 11/04/2026 10:00 PM America/Chicago AS26: BMC Sponsored

TBD

Online

Thursday, November 5, 2026

06:00
06:30
07:00
07:30
13:00
13:30
14:00
14:30
15:00
13:00 - 13:30.
By Piyush Maheshwari & Sameer Raj
Track: Builder
Room: Online
11/05/2026 1:00 PM 11/05/2026 1:30 PM America/Chicago AS26: Microservice-Grade Delivery and Release for Airflow DAGs

At Uber, preparing Airflow to take on workloads from Piper (our Airflow 1 fork operating at nearly one million daily task runs) requires rethinking both DAG delivery and release. Shipping a multi-GB monorepo artifact to isolated Kubernetes executors for every task is neither fast nor efficient.

We’ll share how dependency-aware slim bundles package only the required DAG code, first-party dependencies and generated artifacts into reproducible bundles. We’ll then cover per-DAG version pinning, which separates code distribution from activation and enables pre-production regression gates, controlled promotion and automated rollback.

Attendees will learn practical patterns and trade-offs for monorepo dependency resolution, bundle granularity, Kubernetes execution and safer DAG releases. We’ll also discuss AIP-109, our proposal to contribute DAG version pinning to Apache Airflow.

Online

At Uber, preparing Airflow to take on workloads from Piper (our Airflow 1 fork operating at nearly one million daily task runs) requires rethinking both DAG delivery and release. Shipping a multi-GB monorepo artifact to isolated Kubernetes executors for every task is neither fast nor efficient.

We’ll share how dependency-aware slim bundles package only the required DAG code, first-party dependencies and generated artifacts into reproducible bundles. We’ll then cover per-DAG version pinning, which separates code distribution from activation and enables pre-production regression gates, controlled promotion and automated rollback.

13:30 - 14:00.
By Satej Sahu
Track: Builder
Room: Online
11/05/2026 1:30 PM 11/05/2026 2:00 PM America/Chicago AS26: Adaptive DAGs: Dynamic Query Plan Rewriting and Runtime Task Re-sizing in Airflow

Traditional Airflow DAGs execute rigid, pre-determined computation graphs. Yet in production, data volume and distribution fluctuate wildly: a daily ingestion batch may process 50,000 rows on Monday but 50,000,000 on Black Friday. Statically sized tasks force engineers into a painful compromise: either permanently over-provision worker resources (blowing cloud budgets) or risk out-of-memory crashes, disk spills, and missed SLAs during unexpected volume spikes.

Drawing inspiration from database query optimizer research and recent papers on learned adaptive pipeline scheduling, this technical session demonstrates how to implement Adaptive DAGs in Apache Airflow. We show how to construct workflows that introspect upstream intermediate data statistics to dynamically adjust downstream execution plans, engine targets, and parallelism at runtime. Attendees will learn:

  1. Runtime Plan Profiling: Using lightweight metadata collectors and storage manifests (Iceberg/Delta metadata) to capture data volume, partition skew, and cardinality before heavy compute stages run.
  2. Dynamic Strategy Routing: Leveraging Airflow’s TaskFlow API and Dynamic Task Mapping (.expand()) to conditionally route jobs between lightweight single-node compute (DuckDB/Polars) for small datasets and distributed clusters (Spark/Ray) for large scale.
  3. Adaptive Partition Resizing: Dynamically generating downstream worker parallelism and memory allocations on the fly, eliminating OOM failures and compute waste.

Attendees will leave with production-ready code patterns to build self-tuning, resilient Airflow pipelines.

Online

Traditional Airflow DAGs execute rigid, pre-determined computation graphs. Yet in production, data volume and distribution fluctuate wildly: a daily ingestion batch may process 50,000 rows on Monday but 50,000,000 on Black Friday. Statically sized tasks force engineers into a painful compromise: either permanently over-provision worker resources (blowing cloud budgets) or risk out-of-memory crashes, disk spills, and missed SLAs during unexpected volume spikes.

Drawing inspiration from database query optimizer research and recent papers on learned adaptive pipeline scheduling, this technical session demonstrates how to implement Adaptive DAGs in Apache Airflow. We show how to construct workflows that introspect upstream intermediate data statistics to dynamically adjust downstream execution plans, engine targets, and parallelism at runtime. Attendees will learn:

14:00 - 14:30.
By Shrividya Hegde
Track: Builder
Room: Online
11/05/2026 2:00 PM 11/05/2026 2:30 PM America/Chicago AS26: Self-Healing Airflow: Pipelines That Fix Themselves

It’s 2 AM and a task just failed. Does someone get paged, or does the pipeline quietly diagnose the problem, wait the right amount of time, retry, and only escalate when it truly can’t recover on its own?

Most Airflow pipelines are babysat. This session shows how to make them self-healing instead. We’ll work up from the 80% fix , retries with exponential back off ,through failure and retry callbacks, SLA monitoring, automatic clearing of stuck and zombie tasks, and circuit breakers that stop hammering a broken downstream system. You’ll see how to route alerts to Slack or PagerDuty so humans are interrupted only when it matters, and how to keep secrets out of your DAGs while doing it.

Coming from a testing background, I’ll frame this the way an SRE would: design for failure first. You’ll leave with concrete patterns and a template you can drop into your own DAGs on Monday.

Online

It’s 2 AM and a task just failed. Does someone get paged, or does the pipeline quietly diagnose the problem, wait the right amount of time, retry, and only escalate when it truly can’t recover on its own?

Most Airflow pipelines are babysat. This session shows how to make them self-healing instead. We’ll work up from the 80% fix , retries with exponential back off ,through failure and retry callbacks, SLA monitoring, automatic clearing of stuck and zombie tasks, and circuit breakers that stop hammering a broken downstream system. You’ll see how to route alerts to Slack or PagerDuty so humans are interrupted only when it matters, and how to keep secrets out of your DAGs while doing it.

14:30 - 15:00.
Track: Sponsored
Room: Online
11/05/2026 2:30 PM 11/05/2026 3:00 PM America/Chicago AS26: IBM Sponsored

TBD

Online
15:00 - 15:30.
By Vinod Jayendra, Sean Bjurstrom & Abdul Majid Mohammed
Track: Data Strategy
Room: Online
11/05/2026 3:00 PM 11/05/2026 3:30 PM America/Chicago AS26: Ten Teams, One Airflow Cluster, No Guardrails. Until Now.

Your Airflow environment started as one team’s orchestrator. Now three teams share it, each wants their own connections, variables, and execution roles. Your security review flagged that every DAG runs with the same permissions. Sound familiar?

We’ll tackle multi-tenancy head-on. Most Airflow deployments have a single execution context, one DAGs folder, and limited native isolation between teams. Here are the patterns that work in production.

Runtime resource isolation: per-task credential scoping so each team’s DAGs only access their own data stores and APIs. DAG-level access control with automated tag-based RBAC that syncs identity attributes to Airflow roles on a schedule, so new DAGs inherit the right permissions without manual work. And the deployment tradeoff: single shared instance with RBAC guardrails vs. separate environments per team.

We’ll show how AI agents speed up multi-team operations. Using Airflow 3’s task.agent decorator, we built an automated onboarding workflow that scans a team’s DAG repo, infers required connections and permissions, generates scoped access policies, and configures RBAC. All as auditable, retryable Airflow tasks instead of manual runbooks.

We’ll preview Airflow 3.2’s experimental multi-team support: separate DAGs, connections, variables, pools, and executors per team in a single instance.

You’ll leave with a blueprint for operating Airflow as an internal platform. Self-service onboarding, cost allocation per team, and security guardrails that don’t slow developers down.

Online

Your Airflow environment started as one team’s orchestrator. Now three teams share it, each wants their own connections, variables, and execution roles. Your security review flagged that every DAG runs with the same permissions. Sound familiar?

We’ll tackle multi-tenancy head-on. Most Airflow deployments have a single execution context, one DAGs folder, and limited native isolation between teams. Here are the patterns that work in production.

Runtime resource isolation: per-task credential scoping so each team’s DAGs only access their own data stores and APIs. DAG-level access control with automated tag-based RBAC that syncs identity attributes to Airflow roles on a schedule, so new DAGs inherit the right permissions without manual work. And the deployment tradeoff: single shared instance with RBAC guardrails vs. separate environments per team.

06:00 - 06:30.
By Yunhao Qing
Track: Data & AI Applications
Room: Online
11/05/2026 6:00 AM 11/05/2026 6:30 AM America/Chicago AS26: Airflow Across Four Regions and a Fleet of Agents: How Notion Runs Its Data Platform

Notion runs in four AWS regions today, with more on the way, and every data pipeline has to respect where a customer’s data lives. Airflow on Astronomer is the control plane that holds this together: two Astro deployments, 500+ DAGs, and region-aware routing to EMR, EMR Serverless, Ray on Anyscale, Kafka, Snowflake, Databricks, and our vector database.

This talk covers the platform patterns that let a small team operate all of that: a generated cell and region model shared by every DAG, so new regions and retired cells propagate without touching pipeline code; DAG factories that fan one config out per region; and per-environment, per-region compute isolation, including the custom Anyscale operator we built and the cloud migrations we ran from DAG config alone.

We then zoom into one concrete workload: indexing Jira, Slack, Google Drive, GitHub, and a dozen other connectors for Notion AI. Kafka feeds Spark and Ray embedding jobs that write to region-local vector indexes. This workload shows both where our abstractions held up and where they broke down, including the per-region DAG file explosion we are still paying down.

Attendees will leave with practical patterns for designing region-aware Airflow platforms: how to model regions and cells, generate DAGs without duplicating business logic, isolate compute by environment and geography, and evolve infrastructure as regions are added, migrated, or retired.

Online

Notion runs in four AWS regions today, with more on the way, and every data pipeline has to respect where a customer’s data lives. Airflow on Astronomer is the control plane that holds this together: two Astro deployments, 500+ DAGs, and region-aware routing to EMR, EMR Serverless, Ray on Anyscale, Kafka, Snowflake, Databricks, and our vector database.

This talk covers the platform patterns that let a small team operate all of that: a generated cell and region model shared by every DAG, so new regions and retired cells propagate without touching pipeline code; DAG factories that fan one config out per region; and per-environment, per-region compute isolation, including the custom Anyscale operator we built and the cloud migrations we ran from DAG config alone.

06:30 - 07:00.
By Aleksandr Shirokov & Roman Khomenko
Track: Data & AI Applications
Room: Online
11/05/2026 6:30 AM 11/05/2026 7:00 AM America/Chicago AS26: Beyond Multi-Cluster Airflow: Operating GPU Workloads at Scale

At last year’s Airflow Summit, we shared how we built a multi-cluster orchestration layer on top of Apache Airflow to run ML workloads across multiple Kubernetes GPU clusters.

Once hundreds of ML engineers started running GPU pipelines in production, we discovered that orchestration alone is not enough. Operating multi-cluster GPU infrastructure introduces new challenges: controlling GPU allocation across teams, observing pipelines across clusters, and helping users run workloads efficiently without wasting expensive GPU resources.

In this talk, we’ll show how our Airflow platform evolved from a workflow orchestrator into an operational control plane for GPU infrastructure. We’ll cover custom scheduling strategies that dynamically route workloads across clusters using Airflow policies and resource awareness, integration with HAMI to improve GPU utilization, and AIOps workflows with KeepHQ that detect underutilized CPU, RAM, and GPU resources. We’ll also present powerful dashboards and AI-assisted tools that reduce Time2Market and simplify debugging while keeping infrastructure complexity hidden.

We’ll be happy to share how our platform continues to evolve with Apache Airflow.

Online

At last year’s Airflow Summit, we shared how we built a multi-cluster orchestration layer on top of Apache Airflow to run ML workloads across multiple Kubernetes GPU clusters.

Once hundreds of ML engineers started running GPU pipelines in production, we discovered that orchestration alone is not enough. Operating multi-cluster GPU infrastructure introduces new challenges: controlling GPU allocation across teams, observing pipelines across clusters, and helping users run workloads efficiently without wasting expensive GPU resources.

07:00 - 07:30.
By Christos Bisias
Track: Builder
Room: Online
11/05/2026 7:00 AM 11/05/2026 7:30 AM America/Chicago AS26: Cross-Deployment Monitoring and Dag Dependencies with the Kafka Event Producer Plugin

Organizations can have multiple Airflow instances running on separate hardware, each with its own database, and sometimes in different geographical regions. How can we monitor these instances and get updates on their state without polling their databases or APIs? Is there a way to coordinate Dags from different Airflow instances?

This session answers the above questions by presenting a newly released plugin in the Apache Kafka Provider, which publishes a message to a configured Kafka topic for every Dag run or task instance state change event.

It will cover:

  • what the plugin does
  • what the published messages look like
  • how to enable and configure it
  • how to consume the messages
  • some use cases
Online

Organizations can have multiple Airflow instances running on separate hardware, each with its own database, and sometimes in different geographical regions. How can we monitor these instances and get updates on their state without polling their databases or APIs? Is there a way to coordinate Dags from different Airflow instances?

This session answers the above questions by presenting a newly released plugin in the Apache Kafka Provider, which publishes a message to a configured Kafka topic for every Dag run or task instance state change event.

07:30 - 08:00.
By Ramzi Alashabi
Track: Data Strategy
Room: Online
11/05/2026 7:30 AM 11/05/2026 8:00 AM America/Chicago AS26: Scale Without Scaling the Team: Metadata-Driven Airflow the Business Can Own

Every data team hits the same wall: a handful of engineers writing and maintaining every Airflow DAG, while the business waits in the queue.

We took a different path. Instead of hand-writing pipelines, we generate them from metadata, a declarative description of what each data product needs, compiled straight into asset-scheduled Airflow 3 DAGs. Source dependencies, partitioning, trigger conditions, retries: all derived, not typed. That alone let us stand up hundreds of pipelines without hundreds of hand-maintained files.

But auto-generation only takes you so far. Real orchestration needs human judgment: this product should wait for that one; these steps run in a specific order. So we gave the business a canvas drag-and-drop dependencies on top of the generated graph that compiles back into the same metadata. Analysts compose their own reusable orchestration; engineers stay out of the critical path.

I’ll walk through: • How to compile metadata into an asset-scheduled Airflow DAG, what to derive automatically, and what to leave configurable. • How a visual, drag-and-drop dependency graph maps cleanly back to declarative orchestration, not throwaway clicks. • Where auto-generation ends and human-authored orchestration begins and how to let both live in one source of truth. • What this does to delivery speed when the people who understand the data can ship the pipeline.

You’ll leave with a blueprint for metadata-driven DAG generation plus business-owned orchestration a way to scale pipeline delivery without scaling your DAG-writing team.

Online

Every data team hits the same wall: a handful of engineers writing and maintaining every Airflow DAG, while the business waits in the queue.

We took a different path. Instead of hand-writing pipelines, we generate them from metadata, a declarative description of what each data product needs, compiled straight into asset-scheduled Airflow 3 DAGs. Source dependencies, partitioning, trigger conditions, retries: all derived, not typed. That alone let us stand up hundreds of pipelines without hundreds of hand-maintained files.

06:00 - 06:30. Online
By Yunhao Qing
Track: Data & AI Applications
11/05/2026 6:00 AM 11/05/2026 6:30 AM America/Chicago AS26: Airflow Across Four Regions and a Fleet of Agents: How Notion Runs Its Data Platform

Notion runs in four AWS regions today, with more on the way, and every data pipeline has to respect where a customer’s data lives. Airflow on Astronomer is the control plane that holds this together: two Astro deployments, 500+ DAGs, and region-aware routing to EMR, EMR Serverless, Ray on Anyscale, Kafka, Snowflake, Databricks, and our vector database.

This talk covers the platform patterns that let a small team operate all of that: a generated cell and region model shared by every DAG, so new regions and retired cells propagate without touching pipeline code; DAG factories that fan one config out per region; and per-environment, per-region compute isolation, including the custom Anyscale operator we built and the cloud migrations we ran from DAG config alone.

We then zoom into one concrete workload: indexing Jira, Slack, Google Drive, GitHub, and a dozen other connectors for Notion AI. Kafka feeds Spark and Ray embedding jobs that write to region-local vector indexes. This workload shows both where our abstractions held up and where they broke down, including the per-region DAG file explosion we are still paying down.

Attendees will leave with practical patterns for designing region-aware Airflow platforms: how to model regions and cells, generate DAGs without duplicating business logic, isolate compute by environment and geography, and evolve infrastructure as regions are added, migrated, or retired.

Online

Notion runs in four AWS regions today, with more on the way, and every data pipeline has to respect where a customer’s data lives. Airflow on Astronomer is the control plane that holds this together: two Astro deployments, 500+ DAGs, and region-aware routing to EMR, EMR Serverless, Ray on Anyscale, Kafka, Snowflake, Databricks, and our vector database.

This talk covers the platform patterns that let a small team operate all of that: a generated cell and region model shared by every DAG, so new regions and retired cells propagate without touching pipeline code; DAG factories that fan one config out per region; and per-environment, per-region compute isolation, including the custom Anyscale operator we built and the cloud migrations we ran from DAG config alone.

06:30 - 07:00. Online
By Aleksandr Shirokov & Roman Khomenko
Track: Data & AI Applications
11/05/2026 6:30 AM 11/05/2026 7:00 AM America/Chicago AS26: Beyond Multi-Cluster Airflow: Operating GPU Workloads at Scale

At last year’s Airflow Summit, we shared how we built a multi-cluster orchestration layer on top of Apache Airflow to run ML workloads across multiple Kubernetes GPU clusters.

Once hundreds of ML engineers started running GPU pipelines in production, we discovered that orchestration alone is not enough. Operating multi-cluster GPU infrastructure introduces new challenges: controlling GPU allocation across teams, observing pipelines across clusters, and helping users run workloads efficiently without wasting expensive GPU resources.

In this talk, we’ll show how our Airflow platform evolved from a workflow orchestrator into an operational control plane for GPU infrastructure. We’ll cover custom scheduling strategies that dynamically route workloads across clusters using Airflow policies and resource awareness, integration with HAMI to improve GPU utilization, and AIOps workflows with KeepHQ that detect underutilized CPU, RAM, and GPU resources. We’ll also present powerful dashboards and AI-assisted tools that reduce Time2Market and simplify debugging while keeping infrastructure complexity hidden.

We’ll be happy to share how our platform continues to evolve with Apache Airflow.

Online

At last year’s Airflow Summit, we shared how we built a multi-cluster orchestration layer on top of Apache Airflow to run ML workloads across multiple Kubernetes GPU clusters.

Once hundreds of ML engineers started running GPU pipelines in production, we discovered that orchestration alone is not enough. Operating multi-cluster GPU infrastructure introduces new challenges: controlling GPU allocation across teams, observing pipelines across clusters, and helping users run workloads efficiently without wasting expensive GPU resources.

07:00 - 07:30. Online
By Christos Bisias
Track: Builder
11/05/2026 7:00 AM 11/05/2026 7:30 AM America/Chicago AS26: Cross-Deployment Monitoring and Dag Dependencies with the Kafka Event Producer Plugin

Organizations can have multiple Airflow instances running on separate hardware, each with its own database, and sometimes in different geographical regions. How can we monitor these instances and get updates on their state without polling their databases or APIs? Is there a way to coordinate Dags from different Airflow instances?

This session answers the above questions by presenting a newly released plugin in the Apache Kafka Provider, which publishes a message to a configured Kafka topic for every Dag run or task instance state change event.

It will cover:

  • what the plugin does
  • what the published messages look like
  • how to enable and configure it
  • how to consume the messages
  • some use cases
Online

Organizations can have multiple Airflow instances running on separate hardware, each with its own database, and sometimes in different geographical regions. How can we monitor these instances and get updates on their state without polling their databases or APIs? Is there a way to coordinate Dags from different Airflow instances?

This session answers the above questions by presenting a newly released plugin in the Apache Kafka Provider, which publishes a message to a configured Kafka topic for every Dag run or task instance state change event.

07:30 - 08:00. Online
By Ramzi Alashabi
Track: Data Strategy
11/05/2026 7:30 AM 11/05/2026 8:00 AM America/Chicago AS26: Scale Without Scaling the Team: Metadata-Driven Airflow the Business Can Own

Every data team hits the same wall: a handful of engineers writing and maintaining every Airflow DAG, while the business waits in the queue.

We took a different path. Instead of hand-writing pipelines, we generate them from metadata, a declarative description of what each data product needs, compiled straight into asset-scheduled Airflow 3 DAGs. Source dependencies, partitioning, trigger conditions, retries: all derived, not typed. That alone let us stand up hundreds of pipelines without hundreds of hand-maintained files.

But auto-generation only takes you so far. Real orchestration needs human judgment: this product should wait for that one; these steps run in a specific order. So we gave the business a canvas drag-and-drop dependencies on top of the generated graph that compiles back into the same metadata. Analysts compose their own reusable orchestration; engineers stay out of the critical path.

I’ll walk through: • How to compile metadata into an asset-scheduled Airflow DAG, what to derive automatically, and what to leave configurable. • How a visual, drag-and-drop dependency graph maps cleanly back to declarative orchestration, not throwaway clicks. • Where auto-generation ends and human-authored orchestration begins and how to let both live in one source of truth. • What this does to delivery speed when the people who understand the data can ship the pipeline.

You’ll leave with a blueprint for metadata-driven DAG generation plus business-owned orchestration a way to scale pipeline delivery without scaling your DAG-writing team.

Online

Every data team hits the same wall: a handful of engineers writing and maintaining every Airflow DAG, while the business waits in the queue.

We took a different path. Instead of hand-writing pipelines, we generate them from metadata, a declarative description of what each data product needs, compiled straight into asset-scheduled Airflow 3 DAGs. Source dependencies, partitioning, trigger conditions, retries: all derived, not typed. That alone let us stand up hundreds of pipelines without hundreds of hand-maintained files.

13:00 - 13:30. Online
By Piyush Maheshwari & Sameer Raj
Track: Builder
11/05/2026 1:00 PM 11/05/2026 1:30 PM America/Chicago AS26: Microservice-Grade Delivery and Release for Airflow DAGs

At Uber, preparing Airflow to take on workloads from Piper (our Airflow 1 fork operating at nearly one million daily task runs) requires rethinking both DAG delivery and release. Shipping a multi-GB monorepo artifact to isolated Kubernetes executors for every task is neither fast nor efficient.

We’ll share how dependency-aware slim bundles package only the required DAG code, first-party dependencies and generated artifacts into reproducible bundles. We’ll then cover per-DAG version pinning, which separates code distribution from activation and enables pre-production regression gates, controlled promotion and automated rollback.

Attendees will learn practical patterns and trade-offs for monorepo dependency resolution, bundle granularity, Kubernetes execution and safer DAG releases. We’ll also discuss AIP-109, our proposal to contribute DAG version pinning to Apache Airflow.

Online

At Uber, preparing Airflow to take on workloads from Piper (our Airflow 1 fork operating at nearly one million daily task runs) requires rethinking both DAG delivery and release. Shipping a multi-GB monorepo artifact to isolated Kubernetes executors for every task is neither fast nor efficient.

We’ll share how dependency-aware slim bundles package only the required DAG code, first-party dependencies and generated artifacts into reproducible bundles. We’ll then cover per-DAG version pinning, which separates code distribution from activation and enables pre-production regression gates, controlled promotion and automated rollback.

13:30 - 14:00. Online
By Satej Sahu
Track: Builder
11/05/2026 1:30 PM 11/05/2026 2:00 PM America/Chicago AS26: Adaptive DAGs: Dynamic Query Plan Rewriting and Runtime Task Re-sizing in Airflow

Traditional Airflow DAGs execute rigid, pre-determined computation graphs. Yet in production, data volume and distribution fluctuate wildly: a daily ingestion batch may process 50,000 rows on Monday but 50,000,000 on Black Friday. Statically sized tasks force engineers into a painful compromise: either permanently over-provision worker resources (blowing cloud budgets) or risk out-of-memory crashes, disk spills, and missed SLAs during unexpected volume spikes.

Drawing inspiration from database query optimizer research and recent papers on learned adaptive pipeline scheduling, this technical session demonstrates how to implement Adaptive DAGs in Apache Airflow. We show how to construct workflows that introspect upstream intermediate data statistics to dynamically adjust downstream execution plans, engine targets, and parallelism at runtime. Attendees will learn:

  1. Runtime Plan Profiling: Using lightweight metadata collectors and storage manifests (Iceberg/Delta metadata) to capture data volume, partition skew, and cardinality before heavy compute stages run.
  2. Dynamic Strategy Routing: Leveraging Airflow’s TaskFlow API and Dynamic Task Mapping (.expand()) to conditionally route jobs between lightweight single-node compute (DuckDB/Polars) for small datasets and distributed clusters (Spark/Ray) for large scale.
  3. Adaptive Partition Resizing: Dynamically generating downstream worker parallelism and memory allocations on the fly, eliminating OOM failures and compute waste.

Attendees will leave with production-ready code patterns to build self-tuning, resilient Airflow pipelines.

Online

Traditional Airflow DAGs execute rigid, pre-determined computation graphs. Yet in production, data volume and distribution fluctuate wildly: a daily ingestion batch may process 50,000 rows on Monday but 50,000,000 on Black Friday. Statically sized tasks force engineers into a painful compromise: either permanently over-provision worker resources (blowing cloud budgets) or risk out-of-memory crashes, disk spills, and missed SLAs during unexpected volume spikes.

Drawing inspiration from database query optimizer research and recent papers on learned adaptive pipeline scheduling, this technical session demonstrates how to implement Adaptive DAGs in Apache Airflow. We show how to construct workflows that introspect upstream intermediate data statistics to dynamically adjust downstream execution plans, engine targets, and parallelism at runtime. Attendees will learn:

14:00 - 14:30. Online
By Shrividya Hegde
Track: Builder
11/05/2026 2:00 PM 11/05/2026 2:30 PM America/Chicago AS26: Self-Healing Airflow: Pipelines That Fix Themselves

It’s 2 AM and a task just failed. Does someone get paged, or does the pipeline quietly diagnose the problem, wait the right amount of time, retry, and only escalate when it truly can’t recover on its own?

Most Airflow pipelines are babysat. This session shows how to make them self-healing instead. We’ll work up from the 80% fix , retries with exponential back off ,through failure and retry callbacks, SLA monitoring, automatic clearing of stuck and zombie tasks, and circuit breakers that stop hammering a broken downstream system. You’ll see how to route alerts to Slack or PagerDuty so humans are interrupted only when it matters, and how to keep secrets out of your DAGs while doing it.

Coming from a testing background, I’ll frame this the way an SRE would: design for failure first. You’ll leave with concrete patterns and a template you can drop into your own DAGs on Monday.

Online

It’s 2 AM and a task just failed. Does someone get paged, or does the pipeline quietly diagnose the problem, wait the right amount of time, retry, and only escalate when it truly can’t recover on its own?

Most Airflow pipelines are babysat. This session shows how to make them self-healing instead. We’ll work up from the 80% fix , retries with exponential back off ,through failure and retry callbacks, SLA monitoring, automatic clearing of stuck and zombie tasks, and circuit breakers that stop hammering a broken downstream system. You’ll see how to route alerts to Slack or PagerDuty so humans are interrupted only when it matters, and how to keep secrets out of your DAGs while doing it.

14:30 - 15:00. Online
Track: Sponsored
11/05/2026 2:30 PM 11/05/2026 3:00 PM America/Chicago AS26: IBM Sponsored

TBD

Online
15:00 - 15:30. Online
By Vinod Jayendra, Sean Bjurstrom & Abdul Majid Mohammed
Track: Data Strategy
11/05/2026 3:00 PM 11/05/2026 3:30 PM America/Chicago AS26: Ten Teams, One Airflow Cluster, No Guardrails. Until Now.

Your Airflow environment started as one team’s orchestrator. Now three teams share it, each wants their own connections, variables, and execution roles. Your security review flagged that every DAG runs with the same permissions. Sound familiar?

We’ll tackle multi-tenancy head-on. Most Airflow deployments have a single execution context, one DAGs folder, and limited native isolation between teams. Here are the patterns that work in production.

Runtime resource isolation: per-task credential scoping so each team’s DAGs only access their own data stores and APIs. DAG-level access control with automated tag-based RBAC that syncs identity attributes to Airflow roles on a schedule, so new DAGs inherit the right permissions without manual work. And the deployment tradeoff: single shared instance with RBAC guardrails vs. separate environments per team.

We’ll show how AI agents speed up multi-team operations. Using Airflow 3’s task.agent decorator, we built an automated onboarding workflow that scans a team’s DAG repo, infers required connections and permissions, generates scoped access policies, and configures RBAC. All as auditable, retryable Airflow tasks instead of manual runbooks.

We’ll preview Airflow 3.2’s experimental multi-team support: separate DAGs, connections, variables, pools, and executors per team in a single instance.

You’ll leave with a blueprint for operating Airflow as an internal platform. Self-service onboarding, cost allocation per team, and security guardrails that don’t slow developers down.

Online

Your Airflow environment started as one team’s orchestrator. Now three teams share it, each wants their own connections, variables, and execution roles. Your security review flagged that every DAG runs with the same permissions. Sound familiar?

We’ll tackle multi-tenancy head-on. Most Airflow deployments have a single execution context, one DAGs folder, and limited native isolation between teams. Here are the patterns that work in production.

Runtime resource isolation: per-task credential scoping so each team’s DAGs only access their own data stores and APIs. DAG-level access control with automated tag-based RBAC that syncs identity attributes to Airflow roles on a schedule, so new DAGs inherit the right permissions without manual work. And the deployment tradeoff: single shared instance with RBAC guardrails vs. separate environments per team.