Building a RAG prototype is easy, but keeping millions of enterprise documents continually parsed, chunked, embedded, and indexed in vector databases is hard. Document updates, schema changes, embedding model versioning, and rate limits on embedding APIs create unique engineering hurdles.
In this talk, we present a robust, scalable architecture combining Apache Airflow and Apache Beam on Dataflow for production RAG pipelines. We demonstrate how Airflow schedules and triggers incremental document ingestion, how Beam parallelizes document chunking and embedding generation via batch inference transforms, and how the pipeline safely handles upserts, and embedding version migrations. Attendees will gain a blueprint for building highly scalable, cost-effective vector ingestion pipelines.
Ashwin Sampathkumar
Engineering Manager