AI can now write SQL, generate Python, review pull requests, and build entire applications with a few well-crafted prompts. Yet despite all of that progress, data engineering still feels oddly resistant to the full AI revolution. Pipelines break. Data models drift. Production data needs validation. Someone still has to close the loop.
In this episode of the Data Engineering Central Podcast, I sat down with Hugo Lu, founder and CEO of Orchestra, to talk about what “agentic data engineering” actually means beyond the buzzwords. Rather than another conversation about AI replacing engineers, we dug into the infrastructure that’s still missing before autonomous data platforms become reality.
Hugo shares his unlikely path into data engineering, from investment banking to helping build data systems at Juul, before eventually founding Orchestra.
What started as an effort to simplify orchestration has evolved into a broader vision where AI agents don’t just generate code, but can safely execute work, observe the results, validate changes, and iteratively improve pipelines inside secure environments.
That ability to observe outcomes, what many are calling “closing the loop,” may be the missing ingredient preventing today’s coding agents from becoming truly autonomous.
We also explore why data engineering has not experienced the same AI disruption as traditional software engineering. While AI can produce application code remarkably well, production data systems introduce a completely different set of problems. Branching production data, validating schema changes, testing transformations against realistic datasets, and understanding business semantics all remain difficult challenges that require much more than simply generating code.
The conversation naturally turns toward the future of the profession itself. We discuss whether junior engineers are losing the traditional apprenticeship path, how senior engineers are shifting from writing code to reviewing AI-generated work, and why decades of experience debugging production systems may actually become even more valuable in an AI-first world. Rather than eliminating engineering expertise, AI may simply be changing where that expertise is applied.
Finally, we dive into the rapidly changing data infrastructure landscape. From DuckDB and Polars to serverless compute, Iceberg, semantic layers, AI-native orchestration, and the growing concern over rising LLM token costs, we discuss where the industry appears to be heading and which trends are likely to stick long after today’s hype cycle fades.
If you’ve been wondering what comes after Orchestration, how AI agents will actually manage production data pipelines, or whether data engineering itself is about to undergo its biggest transformation in a decade, I think you’ll enjoy this conversation.
In this episode we discuss
What “agentic data engineering” actually means.
Why AI still struggles to fully automate data engineering.
The importance of closing the loop with production feedback.
Why orchestration may become the operating system for AI agents.
The changing role of data engineers in an AI-first world.
Why junior engineers face a very different career path than previous generations.
DuckDB, Polars, serverless data platforms, and where modern infrastructure is heading.
Whether today’s dependence on proprietary LLMs will create tomorrow’s vendor lock-in.
How Orchestra is building infrastructure for AI-native data platforms.











