LDX3 New York: every AI conversation came back to proof
We spent September 15 and 16 at LDX3, LeadDev’s New York conference. Returning attendees told us the agenda usually covers a lot of ground. This year, it felt like people could only talk about one thing: AI and its impact.
And almost every conversation we had came back to one question: can you prove AI moved the business? Here’s what we took home from New York.
The question leaders kept asking
In the DirectorPlus roundtables and senior-leader talks, the conversation kept turning to impact. We set out to deliver this initiative this quarter. Did it happen? Are we meeting the business goals we committed to?
More AI-generated code can make that question much harder to answer, because code activity goes up either way.
Interestingly, none of the leaders we spoke with framed engineering measurement as a way to justify headcount cuts. That matches how we think about these signals. Nobody should use them to rank individuals or set compensation.
Team changes did come up. One leader’s org had ended contracts with offshore developers and was working out how to keep quality consistent in the future with a smaller team and the same amount (or more) of code coming through.
No metric stands on its own
Jules: I heard a fair amount of DORA-is-dead talk, which I don’t really agree with, but I do think DORA needs to evolve. The framework is about 12 years old, and software development has changed a lot since then. The way we measure software delivery has to change for AI-accelerated development.
My own takeaway: no metric means much by itself. You need a more comprehensive picture to really understand what’s happening in your code.
For example, shorter lead time and higher deployment frequency are good news only if your change failure rate holds steady or drops. If you focus on those first two numbers alone, you can end up with something we’re calling velocity theater: dashboards that look like progress while the code tells a different story.
Failure rates can hide extra work too. In one customer example Aaron shared recently, deployments rose from 1,297 to 2,007 across two quarters, and reverts rose from 11 to 18. The change failure rate barely moved. But that’s still seven more broken changes someone had to find, roll back, and fix.
The SPACE framework came up too. It comes from the same research lineage as DORA and makes the same case: you can’t measure engineering work with one number. The dimension I’d watch most right now is collaboration, which covers code review, because that’s where AI adds the most obvious strain.
Your heaviest reviewer often looks like your least productive engineer, because every hour of review is an hour they didn’t spend shipping their own features. As AI pushes more code through the pipeline, that load lands on the same small group.
That’s why I want Flux to explain the number. Say lead time jumps 42%. Flux can trace it to one contested PR that sat in review for 12 days and collected 43 comments, and now you know where to look on Monday morning. That’s exactly what our lead time analysis is built to show.
Software factories
Aaron: Software factories came up in a lot of conversations and talks at LDX3. By software factory, I mean the newer definition, a staged delivery system where you put in a definition of the software you want (anything from a full spec to a rough description) and software comes out. Humans still own the approval gates. We talked with one team whose system one-shots low-risk PRs.
In my conversations with other leaders, I met folks building their own factory or exploring third-party ones, but I didn’t meet anyone running a whole shop that way. There was a stretch when people talked as if full automation was inevitable, and coming tomorrow. In most industries, going all-in is close to impossible, so leaders are treating the factory approach as one piece of the team and working out where it fits. Right now, the range seems to be from 8%–30% of software workloads. The teams furthest along start with front-end changes, bugfixes, and low-risk improvements, and keep human eyes on core business logic the longest.
So what should you watch out for when a team starts routing low-risk PRs through a factory? I’d start with change size, time to first review, and PR cycle time, measured against the team’s own previous 12 weeks. At one customer, the median wait for a first review ran 6.4 hours across a quarter and 25.6 hours during a single spike week. We’re calling code that sits in a queue review debt, and it compounds pretty quickly as AI-assisted coding adoption grows.
Along with review bottlenecks, you need to consider code quality and risk. On my own team, agentic development brought a 48% jump in throughput over two quarters, followed by a 16% rise in stability issues. That 16% is what quality drift looks like, and it’s the thing I’d hate for a leader to miss. AI tends to create plausible code, polished enough to pass review. The problems in that code typically surface across many changes over weeks, a layer no single PR review can see. By the time you catch it, it’s usually an on-call alert traced back to a hastily-reviewed dependency. And there’s a decent chance a customer noticed first.
What’s next for engineering leaders
Proving AI impact is why we expanded the Flux platform recently. Now you can see whether your AI investment is paying off and find (and address) areas where it’s not working as well as you’d hoped. On October 20 at 11 AM ET, we’re hosting a webinar on the five blind spots behind unproven AI investments. Register now to join us and bring your questions. If you can’t make it or you’d like to see what Flux shows about your own codebase before then, request a demo and we can walk through your specific questions.