The AI bet is real; can you prove it’s paying off?
Every engineering leader I talk to right now is a variation of the same conversation. The board wants to know if the AI investment is paying off. The CFO wants a number. And most leaders don’t have concrete, defensible answers.
They adopted AI coding tools and shipped a lot of code. But their legacy approaches for management at best can’t scale for the agentic era, and at worse can’t deliver key metrics they now require. Solving for this was the founding premise of Flux: looking at code instead of tickets.
Tickets tell you what someone planned to do. Story points tell you how big a team thought a task was before anyone wrote a line of code. Adoption percentages tell you how many people opened an AI tool this week. None of that tells you whether the work actually shipped, whether it’s any good, or whether the money your org invested in those tools is truly delivering AI-accelerated development.
That gap is why we built what we’re announcing today, and it’s the same gap we wrote about back in June: AI solved the code creation problem. Visibility is where we win or lose next.
Activity growth isn’t the same as outcome delivery
DORA’s 2025 research found that 90% of software professionals now use AI daily, up from 76% the year before. That’s about as close to universal adoption as this industry gets. DORA’s own conclusion: AI amplifies what’s already there. It makes strong organizations stronger and struggling ones worse, faster than either of them may have expected.
Your team’s already using AI (we’re pretty sure). The real question is whether that use is turning into something you can put in front of your board. When we surveyed engineering leaders on what AI-generated code actually looks like in production, nearly a third told us they can’t keep up with what’s changing in the codebase week to week.
In working with Engineering leaders on this, the same five gaps come up consistently independent of industry or company size. We started calling them blind spots, because that’s what they are: places where leaders know activity is happening but can’t tell if it really matters.
Five blind spots undermining business outcomes
The five blind spots are velocity theater, review debt, hidden work, quality drift, and unproven spend.
Each one of these blind spots is distinct. An organization can be doing great on one and be exposed on another at the same time, which is exactly why a single blended health score isn’t useful.
Velocity theater might be the easiest blind spot to fall victim to. Commit and PR volume goes up, and it feels like progress. But volume isn’t the same thing as shipped value, and I’ve seen teams celebrate a busy quarter with no substantial product delivery to show for it.
Review debt is what happens when AI writes code faster than people can responsibly review it. We’ve all seen this happen as code volumes increased. Sign-off starts moving faster than careful code review could possibly allow, and the bottleneck still grows. Often, that burden gets handed to senior engineers, who are expected (somehow) to catch everything everyone else missed.
Hidden work is the refactor that never got a ticket, the architecture change that ate a week of engineering time nobody planned for. It’s real effort that’s typically invisible to tickets because it wasn’t planned.
Quality drift is the kind of failure that’s hard to detect until an incident forces it into view. A retry loop with no backoff ships clean, passes code review, and works fine in staging. Three weeks later it’s hammering a downstream service during a traffic spike, and nobody connects it back to the PR that introduced it.
Unproven spend is the hardest conversation engineering leaders are having with CFOs lately: money going into AI tools, seats, token spend, and initiatives with no clear way to show what it delivered.
What we built to expose these blind spots
Today we’re releasing five new capabilities to address each blind spot: verified velocity, trusted review, auditable work, continuous quality, and defensible spend.
All five capabilities are possible because Flux analyzes the code itself, the commits, the pull requests, the history of what got built. That’s the same code analysis Flux already runs, so there’s no new instrumentation to bolt on and no process to change. You get the answers you need by analyzing your org’s code.
“I’ve spent the last year trying to prove where our engineering budget actually goes,” said Gunter Ollmann, CTO at Cobalt, the leader of human-led, AI-powered offensive security. “Now I have evidence I can hand to finance for R&D credit substantiation, and to the board when they ask if these bets are paying off.”
That’s the point: evidence you can actually stand behind.
Where to start
These capabilities are available now. If you’re tired of guessing at the answer your board is asking for, join Aaron and Jules on October 20 at 1 PM ET/10 AM PT for a webinar. We’ll walk through all five blind spots and what resolving each one looks like in practice, using evidence straight from the code.
If you’d rather see it against your own codebase before the webinar, you can also request a demo.