Are AI coders more trouble than they’re worth?

The majority of UK and US software development teams polled by independent research firm Coleman Parkes say artificial intelligence (AI) agents struggle to find issues in complex code. 

The poll of 300 software developers and engineering leads, published in the Overcoming the limitations of coding agents in complex software systems report from Undo, found that 93% of software development teams have experienced AI hallucinations, leading to incorrect diagnosis of problems in code. Over half (55%) say agents introduce incorrect code too frequently, creating rework that delays delivery cycles. In fact, 94% admit they are losing productivity due to having to analyse AI-generated code at least once a month.

Undo, which provides tools for debugging code, found that developers tend to spend twice as long debugging code as they do writing it, averaging 16.9 hours a week because software engineers are unable to keep up with the coding agents they use. The poll found that more than a third (35%) of AI-generated code reaches production before software development teams fully understand what the code actually does.

This leads to coding errors being introduced in production systems. 

While AI coding agents can potentially dramatically reduce the cost of creating code, the poll from Undo suggests that bottlenecks have been pushed downstream, which means engineers are now struggling to understand the code being produced, leading to them having to spend more time on debugging and incident investigations.

According to Undo, if the promised productivity gains of using AI coding are to be realised, AI must be applied to the software delivery cycle, not just code generation.

Undo warned that the difficulty of understanding why an application behaves the way it does, coupled with AI’s tendency to give confident answers from incomplete evidence, makes it prone to hallucinate the cause of bugs and instabilities. This means engineers still need to step in to solve these downstream problems, and this is where they are now spending most of their time.

As the authors of the report point out, relying on human effort to lead AI to fix every bug or figure out the cause of unexpected behaviour is unsustainable given the pace at which agents are generating code. Undo said this means AI agents need to be able to gather evidence and data to reach accurate conclusions independently.

“When code is obviously broken, the cause is usually easy to find,” said Greg Law, founder and CEO of Undo. “Where engineers struggle is with code that’s almost – but not quite – right. Those are the times they lose days trying to unravel what went wrong and why.

“Their challenge is that while agents are great at writing reams of code quickly, they’re less capable at debugging it. The result is engineers are being buried in an avalanche of code that’s well beyond human capacity to debug. That’s why we have to give them a way to make AI better at debugging, by feeding agents with the rich context of what code actually does at runtime.”

Original source Are AI coders more trouble than they’re worth?

Back to home