Brenton O’Callaghan at Avantra describes how to use AI to automate root-cause analysis and reduce time-to-resolution

Spare a thought for enterprise ops engineers. Never easy, their jobs are only getting harder by the day. In a world where IT systems are growing in complexity almost on the hour, ops teams are under huge amounts of pressure to resolve incidents at speed and without any additional resources. With the strain beginning to show, it’s time for a new approach to diagnosing the root causes of system faults.
Incident response in the age of technology sprawl
Ongoing digital transformation is making core enterprise processes across finance, HR, supply chain, and procurement more complex than ever. Traditional and AI systems alike are proliferating, and increasingly span multiple clouds and on-premise data centres. Vendor sprawl is a real and present danger.
The numbers speak for themselves. Eighty-nine per cent of businesses operate a multi-cloud strategy, while the average large enterprise now uses approximately 897 disconnected applications across various business units. If something goes wrong, it falls on ops engineers to identify the fault, diagnose its cause, and find a fix – far from simple in an environment comprising this multiplicity of platforms and applications. All too often, diagnosing a fault can be like searching for the proverbial needle in a haystack.
However, delays to fault resolution can quickly show up in everything from lost productivity, disruption to the customer experience, and brand damage. Not all faults are equal, of course, but the worst can have truly catastrophic effects. According to a study by PagerDuty, for instance, some organisations report losing more than $1 million per hour during unplanned disruptions. This is nothing short of a disaster, and is clearly a boardroom issue.
No wonder ops teams are feeling under pressure. Indeed, the same study reveals that 42% of business leaders and IT decision-makers cite burnout as a major consequence of IT disruption.
Where visibility tooling falls short
In recent years, organisations have invested heavily in advanced visibility, detection, and alerting tooling. The global observability tools and platforms market size was estimated at $2.71 billion in 2023 and is projected to reach $5.4 billion by 2030, growing at a CAGR of 10.7% from 2024 to 2030.
Observability systems are great at automatically determining the state of a given system based on its near-real-time inputs and outputs. However, visibility only addresses part of the problem. Detection without diagnosis just moves the bottleneck further down the chain. As technology estates grow in complexity, the diagnostic burden compounds faster than teams can address it through headcount alone.
A recent study from Monte Carlo shines a spotlight on this problem. According to its analysis, engineers spend an average of 19 hours resolving a single incident, with the majority of that time spent on diagnosis alone. In some cases, the cause is never found. Research from Splunk found that only 38% of technology executives report consistently identifying the root cause of a downtime incident. If you can’t diagnose the fault, what’s to stop it from happening again?
Automating root-cause analysis
The industry’s focus on detection and alerting has obscured the true bottleneck. The real question is what it takes to close the gap between an alert firing and an engineer knowing exactly what to do. It’s here that modern AIOps can really come into its own.
The next stage of operations optimisation will laser in on root-cause analysis and fault diagnosis. A new generation of AI-powered intelligence engines are emerging that promise to automatically investigate incidents and present a structured diagnosis the moment an issue is detected.
These tools are able to instantly correlate logs, system data, and alerts across multiple systems automatically. Rather than waiting hours or days diagnosing an issue, ops teams have the diagnosis served up to them instantly on a platter, with the AI also recommending next steps and practical guidance.
Although still at the cutting edge of technology, early deployments of AI-assisted triage and diagnosis are impressive, with users reporting up to a 60% reduction in mean time to insight and resolution, a 2-3x increase in automated tasks, and 10-20 hours in monthly time savings per senior engineer. That’s time that can now be spent on supporting innovation and other revenue-generating activities.
From downtime to delivery
With ops teams being asked to do more with less, managing systemic complexity and leveraging automation to the absolute max are essential. By going beyond detection and response using AI-powered intelligence, enterprise ops teams can eliminate the hours between detection and action entirely, freeing up resources while reducing the time to resolution. That means less stress for the team, and less downtime for the business.
Organisations that embrace AI-powered diagnosis will clearly benefit from a significant competitive advantage. The competitive advantage is clear: the only question is how long you can afford to wait.
Brenton O’Callaghan is CPO at Avantra
Main image courtesy of iStockPhoto.com and Parradee Kietsirikul


© 2025, Lyonsdown Limited. Business Reporter® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543