As AI becomes embedded into services across the enterprise, a new risk is emerging. João Freitas at PagerDuty explains that systems that appear to be working are quietly making the wrong decisions, and those errors compound

AI is no longer confined to the cutting edge. The conversation is not whether to use the technology, but how to use it.
The approach many teams have taken is to force AI into operational models that weren’t built with AI in mind. Legacy workflows, fragmented ownership, and manual processes were designed for a time before large language models (LLMs) became prevalent in the industry. Simply layering AI on top without reworking these foundations pushes too much of the burden on humans to triage and route alerts, run diagnostics, and perform remediation, with the outcomes resulting in manual toil and alert fatigue.
The risk of retrofitting AI into digital operations is that it accelerates bad processes rather than fixes them. Maturity will require a new approach to labour, split between humans and machines, and incident management is one area where this split is most tangible, with machines potentially doing most of the work.
When incidents strike, the blast radius extends beyond customers. Subject matter experts (SMEs) and incident responders are pulled out of bed or dragged away from important projects. This firefighting results in mental fatigue as teams are forced into emergency mode. When asked where outages hurt most, 42% of business and IT leaders cite developer morale and burnout as one of the greatest impact areas.
AI-first operations mitigates the burnout issue by reducing the number of people needed to resolve incidents. In a human + AI model, AI can handle low-risk or well-understood incidents, which keeps a human in the loop while ensuring they’re only interrupted for high-severity or novel situations. Even then, AI agents accelerate context gathering, diagnosis, mitigation and decision making.
An SRE (site reliability engineering) agent is a good example of this model in practice, where AI can take the lead. When an incident triggers, the AI agent can gather context in seconds by pulling logs, metrics, changes, and incident history. It can summarise what has happened, identify root causes and recommend next steps. That’s a lot of time and effort saved for human responders before they’ve even joined the call.
The journey from human + AI to autonomous operations
Before humans can fully enable AI to take control of the incident response, they need to trust it. An incident management model that combines AI and humans can be built around three tiers, based on how familiar, repeatable, and risky the incident is. A good starting point is for AI to take the lead on routine work, with humans retaining decision accountability where risk, complexity, or novelty demand it.
Tier One: Well-understood incidents. AI agents can autonomously resolve common issues with clear patterns, known fixes, and low risk within agreed guardrails. Examples include alert enrichment, noise reduction, initial triage, context gathering, pattern matching against previous incidents, routing recommendations, and approved remediation.
Tier Two: Familiar, but less certain, incidents. These incidents are issues the organisation has seen before, but the cause or best response may not be immediately clear. In these cases, AI can take the lead by gathering context, correlating data, identifying likely causes, and recommending next steps. A human should remain in the loop to validate findings and evidence and approve actions.
Tier Three: Novel, complex or high-impact incidents. Human-led remediation is still important, particularly where incidents carry customer, commercial, security, or reputational risk, or demand the kind of judgement, ethical reasoning, or empathy that AI can’t replicate. AI has an important supporting role, helping teams collect data, summarise what has happened, document decisions, and manage routine communications.
Process, trust and mindset
Delivering successful AI-first operations is more than just deploying technology. AI agents only fulfil their potential when enterprises define the processes, permissions, and trust levels that allow them to act safely and usefully during incidents. If an AI agent is expected to act as a first responder, the incident management workflow needs to assign it a clear role from the start. This includes defining which actions the agent can take independently, which require human approval, and which must be escalated immediately.
Governance underpins all of this. Access controls, audit trails, and clear escalation paths are essential to making the human + AI model safe and accountable. Those parameters should vary depending on incident severity, customer impact, confidence level, service criticality, and potential blast radius.
Trust is just as important. SREs, engineers, service owners, and executives all need to understand how the human + AI operating model works, where accountability sits, and why the AI agent is being introduced.
Technology alone won’t shift behaviour. Culture and mindset need to change too. The organisations that progress fastest will be those willing to test the model in controlled, lower-risk scenarios, learn from the results, and gradually expand the agent’s responsibilities as trust grows.
Adapting to an AI-first culture
To get started, teams should look for scenarios that are well-understood and repeatable. Routine incident triage, system health checks, and compliance monitoring are all ripe for automation. Teams must start small, build confidence, and demonstrate success early on, as this makes it easier to get buy-in for broader deployment.
AI agents need access to the right operational context, including observability data, incident history, service topology, runbooks, knowledge bases, and escalation policies. Then, teams can move gradually through trust levels: observe, recommend, act with approval, then act autonomously in defined scenarios. Organisations should invest in training engineers so that they’re confident in supervising, tuning, and improving AI-powered workflows.
Measurement matters too. Rather than treating AI agents as a black box, it’s important to track the impact they’re having through metrics such as time to triage, after-hours interruptions, escalation volume, and developer time returned. If the numbers aren’t moving, that’s a signal to re-tune, not retreat.
What separates the leaders is how they redesign work around the strengths of AI agents and humans, not how quickly they deploy. The result is a more resilient, intelligent, and sustainable operating model.
João Freitas is Chief AI Officer at PagerDuty
Main image courtesy of iStockPhoto.com and da-kuk
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543