AI control becomes a new enterprise resilience priority

Artificial intelligence is moving deeper into enterprise infrastructure, bringing a broader risk horizon with it. As AI progresses from generating text in isolated environments toward initiating transactions, influencing decisions, and contributing to production code, the consequences of an error may extend across interconnected systems.

Splunk’s 2026 research, released by Cisco in partnership with Oxford Economics, estimates that unplanned downtime has reached an annual $600 billion impact across Global 2000 companies, with average downtime costs reaching $15,000 per minute.

Recent cloud disruptions illustrate how the way a change is introduced can affect the scale of an incident. During a significant Google Cloud outage in June 2025, a new feature was activated globally rather than through a gradual rollout, according to Google’s subsequent incident report. When the change triggered failures, the broad deployment meant the impact was immediate across multiple regions and services, including third-party platforms dependent on Google Cloud infrastructure.

Google later identified progressive rollouts as one of the safeguards that should have been used. The incident provides a useful example of a broader resilience principle: limiting initial exposure to a change can reduce its potential blast radius and give teams an opportunity to intervene before a problem propagates across a production environment.

That exposure is also contributing to a broader governance discussion. The same Splunk research found that 68% of technology leaders surveyed expressed concern about unpredictable AI-agent behavior, while every technology leader surveyed reported experiencing some form of AI-related downtime. The Google example is particularly relevant in the context of how quickly AI is changing software development.

Google CEO Sundar Pichai recently stated in Alphabet’s annual report, “Today, nearly 75% of all new code at Google is AI-generated and approved by engineers, up from 50% last fall.” As AI accelerates the volume and pace of software changes, incidents like Google’s illustrate why the controls governing how those changes reach production may become increasingly important.

These findings suggest that observability alone may address only part of the operational question. Monitoring can identify unusual behavior or deteriorating performance, while an intervention mechanism can give teams a means of stopping a defined process when predetermined conditions are reached.

The question, therefore, is expanding from how organizations observe AI systems to how they maintain operational control over them. As automated software becomes more capable and more deeply connected to production environments, mechanisms for pausing, disabling, or reversing AI-enabled functions may become relevant to resilience planning alongside monitoring and incident response.

The emerging consideration is less about whether enterprises can introduce AI and more about how they can control, contain, and reverse AI-enabled behavior once those systems become part of everyday operations.

Egil Østhus, CEO of Unleash, places this question within the evolution of software release management. Unleash is an open-source feature management platform that separates code deployment from the decision to activate or change software behavior in production, allowing organizations to control applications, services, and AI-enabled capabilities at runtime.

That distinction becomes particularly relevant as AI-assisted development increases the volume of code changes moving toward production. “Race car brakes are about speeding you up, not slowing you down,” says Østhus, “Teams can move faster when they know they can reverse a problematic change instantly, rather than waiting for a fix to make its way through production.”

Within that model, Unleash provides a control layer for managing how software behaves after it has reached production. Rather than treating deployment as the final point of control, organizations can limit the exposure of a new capability and intervene immediately if its behavior falls outside acceptable parameters.

This is precisely the safeguard Google identified as missing following its 2025 outage. In its postmortem, Google stated that the failed code path “did not have appropriate error handling nor was it feature flag protected,” adding: “If this had been flag protected, the issue would have been caught in staging.”

For AI-enabled applications, applying this kind of control at runtime can reduce the blast radius of problematic behavior and provide an immediate path to containment or fallback without waiting for another deployment.

“A team could, for instance, release an AI function to a limited user group, monitor its behavior, expand its availability when predefined conditions are met, or deactivate it if an established threshold is exceeded,” Østhus explains. “This separates the decision to run an AI capability from the permanence of the code containing it.”

That separation also turns an abstract governance requirement into an operational control. Under the EU AI Act, high-risk AI systems must provide appropriate human oversight, including the ability to intervene in their operation or interrupt them through a “stop” button or similar procedure that brings the system to a safe state.

For financial institutions, DORA already adds another operational imperative: major ICT incidents must be initially reported within four hours of being classified as major and no later than 24 hours after detection. In that environment, the question is not simply whether an organization has a policy saying a human can intervene.

It is whether that person has a technical mechanism to stop problematic AI behavior immediately, preserve the underlying service where possible, and create an auditable record of what was changed and when.

For organizations running AI in business-critical systems, the control mechanism itself also needs to be resilient. Unleash can run the decision logic behind an AI kill switch within a customer’s own environment, close to the applications and services it controls. That means intervention does not depend on waiting for a new deployment to propagate or on maintaining connectivity to an external control service.

If an AI capability needs to be stopped, restricted, or redirected to an established fallback, the change can take effect immediately where the software is running. This makes runtime control part of the resilience architecture itself, rather than another external dependency during an incident.

Østhus’s broader argument connects this capability to the changing economics of software development. AI tools can make code generation faster and easier, potentially increasing the number of changes developers introduce into production environments.

That acceleration can create a corresponding need for mechanisms capable of handling mistakes at comparable speed. From this perspective, the role of feature management is evolving alongside the software it controls. The objective is to establish how specific capabilities can be introduced, restricted, monitored, and withdrawn as conditions change.

As AI becomes more embedded in business-critical systems, rapid intervention may increasingly become part of operational resilience, corporate governance, and regulatory preparedness. The scale of downtime costs, the reach of cloud infrastructure, and concerns surrounding autonomous AI behavior all suggest that control mechanisms may warrant consideration alongside the systems themselves.

Original source AI control becomes a new enterprise resilience priority

Back to home