News

Introducing AI into ZOIS Platforms stream

07/11/2025

This article explores how AI-related topics fit into the ZOIS Platform Workstream. At the Vienna event, speakers offered two contrasting complementary perspectives: Stephan Kasulke described how Agentic AI could enhance resilience in Zero Outage infrastructures, while Christoff Schacher warned about the new risks that come with intelligent systems. These views reflect the dichotomy of AI: it can strengthen resilience and performance, yet also introduce new vulnerabilities like unpredictable model behavior, data bias, and the challenge of maintaining consistent reliability. AI doesn’t just make systems smarter; it makes keeping them stable more complex.

To understand where AI fits into the picture, we need to take a look at the foundation of the ZOIS layered model, which divides operations into three horizontal layers: Basic IT Infrastructure (data centers, cabling, storage, and computing capacity), IT Systems and Services (operating systems, infrastructure controllers, and services) and Business Applications (functionality supporting business processes). There are also vertical dimensions defining whether an infrastructure is dedicated to one client, shared among a few, or public, serving many users.

The Platform Workstream expands the layered model, describing the complexity and interaction of the two fundamental layers, IT Infrastructure and IT Systems. In the three layer ZOIS diagram the Platform Workstream corresponds to these two layers and explains the building blocks and relationships between them, defining architectural policies that ensure resilience, security, availability, and performance. It explains how infrastructure and systems interact to support reliable service delivery.

When it comes to AI, most activity takes place in the IT Systems and Services layer. But while AI depends on hardware, such as GPUs or data storage performance, the artificial intelligence ‘happens’ in the middle layer. In practice, this is also where the biggest testing and reliability challenges emerge because the combination of AI logic and complex distributed systems often leads to unexpected behaviors that surface level test design might overlook.

To explain why AI fits there, we need to start with its ever-changing definition. Decades ago, systems that could play chess or give expert advice were seen as examples of AI. Today, most practitioners would no longer call those as such. They are just examples of what we have mastered. The definition of AI evolves as our mastery grows, and understanding this shift is essential as we integrate AI into the Zero Outage framework.

AI spans a wide range of technologies, from symbolic reasoning to sub-symbolic learning such as supervised, unsupervised, and reinforcement learning. Yet today, when most people mention AI, they usually mean transformer-based deep learning models or LLMs behind systems like ChatGPT, Gemini, Llama, or Grok.

It’s also crucial to remember that AI is not your friend, not your enemy, not your teammate, and not your partner. While developing ZOIS, we should make a deliberate effort to avoid misleading anthropomorphic language that make AI sound human-like.

Instead, we apply a simple rule: AI is software. Not every software is AI, but any AI system is software. To test this idea, we can perform a simple substitution: take any sentence that includes the word AI and replace it with software. Then ask yourself if the sentence still makes sense.

Here’s an example that Michael Bolton once gave. A McKinsey study claims that “software developers can complete coding tasks up to twice as fast with GenAI.” If we replace AI with software, the statement becomes “software developers can complete coding tasks up to twice as fast with software.” It suddenly sounds a bit less reasonable. But replacing ‘AI’ in a phrase “AI isn’t the enemy, it’s your personal growth partner” gives us “Software isn’t the enemy, it’s your personal growth partner”, revealing the metaphor’s emptiness. This exercise is a practical way to keep our language clear and grounded.

Software itself combines instructions and data, both of which are forms of information. Some data is human-created, it’s natural, manual and attended. Other data is generated by software and it’s synthetic, artificial or automated. Recognizing this distinction helps us reason about how AI-enabled systems produce and consume information.

This brings us to the data volume dichotomy. Humans operate within a narrow bandwidth, because the amount of information we can process at once is limited. The data we can handle takes the form of artifacts or snippets. Those are small, interpretable fragments of information. Machines, on the other hand, work on an entirely different scale, processing datasets or corpora that far exceed human capacity.

This distinction matters because intelligence, whether human or artificial, relates to how entities acquire, process, and apply knowledge and skills. Humans have intelligence, but there’s no precise, agreed-upon definition of it. We interact with reality through skills and store knowledge as ideas. We can generate and interpret small artifacts like texts, designs or code snippets, all of which serve as proxies for our intelligence.

Software works differently. It exists as data, or code and instructions, and can process and produce synthetic artifacts derived from vast datasets. Its interaction with reality happens through data flows and control loops, such as sensors and actuators. Even something as simple as a screen display is part of that interaction, because its data represents reality in digital form. This describes running software, but there is also software that is not running. For example, when data and code are stored on a disk or software in a dormant state.

Software interacts with reality through data. For example, showing an image or text on a screen. Artificial intelligence is software capable of acquiring, processing, and applying knowledge and skills. Of course, knowledge and skills are human concepts, but in software we can still find representations of these ideas. In this sense, skills refer to the software’s ability to interact with the external world, while knowledge represents the internal structures or the data and the instructions contained within the system.

One major difference between humans and software is scale. Software, unlike humans, can process entire datasets and corpora. It can read them, generate them, and use them to produce synthetic artifacts. And when it comes to judging artificial intelligence, the only real way we can do that is by examining these synthetic artifacts. They are our proxies for understanding what AI does and how it behaves. These short reports are synthetic artifacts or snippets that serve as our window into the system’s behavior.

Here is an example. When we give ChatGPT a prompt, an enormous factory of GPUs goes to work. The system takes our prompt, converts it into tokens and processes a massive dataset, applying mathematical transformations to predict the next token in the sequence. Then it converts those tokens back into words, forming the response we see on screen. As humans, we can’t possibly evaluate what happens inside that overwhelming process. But what we can evaluate is the response. And this snippet becomes our proxy for AI’s intelligence because it’s the only part we can meaningfully review. Elon Musk once said that AI existed long before 2022. What changed was that chatbots connected humans directly to it. Through tools like ChatGPT, we can now see and assess the snippets AI produces.

This perspective of looking at AI through its synthetic artifacts is precisely what we are planning to include in the standard. By describing AI in this grounded, technical way, we can also eliminate the kind of senseless language that sometimes creeps into standards.

As I mentioned before, there’s a clear dichotomy in how we view AI. On the one hand, it can dramatically reduce human effort through automating, accelerating, and scaling tasks that once required large teams. On the other hand, it introduces new risks and challenges that must be carefully managed. AI-enabled systems can enhance resilience and, in some cases, operate autonomously but the role of humans must be clearly defined. Phrases like “always keeping a human in the loop” sound reassuring yet mean little without specifics. Let’s take Python for example. A human could initiate the loop, intervene at every iteration, act only after the loop, or randomly pick certain iterations. Each of these patterns represents a completely different kind of human involvement.

As both Wikipedia and Human Rights Watch point out, there are three distinct categories of human participation. Human-in-the-loop means the process requires explicit human approval for each iteration and the system cannot continue without human confirmation. Human-on-the-loop means the process runs autonomously, but a human can intervene or abort execution, like pressing a “break” command. Human-out-of-the-loop means the process runs fully autonomously, without any human input or control. The Zero Outage standard must define these concepts precisely, for it’s a technical requirement for clarity, accountability, and safety.

AI systems are both non-deterministic and probabilistic, their outputs can vary even under similar conditions. This inherent uncertainty means that when we use ensembles of AI models to support resilience, we must expect different models to produce different results. Consequently, we need to define and assess a range of quality characteristics for AI systems within the Platform Workstream, including adaptability, evolution, autonomy, safety, freedom from inappropriate bias, reward hacking, and unintended side effects, transparency, explainability, and interpretability among others.

Platforms must leverage AI/ML capabilities to improve observability, predictive maintenance, automated root-cause analysis and self-healing capabilities. At the same time AI systems must comply with requirements for transparency, verifiability and fallback mechanisms to prevent new single points of failure. Integration of AI into the platform should include resilience testing of the AI itself, robust data pipelines and documented escalation rules for human oversight, specifying when and how humans must intervene. It is of course still work in progress and we need to do a lot of things to improve the Platform Workstream.

Download presentation
Watch the post-event recording featuring Iosif Itkin and Anna-Maria Lukina: