The Gap Between What Executives Sign and What They Actually Receive
When an SLA lands on an executive's desk, it looks complete: uptime percentages at 99.9%, response times measured in minutes, penalty clauses for breach. This format creates a sense of protection, but it often protects the wrong layer of the engagement. Traditional SLAs were designed for static software systems, where the outcome is known in advance and the only real question is whether the server is running. AI systems are fundamentally different: a model can be available 100% of the time and still produce misclassifications, skewed recommendations, or decisions poorly aligned with a Saudi market context it was only partially trained on.
The real gap surfaces months into operation, when the operations team notices that model accuracy has drifted, or that recommendations have become less relevant to local business context, yet no clause obligates the vendor to act because 'the system was available.' This is not a breach in the literal contractual sense, but it is a breach of the purpose the enterprise paid for. The executive who signed in good faith finds no contractual lever to demand performance correction, because the agreement was never built to measure decision quality, only infrastructure operation.
This distinction is not academic. Enterprises relying on AI systems for risk pricing, customer segmentation, or credit decision support absorb the consequences of accuracy drift directly on the business outcome, while the vendor remains fully protected contractually. The first serious step in handling any AI SLA is a clear separation between technical availability guarantees and functional performance guarantees, treating them as two distinct contractual categories rather than one clause wrapped in unified language.
What Can Actually Be Measured in an AI SLA
Not every dimension of AI performance can be measured with enough rigor to embed in a contract, but that does not justify abandoning measurement altogether. The first clearly measurable category is traditional availability metrics: uptime, response latency, technical error rate. These are useful but say nothing about decision quality. The second, more critical category for Saudi enterprises, is functional performance: classification accuracy against a pre-agreed test sample, the permitted drift threshold before retraining is triggered, and the alignment rate between system output and expert human judgment on defined test cases.
The third category, hardest to draft but most valuable, concerns stability over time: does the vendor commit to periodic model re-evaluation? Is there a documented mechanism for retraining once drift exceeds a defined threshold? Who bears the retraining cost if drift stems from changes in the enterprise's own data, versus who bears it if drift stems from limitations in the original model? These require specific contractual answers, not general promises of 'ongoing support.'
Serious enterprises require the vendor to provide periodic performance reporting in a form suitable for independent audit, not an internal report the vendor issues about itself. This distinction between self-reported and independently verifiable measurement is what turns an SLA from a formal protection document into an actual governance instrument that can inform contract renewal or performance escalation decisions.
What Cannot Honestly Be Guaranteed, and Why Trying to Is a Warning Sign
Any vendor offering a guarantee of '99% accuracy in all cases' or 'zero errors' for a live AI system handling real, changing data is making a promise that cannot honestly be kept, and this should register as a warning sign rather than a trust signal. AI model performance is shaped by input data quality, by shifting user and market behavior over time, and by edge cases that cannot be fully anticipated in advance. A serious agreement acknowledges this constraint explicitly, defining expected performance within stated conditions, such as data quality, volume and distribution, rather than promising an absolute outcome.
Another element that cannot be directly guaranteed contractually is the quality of human decisions made based on system output. A vendor can guarantee that the system will produce a classification or recommendation at a certain accuracy level, but cannot guarantee that the enterprise's internal team will interpret that recommendation correctly or integrate it into an appropriate business context. This places a parallel responsibility on the enterprise itself: building internal capability to interpret and verify system outputs, rather than assuming the contract resolves this on its own.
Attempts to draft overreaching guarantees often mask either a genuine lack of understanding of the model's own limits, or a desire to close the deal quickly without engaging the technical complexity honestly. Saudi enterprises that scrutinize such offers carefully, and ask vendors to state plainly what cannot be guaranteed, typically end up with a more realistic and mature partnership than those who accept absolute promises without unpacking them.
A Practical Framework for Building an SLA That Actually Holds
The framework we recommend begins by classifying every proposed clause into one of three categories: independently measurable, internally measurable subject to periodic audit, and not guaranteeable but requiring clear disclosure of its limits. This classification shifts the SLA discussion from a debate over legal language into a structural conversation about what the enterprise can actually rely on in its operational decisions. Any clause that cannot be clearly placed in one of these three categories needs redrafting before signing, not after.
The second element is the escalation mechanism when performance drifts. A sound agreement specifies precisely: who detects drift, using what tool, within what timeframe, and what the sequenced steps are from alert to retraining to compensation, if applicable. Absent this mechanism, the enterprise discovers the problem through its business impact rather than through an agreed early-warning system, a fundamental difference between proactive governance and after-the-fact crisis management.
The third element concerns intellectual property over data and the customized model, and the enterprise's right to obtain an independently auditable performance record separate from the vendor's own tooling, particularly relevant when considering a future vendor change. An enterprise that builds this framework before signing, not after a problem surfaces, enters negotiation with a clear understanding of what it is actually purchasing, and that alone changes the quality of dialogue with any serious vendor.
Common Mistakes Enterprises Discover Too Late
The first and most frequent mistake is signing a global template SLA without adapting it to the Saudi data context, particularly when a model trained primarily on other markets' data is then deployed locally without performance standards specific to this new context. The result is performance that looks acceptable in the vendor's global reporting but is weak in local application, and the agreement gives the enterprise no lever to demand a different standard because none was defined from the outset.
The second mistake is leaving the definition of 'acceptable performance' deliberately or inadvertently vague, using phrases like 'satisfactory performance' or 'per best practices' instead of specific, referenceable figures and criteria. This vagueness feels flexible at signing but becomes an unresolvable point of contention exactly when it matters, because each party interprets it to its advantage. The third mistake is the absence of any clause covering retraining or future maintenance, treating the project as 'finished' at delivery, when AI systems require periodic maintenance to sustain accuracy, and the absence of this clause means unexpected costs surface months later.
The fourth mistake, more administrative than technical, is failing to assign anyone inside the enterprise clear ownership of monitoring SLA compliance after signing. Even the strongest agreement, without an internal owner tracking it, gradually becomes an archival document rather than an active governance tool. Avoiding these mistakes requires less advanced legal expertise than a clear system of the right questions asked before signing.
Knowing Whether Your Enterprise Is Ready for This Conversation, and the Next Step
This conversation is relevant to your enterprise if you are in one of three situations: preparing to negotiate a new AI contract and wanting to enter with precise questions rather than general trust in the vendor; holding an existing agreement you suspect protects you against technical downtime risk only, not performance quality; or facing a noticeable decline in an existing system's accuracy with no clear contractual lever to demand correction from the vendor. If none of these apply, the discussion is theoretical for you right now, which is entirely fine, and best revisited as an actual project approaches.
Not addressing this gap does not necessarily mean immediate disaster, but it does mean quiet risk accumulation: business decisions gradually built on outputs from a system whose performance is degrading unmonitored, a vendor relationship lacking any objective tool for discussing performance, and a correction cost that is always higher later than the cost of designing it correctly upfront. This is not a call to fear the technology, but a realistic account of how contractual risk accumulates quietly when an SLA is treated as a formality rather than an active governance instrument.
The logical next step is not signing a new contract or switching vendors, but a focused review session of an existing agreement or proposed draft, examining three core questions: what is actually measured, what is assumed without evidence, and what the escalation mechanism is when performance declines. This review gives you a clear map before any deeper financial or contractual commitment, and it is the natural starting point for a serious advisory relationship with someone who understands the difference between system uptime and the quality of the decision it produces.

