The question everyone asks, and why it is now the wrong one
In almost every AI procurement conversation we sit in, the first governance question is the same: will they train on our data?
It is a reasonable question and it has become a settled one at the model layer. Anthropic’s commercial terms state plainly that it may not train models on customer content from its services. OpenAI’s enterprise terms state that it does not train on business customers’ data by default, with thirty-day deletion for enterprise conversations and zero-data-retention available for eligible API endpoints.
So the frontier providers have answered it. The risk did not disappear — it moved.
It now sits with the layer between you and the model: the application vendor, the systems integrator, the agent platform. Their terms are not the model provider’s terms. Their telemetry pipelines, evaluation datasets and support tooling may retain your content in ways the underlying provider’s contract would prohibit. That is where the clause you actually need to read has gone.
Nine things to establish before signing
1. Who are the sub-processors, by name?
Not leading cloud and AI providers. Names, and the right to be notified before the list changes. If a vendor will not name them, you cannot assess residency, you cannot assess the training boundary, and you cannot answer your regulator’s third-party due-diligence question.
2. Does the no-training commitment flow through?
A vendor building on a model with zero data retention does not automatically inherit that posture. Ask specifically whether zero retention is enabled on the endpoints your workload uses, and whether the vendor’s own logging, evaluation and support systems retain prompts and outputs. The second half is where the real exposure usually is.
3. Where does inference actually happen?
There is a large difference between hosted in the region, sovereign-ready, and runs inside an account you control. Microsoft’s Saudi Arabia East region becomes available to customers in Q4 2026 and is described as sovereign-ready, with sovereign services still being explored. Sovereign-ready is not sovereign. If a regulated workload may need to move in-country or on-premise, that has to be architecturally true from the start.
4. Can the deployment survive a supervisor’s question?
The CBUAE’s February 2026 guidance asks licensed institutions to secure audit rights over AI vendors, including regulator access. Whatever your sector, that is a good template. If the contract cannot produce an audit trail of what the system did and on whose authority, the governance story ends at the contract.
5. Is it actually an agent?
Gartner estimates that of the thousands of vendors marketing agentic AI, roughly one hundred and thirty are genuine — the rest are assistants, RPA and chatbots rebranded. It calls this agent washing. The test is not whether the system uses a language model. It is whether it can take an action in a system of record within a policy boundary, and whether you can see and constrain that boundary.
6. What is the pricing exposed to?
Per-seat pricing for assistive tools is under structural pressure. Gartner expects more than half of enterprises to stop paying for assistive intelligence by 2028 in favour of outcome-focused workflow platforms, and expects software companies bolting AI onto legacy applications to face margin compression of up to eighty percent by 2030. Separately, about twenty percent of McKinsey’s 2026 respondents said AI operating costs were constraining their usage. Model a consumption scenario at three times your expected volume before you sign, and ask what happens to the price when the vendor’s own model costs move.
7. Was the Arabic tested, or claimed?
This deserves its own section, below.
8. What does the evaluation look like — theirs and yours?
Ask for the evaluation set, the failure cases, and the conditions under which stated latency or accuracy figures were measured. A number without test conditions is marketing. And build your own set: fifty real cases from your own operation, with the right answers written down before the demo, is worth more than any benchmark.
9. What happens when you leave?
Export format for conversations, records and any tuned artefacts. Deletion timeline, in writing. Whether anything you contributed persists in the vendor’s evaluation or improvement pipeline after termination.
On Arabic, specifically
The regional market has spent three years treating Arabic capability as the differentiator. On the published evidence, that has largely stopped being true for general Arabic.
BALSAM, a benchmarking platform covering 78 tasks and 52,000 examples, found large closed models outperforming Arabic-centric models by sizable margins. More tellingly, a 2025 survey of Arabic LLM evaluation authored by researchers at the Technology Innovation Institute in Abu Dhabi — the institution behind Falcon — reached the same conclusion about frontier multilingual models outperforming Arabic-specific alternatives.
When TII launched Falcon-H1 Arabic in January 2026 with strong Open Arabic LLM Leaderboard scores, the comparison set was regional and open-weight models. It did not benchmark against the frontier closed models. Leading Arabic model in that context means leading among open models, which is a real achievement and a different claim.
So where does Arabic still differentiate?
- Dialect and code-switching. The evaluation literature identifies these as unsolved and under-tested. Gulf dialect, and Arabic-English switching inside a single sentence, is the normal case in UAE customer service and almost never what a benchmark measures.
- Cultural alignment, which is separate from linguistic competence.
- Deployment locus. Open-weight Arabic models can run on-premise or in-country. Frontier closed models generally cannot. For a regulated workload the differentiator is sovereignty, not Arabic skill.
- Cost at scale, where a smaller open model may be economically preferable regardless of frontier superiority.
The practical consequence: stop asking vendors whether they support Arabic. Everyone says yes. Give them twenty recordings or transcripts from your own operation — real ones, with dialect, with names, with code-switching, with the noise — and score the output yourself. It takes an afternoon and it settles the question that six months of demos will not.
The governance gap nobody has closed
One number is worth holding onto. In Deloitte’s 2026 research across 3,235 IT and business leaders in 24 countries, only twenty-one percent reported mature governance models for agentic AI — while seventy-four percent anticipated moderate-to-extensive agent deployment by 2027.
Four in five organisations are deploying systems that take actions without having decided how those actions are bounded. The missing mechanisms are consistent: clear decision boundaries between what an agent may do autonomously and what needs human approval; real-time monitoring of behaviour rather than periodic review; and audit trails of the full chain of actions, not just the final output.
Gartner’s related projection is that by 2030 half of agent deployment failures will trace to insufficient runtime enforcement. That is a governance-platform problem, and it is not something you can add after the vendor is embedded.
Our own bar
We evaluate platforms on behalf of clients, and we hold ourselves to the same four tests we would apply to any vendor:
- Deployable in your own cloud or in-country — not because every workload needs it, but because the ones that do cannot wait for a roadmap.
- Arabic tested rather than claimed, on your material, scored by your people.
- A contractual training boundary, flowing through every layer between you and the model.
- Pricing that survives a Gulf mid-market deal at realistic volume, not a headline seat price.
We name platforms once an agreement is in place, and not before. A recommendation made before the commercial terms are settled is not advice — it is a referral with the disclosure missing.
Sources: Anthropic Commercial Terms of Service; OpenAI Enterprise Privacy; Gartner on agentic AI project cancellation and agent washing (25 June 2025); Gartner on assistive AI and outcome-focused workflow (2 April 2026); Deloitte, Agentic AI is scaling faster than guardrails (24 April 2026); BALSAM: A Platform for Benchmarking Arabic Large Language Models; Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps (TII).
