Why Saudi LLM evaluation is a separate exercise
Saudi LLM provider evaluation is materially different from generic LLM evaluation in ways that off-the-shelf benchmarks and Western enterprise procurement frameworks systematically miss. The differences arise from four structural facts: Arabic linguistic complexity that generic English-anchored evaluation cannot capture, a sovereignty-first procurement environment that rewards Saudi-controlled deployment over commercial flexibility, an operational reality where most enterprise workloads are heavily code-switched between Arabic and English, and a regulatory layer (PDPL, SDAIA, sector regulators) that constrains which deployment architectures are even viable for which workloads.
A foreign company evaluating Saudi LLM providers using generic LLM-leaderboard methodology — MMLU scores, HumanEval scores, generic latency benchmarks, headline $/M-token pricing — produces a procurement decision that looks rigorous and is operationally wrong. The right evaluation is structurally different and takes longer. This guide walks through the four-question framework that consistently produces correct procurement decisions, the bake-off methodology that surfaces real differentiation, the candidate model landscape in 2026, and the decision rules for matching providers to workloads.
Question 1 — Arabic-first capability vs. Arabic-as-second-language
The capability gap between models trained from inception with Arabic in the corpus weighting and models with Arabic added through post-training is meaningful and operationally consequential. Allam (34-billion-parameter Arabic-first foundation model trained on 8PB of corpus weighted toward Arabic and culturally relevant content) sits at one end of the spectrum. Falcon and Jais (regional Arabic-first models from UAE-anchored programs) sit nearby. GPT-4-class models with Arabic post-training, Claude with regional alignment work, Gemini with Arabic-aware fine-tuning sit in the middle. Generic English-trained models with no Arabic post-training sit at the unusable end for serious Saudi enterprise deployment.
For high-fidelity Arabic generation — government correspondence, legal drafting, religious-text handling, MSA-vs-dialect navigation, code-switched business communication — the Arabic-first models reliably outperform on Arabic-specific evals (AraBench, Arabic MMLU, dialect-specific benchmarks) even when the US-anchored models score materially higher on English benchmarks. The English-benchmark gap does not predict Arabic-deployment quality.
For mixed-language enterprise workloads (English-language RAG over Arabic source documents, multilingual customer service, English-coding workloads with Arabic UI) the US-anchored models with Arabic post-training are usually adequate and sometimes superior on the English-heavy components. The procurement question becomes: what is the language mix of the actual workload, weighted by criticality?
The framework: weight Arabic-first capability heavily for sovereign-cloud deployments, government workloads, religious/cultural workloads, and customer-facing Arabic-only applications. Weight English-and-coding capability heavily for developer tooling, English-source-document RAG, and international-team-facing applications. Most enterprise deployments are mixes; the right answer is often a multi-model architecture.
Question 2 — Sovereignty posture
Where does the model run, where are the weights stored, and who has access to inference logs? This is the question that most foreign vendors under-weight and that Saudi procurement increasingly leads with.
The sovereignty spectrum in 2026, from maximalist to minimal:
Maximalist sovereignty — Allam running in SDAIA-controlled Hexagon DC with sovereign-cloud designation, Saudi-controlled inference logs, Saudi-controlled fine-tuning data, Saudi-controlled alignment posture. The default for government and defense workloads.
Strong sovereignty — Hyperscaler-hosted models running in named KSA regions with PDPL residency commitments, regional-addendum binding under Saudi law, named regional operational personnel. Bedrock-Anthropic in AWS Riyadh, Vertex-Gemini in Google KSA region, Azure-OpenAI in Microsoft KSA region. Adequate for most regulated enterprise deployments.
Intermediate sovereignty — Hyperscaler-hosted models running in regional data centers (Bahrain, UAE) with PDPL residency wrappers but corporate parent foreign-controlled. Adequate for general enterprise but not for government or critical-infrastructure workloads.
Weak sovereignty — US-hosted API access with cross-border data flow, no PDPL residency commitment, foreign-controlled inference logs. Effectively non-viable for any regulated Saudi deployment.
The decision rule: identify the binding sovereignty constraint on the workload (procurement clause, regulatory framework, customer expectation) before evaluating models. Workloads with sovereign-cloud mandates eliminate everything below maximalist sovereignty regardless of capability.
Question 3 — Eval performance on Saudi-relevant tasks
Generic MMLU and HumanEval scores tell you nothing about Saudi enterprise fitness. The relevant evals for Saudi deployment are:
- Arabic comprehension and generation (AraBench, Arabic MMLU, ArabicNLI). Baseline capability across MSA reading, generation, and reasoning.
- Dialectal handling (Saudi-Najdi, Hejazi, Eastern-Province, where dialect-specific evaluation infrastructure is improving). Differentiating for consumer-facing applications.
- Religious-text accuracy (Quranic citation accuracy, Hadith reference handling, fiqh-aware reasoning, Islamic-finance terminology). Dispositive for religious-tourism, religious-education, and Islamic-banking applications.
- Legal-Arabic drafting (formal legal vocabulary, Saudi-specific legal terminology, contract drafting fluency). Dispositive for legal-services applications.
- Code-switching (Arabic-English mixed instructions, mixed-language document processing, cross-language reasoning). Critical for most enterprise workloads.
- Instruction-following in Arabic vs. English (parity check on whether Arabic-instruction operation degrades). Important for sovereign-mandate deployments.
- Hallucination and refusal calibration in Arabic (whether the model’s safety posture in English carries over to Arabic). Important for production safety.
Insist on evals that match your actual workload. Generic LLM leaderboard scores are usually misleading; workload-matched evaluation produces operationally correct decisions.
Question 4 — Total cost vs. total integration burden
Headline $/M-token pricing for Saudi LLM providers is mostly noise. The signal is total cost over a 24-36 month deployment, including:
- API call costs at projected production volume
- Engineering time to integrate (the largest hidden cost — usually 3-10x the API cost in year one)
- Custom fine-tuning costs (data preparation, training compute, alignment work, evaluation)
- Compliance overhead (PDPL, SDAIA, sector regulator engagement specific to the model choice)
- Vendor lock-in cost (the cost of migrating off the chosen provider 24 months in if performance disappoints)
- Relationship cost (the political and commercial cost of choosing one Saudi anchor over another, where Allam-via-Humain is a different relationship from Claude-via-AWS-via-AWS-Saudi-team)
Allam through Humain Chat is operationally simple but locks deployment into a specific architectural pattern. Anthropic Claude through AWS Bedrock-Riyadh is technically flexible but demands integration engineering. OpenAI through Azure-KSA is mid-flexibility, mid-integration. The right total-cost answer depends on the workload’s integration complexity, fine-tuning needs, and exit-flexibility requirements.
How to actually run the bake-off
A 90-day pilot across exactly three providers — not two, not five. Three is enough to surface differentiation while staying logistically manageable.
- Weeks 1-3: Set up production-equivalent evaluation infrastructure. Define the workload. Build the evaluation set: 200-500 prompts drawn from real workload, scored by 3+ Saudi-Arabic-native evaluators on accuracy, fluency, cultural appropriateness, refusal-behavior calibration. This is the most important step; everything downstream depends on it.
- Weeks 4-9: Run the three candidate providers against the evaluation set. Score on accuracy, latency under production load, integration complexity, and total cost projection. Native-speaker review remains the dominant scoring mechanism for Arabic quality.
- Weeks 10-12: Production-shadow deployment of the top two candidates against real workload (with appropriate user-experience safeguards). Compare outputs, latency, cost, and operational behavior.
- Decision and contract negotiation: Decision is now empirically grounded; contract negotiation focuses on commercial terms with the chosen provider.
Skip native-speaker review at your peril. No public benchmark replaces it.
The 2026 candidate model landscape
For Saudi-deployment LLM evaluation in 2026, the candidate set typically includes:
- Allam (Saudi-sovereign, Arabic-first, the SDAIA-aligned default for sovereign and government deployments)
- Anthropic Claude (commercial deployments via AWS Bedrock-Riyadh, strong reasoning, strong safety posture, robust Arabic post-training)
- OpenAI GPT family (commercial deployments via Azure-KSA, strong general capability, dominant developer ecosystem)
- Google Gemini (commercial deployments via Vertex-KSA, strong multimodal, increasingly strong Arabic capability)
- Meta Llama with Arabic fine-tunes (open-weight option for self-hosted deployments, strong for cost-sensitive use cases, requires fine-tuning investment)
- Mistral and Falcon-derived options (regional alternatives with different alignment posture, useful for specialized use cases)
The right answer depends on workload. Sovereign deployments default to Allam. Commercial deployments default to one of the hyperscaler-hosted options with Arabic post-training. Specialized deployments often use multi-model architectures.
Where the candidate models physically run
The sovereignty analysis in Question 2 becomes concrete when mapped to actual in-Kingdom infrastructure, because each provider’s physical footprint determines which sovereignty tier it can genuinely deliver:
- Allam runs on SDAIA-managed infrastructure, anchored by the SDAIA sovereign AI factory in Riyadh (up to 5,000 Blackwell GPUs, deploying) and the Hexagon data center — the 480 MW Riyadh facility, operational in early 2026, that supports the National Data Lake spanning 430+ government systems. This is the maximalist-sovereignty tier made physical.
- Claude via Bedrock-Riyadh rides the AWS × Humain $5.3B cloud region reaching service in 2026.
- GPT family via Azure-KSA sits on the Microsoft × Humain $1.5B Azure expansion.
- Gemini via Vertex anchors to the Google Cloud × Humain $10B global AI hub in Dammam, with the 300 MW campus (NVIDIA silicon plus Google TPU) targeting 2027.
- Inference-specialist silicon shapes the latency and cost tier for production Arabic workloads: Groq’s LPU deployment under the $1.5B Humain commitment targets real-time Arabic inference, with the Aramco Digital cluster serving EMEA and South Asia; SambaNova’s SN40L infrastructure with SDAIA ($140M, deployed since 2024) supports Arabic model development; Qualcomm’s AI200/AI250 rack services bring 200 MW of inference capacity starting 2026 on a hybrid edge-cloud pattern.
The practical consequence: a provider whose Saudi region is operational today and one whose region reaches general availability in 2027 are different procurement risks even at identical benchmark scores. Ask every candidate for the operational status of the specific region that will serve your workload, and verify against facility-level reporting rather than marketing language.
Pricing the shortlist
In-Kingdom inference pricing in 2026 anchors as follows (per million tokens, input and output): Claude Sonnet-class via Bedrock-Riyadh at $3.00-3.30 and $15.00-16.50, Opus-class at $15.50-17.20 and $77-84, Haiku-class at $0.80-0.95 and $4.00-4.60; GPT-4o via Azure-KSA at $2.55-2.85 and $10.20-11.40, GPT-4o-mini at $0.16-0.19 and $0.62-0.75; Gemini 1.5 Pro via Vertex-Dammam at $1.25-1.45 and $5.10-5.80, Flash at $0.075-0.085 and $0.30-0.36. Regional premia over global baselines run 0-5% on Bedrock, 0-8% on Azure, and near parity on Vertex. Committed-throughput structures pull effective per-token cost down 25-55% for high-utilization workloads.
Worked against a representative enterprise workload — 1B input tokens and 200M output tokens monthly, typical of a mid-scale customer-service deployment — the monthly API spend ranges roughly $2,300-2,600 on Gemini 1.5 Pro, $4,600-5,100 on GPT-4o, and $6,000-6,600 on Sonnet-class Claude, before committed-throughput discounts. The spread looks decisive and is not: at these volumes the annual API delta between the cheapest and most expensive candidate is under $55K, while the year-one integration engineering delta between a well-matched and a poorly-matched provider routinely runs 3-10x the API cost. Pricing eliminates candidates only at the extremes; the Question 4 integration analysis does the real work.
Worked example — a ministry citizen-service deployment
A ministry deploying an Arabic citizen-service assistant illustrates the full framework in sequence. Question 2 resolves first: a government workload carries a sovereign-cloud mandate, which eliminates everything below the maximalist tier and defaults the citizen-facing component to Allam on SDAIA-controlled infrastructure. Question 1 then splits the workload: the citizen-facing surface is Arabic-first with heavy Najdi and Hejazi dialect exposure — Allam territory — while the internal analytics layer (English-language reporting over interaction logs, code-switched staff tooling) is better served by a hyperscaler-hosted model in a named KSA region at the strong-sovereignty tier. Question 3 dictates the evaluation set: 300-500 real citizen interactions scored by native evaluators on dialect handling, refusal calibration in Arabic, and escalation accuracy — not MMLU. Question 4 prices the two-model architecture against a single-model alternative and consistently favors the split: forcing the analytics layer onto the sovereign stack costs integration flexibility, while forcing the citizen surface onto a commercial model fails the sovereignty constraint outright. The result is the multi-model architecture the framework predicts — Allam where the mandate binds, a Bedrock- or Azure-hosted model where it does not, and a routing layer as the primary integration investment.
Contract clauses that matter
Five clauses separate a defensible Saudi LLM contract from a liability. Inference-log residency and access: specify where logs live, who can read them, and the deletion schedule — this is the operational core of the PDPL posture and the first question a SDAIA-aligned auditor asks. Fine-tune ownership: custom weights trained on your data should be contractually yours, including export rights on exit; providers differ sharply on default terms. Region-pinning: the contract must name the serving region and require consent for failover outside the Kingdom — silent failover to a non-KSA region is a compliance breach you inherit. Benchmark-refresh rights: given model-version churn, secure the right to re-run the acceptance evaluation on version upgrades, with rollback rights on regression. Exit assistance: migration support, prompt-and-eval portability, and committed-capacity unwind terms, priced at signature rather than negotiated under duress 24 months in.
Evaluation pitfalls
The recurring failures: scoring Arabic capability with English-benchmark proxies; running the bake-off on synthetic prompts instead of production transcripts; letting a single bilingual staffer stand in for a native-evaluator panel; treating the provider’s hosted-eval sandbox as production-equivalent (latency and refusal behavior differ under load); ignoring version-pinning, so the model evaluated is not the model deployed; and negotiating price before sovereignty tier, which cedes the only leverage that matters. Each is cheap to avoid during the 90-day pilot and expensive to discover after contract signature.
The decision rule
The synthesis: identify the binding sovereignty constraint, identify the binding language-mix profile, identify the binding integration constraint, and pick the model that best fits all three rather than the model that scores highest on any single dimension. This produces operationally correct decisions. Optimizing on any single dimension produces decisions that look defensible at procurement and disappoint at deployment.
For deeper reading: How to deploy Allam · How to evaluate Arabic LLM · How to pick Saudi cloud provider · Player Directory.