The Saudi LLM landscape

Saudi LLM deployment splits between sovereign-controlled foundation models (Allam, the 34B-parameter Arabic-first model under SDAIA) and frontier-model access through hyperscaler partnerships (OpenAI via Azure, Anthropic via AWS Bedrock, Meta Llama open-weights, xAI Grok via the 500 MW Humain JV, Mistral selectively). Open-source Chinese models (DeepSeek and adjacent Chinese-origin model families) have limited official adoption due to US export-control alignment under the November 2025 framework.

The ranking below maps the LLM availability landscape: sovereign-owned versus frontier-imported, Arabic-native versus English-trained-with-Arabic-finetune, commercial-license versus open-weights. Each model occupies a distinct position in Saudi enterprise and government deployment patterns. Reading the ranking requires holding multiple dimensions simultaneously rather than collapsing to a single benchmark.

Reading the top entries

Humain anchors the ranking because it operates the deployment infrastructure for most of Saudi Arabia’s LLM workloads, including hosted Allam, hosted xAI Grok at the 500 MW JV site, hosted Anthropic and OpenAI workloads through hyperscaler integrations, and the Humain Chat consumer interface. Humain is not itself an LLM developer (with the partial exception of the Humain Chat layer) but the deployment infrastructure that connects Saudi users and enterprises to LLM capability. Its top position reflects deployment scale rather than model authorship.

Allam is the sovereign-controlled Arabic-first foundation model under SDAIA. The 34B parameter scale is intentionally sized for deployment efficiency — Allam runs on a single H100/H200 node for inference, which simplifies deployment compared to 70B+ frontier alternatives. The training data comprises 8 PB Arabic-weighted corpus with English supplementation, producing meaningful Arabic-language capability advantages over English-trained-with-Arabic-finetune frontier models on dialectal nuance, classical Arabic, and culturally-specific content. For Arabic-first applications and government workloads requiring full sovereignty, Allam is the structural choice.

SDAIA appears alongside Allam because SDAIA is the operating authority for Allam plus the broader sovereign AI program. SDAIA operates the National Data Lake (430+ government systems integrated), the Hexagon government data center hosting Allam workloads, and the regulatory framework governing AI deployment under PDPL. SDAIA’s position on the LLM ranking captures its operational role rather than separate model authorship.

Meta AI (Llama family) and Mistral AI provide the open-weights tier of LLM availability in Saudi Arabia. Llama 3.x and Llama 4 (when released) provide capable open-weights alternatives that Saudi enterprises and government agencies can fine-tune for specific applications and deploy on Saudi-domiciled infrastructure. Mistral provides European open-weights alternatives with selected commercial licensing options. Both serve the use cases where open-weights customization is preferred over hosted frontier-model API access.

The frontier-tier availability through hyperscalers

Below the sovereign and open-weights tiers, frontier-tier LLM availability runs through hyperscaler partnerships. OpenAI GPT (4o, 4.5, future generations) is available through Azure OpenAI Service in the Microsoft Saudi region (Q4 2026 launch). Anthropic Claude (Sonnet, Opus, future generations) is available through AWS Bedrock in the AWS Riyadh region (operational). Google Gemini (Pro, Ultra, future generations) is available through Vertex AI in the Google Cloud Dammam hub (operational ramp). xAI Grok (Grok 3, Grok 4, future generations) is available through the 500 MW Humain JV site (operational scale through 2026-2027).

The frontier-tier availability spans the major US frontier labs without obvious gaps. Saudi enterprises and government agencies have credible access to the leading frontier models for English-language workloads. The constraint is less model availability than data residency and sovereignty — frontier models running on hyperscaler infrastructure require careful workload classification to ensure sovereignty-sensitive data does not flow to non-Saudi infrastructure.

The Arabic-language capability differentiation

A persistent question in Saudi LLM deployment is the Arabic-versus-English capability trade-off. Frontier models from US labs are predominantly English-trained with Arabic supplementation. They handle Arabic competently but with English-trained biases, occasional dialectal mishandling, and weaker performance on classical Arabic, religious texts, dialectal Arabic and culturally-specific nuance. Allam is Arabic-first trained with English supplementation, producing meaningful Arabic-language capability advantages on the same dimensions.

For most enterprise English-language applications, frontier models retain capability advantages over Allam given the larger parameter count and broader English training. For Arabic-language applications targeting Saudi consumers and citizens, Allam provides the structural choice. Mixed Arabic-English applications often deploy routing-based architectures that combine both: Allam for Arabic-language workloads, frontier models for English-language workloads, with intelligent routing based on input-language detection.

The agentic and multimodal dimensions

Through 2025-2026, the LLM ranking is increasingly differentiated by agentic and multimodal capabilities beyond pure text generation. Frontier models are advancing fastest on agentic capability (multi-step planning, tool use, autonomous task execution) and multimodal capability (vision, audio, video, increasingly embodied). Saudi sovereign deployments through Allam are following the trajectory but at slightly slower cadence reflecting smaller R&D scale than frontier US labs.

The agentic dimension matters because increasingly enterprise AI deployment requires not just text generation but task execution. SDAIA’s citizen-services AI applications, Humain’s enterprise tooling and Aramco’s industrial AI workloads all benefit from agentic capability. Watch the trajectory of Saudi sovereign agentic capability through Allam’s evolution and through dedicated agentic-AI initiatives.

What the ranking misses

The ranking captures the major LLM availability and undercounts smaller specialized models. Domain-specific Saudi-developed models (legal AI, medical AI, financial AI) operating at smaller scale contribute to total Saudi LLM capability. Open-weights model derivatives (Saudi-fine-tuned Llama variants, custom Allam fine-tunes for specific workloads) operate within enterprise and government deployments without appearing as headline LLM positions.

The ranking also undercounts the embedded LLM dimension. Enterprise SaaS applications (Salesforce Einstein, ServiceNow Now Assist, Microsoft Copilot, Google Workspace AI) embed LLM capability without exposing the underlying model as a standalone choice. The cumulative embedded LLM consumption in Saudi Arabia is substantial but distributed across enterprise SaaS rather than concentrated in headline LLM deployments.

What changes the ranking

Three forcing functions reshape the LLM ranking through 2027. First, Allam’s evolution to larger parameter scale or capability tiers — if SDAIA invests in next-generation sovereign models at frontier scale, the sovereign tier deepens. Second, frontier model trajectory at the major US labs — capability advances at GPT, Claude, Gemini and Grok shift the relative attractiveness of frontier-versus-sovereign deployment. Third, open-weights model trajectory — if Llama, Mistral or other open-weights models close the frontier capability gap, more deployment shifts toward open-weights customization rather than hosted frontier APIs.

The methodology disclosure for LLM ranking

The Saudi LLM ranking weights five composite factors: cumulative deployment scale within Saudi infrastructure, sovereignty positioning (sovereign > frontier-hosted > open-weights deployed), Arabic-language capability depth, strategic anchor relationships with Saudi sovereign AI architecture and operational maturity. The ranking captures the operationally relevant LLM availability layer rather than pure capability benchmarks. Frontier capability matters but Saudi-specific deployment depth matters more for the ranking purpose.

Two recurring data-quality issues affect the methodology. First, LLM benchmarks vary in their Arabic-language coverage and consistency; comparing models requires multiple benchmarks rather than a single composite. Second, the boundary between LLM and broader generative AI capability is increasingly blurred as multimodal capabilities become standard; the ranking captures text-LLM availability with emerging multimodal capability noted where relevant.

The reasoning-model trajectory

Reasoning models (OpenAI o-series, DeepSeek R-series, broader chain-of-thought architectures) are reshaping LLM capability through 2025-2026. Saudi LLM deployment increasingly considers reasoning-model availability alongside conventional LLM capability. Frontier reasoning models are available through hyperscaler partnerships at the same access tier as their base-model counterparts. Sovereign reasoning capability through Allam-anchored architectures is at earlier stage but advancing.

For high-stakes Saudi deployments (legal AI, medical AI, complex government decision support), reasoning capability is becoming a hard requirement. The choice between hyperscaler-hosted reasoning capability and sovereign-hosted reasoning capability follows the same architectural trade-offs as the broader LLM choice. Watch the trajectory of Saudi sovereign reasoning capability as Allam evolves and as adjacent SDAIA R&D produces dedicated reasoning architectures.

The DeepSeek question

A persistent question in Saudi LLM landscape is the role of Chinese-origin open-weights models including DeepSeek. The November 2025 BIS framework’s Chinese-equipment exclusion applies to hardware in approved AI facilities; the model-weights side of the question is more nuanced. DeepSeek’s open weights are technically available globally including Saudi Arabia but official Saudi enterprise and government deployment has been limited reflecting US export-control alignment and broader sovereignty considerations.

Selected research and capability-evaluation deployments of Chinese-origin open weights occur at Saudi research institutions and enterprises but at smaller scale than the headline US frontier and Saudi sovereign tier. The trajectory through 2026-2027 depends on broader US-China AI competition dynamics. If Chinese open-weights capability advances meaningfully relative to US frontier models, Saudi pragmatic deployment could increase; if the gap remains, Saudi deployment will likely stay limited to research and evaluation contexts.

The Allam evolution trajectory

Allam’s evolution through 2026-2028 is the single most consequential variable for the Saudi sovereign LLM tier. Plausible trajectories include: scaling Allam to larger parameter counts (70B+) while maintaining Arabic-first training; developing successor models with multimodal capability (vision, audio); developing frontier-tier Saudi-controlled models that close the capability gap with US frontier labs at meaningful scale; or layering specialized fine-tunes for specific Saudi vertical applications (legal, medical, financial, industrial).

Each trajectory has different capital requirements, talent requirements and timeline implications. The fastest-cycle path is specialized fine-tunes (months to years); the slowest-cycle path is frontier-tier sovereign models (multi-year R&D investment at scale comparable to US frontier labs). The most likely trajectory through 2027 combines incremental Allam scaling, multimodal extensions and specialized fine-tunes — producing meaningful capability progression without attempting to close the full frontier-capability gap. Watch SDAIA’s R&D disclosures for the trajectory signals.

The Saudi-specific evaluation considerations

Saudi enterprise and government evaluation of LLM capability emphasizes considerations that differ from typical Western enterprise evaluation. Arabic-language fidelity across dialectal, classical and Modern Standard Arabic. Cultural and religious sensitivity in generated outputs. Sovereignty controls over training-data exposure and inference-data flow. Integration with Saudi-domiciled data infrastructure. Alignment with PDPL and broader Saudi regulatory frameworks. Compatibility with Allam-anchored architecture for hybrid deployments. Compatibility with frontier model architectures for hybrid deployments where appropriate.

The Saudi-specific evaluation criteria affect which models actually get deployed at scale even where capability benchmarks suggest different rankings. A frontier model that scores higher on English-language benchmarks but lacks Arabic-language fidelity often loses to Allam in Saudi deployment selection. A model with comparable capability but better sovereignty controls often wins selection over a model with marginal capability advantage but weaker sovereignty positioning.

The cost-of-inference dimension

LLM deployment economics increasingly hinge on cost-of-inference rather than cost-of-training. Sovereign-controlled inference (Allam on Hexagon, Humain Chat on Humain campuses) carries fixed-cost economics — the underlying infrastructure is paid-for and incremental inference is at marginal cost. Hyperscaler-hosted frontier model inference carries usage-based economics with per-token pricing that scales with consumption. For high-volume Arabic-language workloads, sovereign-hosted Allam frequently delivers superior unit economics; for low-volume or English-language workloads, hyperscaler-hosted frontier models frequently deliver superior capability per dollar.

The cost-of-inference dimension shapes Saudi enterprise LLM deployment patterns. Saudi banks running customer-service AI in Arabic frequently deploy Allam-anchored architectures. Saudi enterprises running internal English-language productivity AI frequently deploy frontier-model API access. Mixed workloads run routing-based architectures that send each request to the cost-optimal provider. The architecture choice cumulates over multi-year deployment into meaningful cost differentials.

The fine-tuning ecosystem

Beyond foundation model availability, Saudi LLM deployment increasingly depends on fine-tuning ecosystem maturity. Allam supports fine-tuning through SDAIA-coordinated processes for government workloads and increasingly through commercial channels for enterprise applications. Open-weights models (Llama, Mistral, selected smaller models) support full fine-tuning on Saudi-domiciled infrastructure. Hosted frontier models support hosted fine-tuning through hyperscaler APIs (Azure OpenAI fine-tuning, AWS Bedrock fine-tuning, Vertex AI fine-tuning).

The fine-tuning architecture cascades into deployment sovereignty. Fully sovereign deployments require open-weights or sovereign-controlled foundation models that can be fine-tuned within Saudi infrastructure boundaries. Partial-sovereignty deployments use hosted fine-tuning with sovereignty-controlled access. Frontier-tier deployments accept the hosted-fine-tuning model with hyperscaler operational control.

The token-volume dimension

Saudi LLM deployment scales with cumulative token volume across applications. Government services (citizen-facing chatbots, internal productivity, briefing generation), enterprise applications (customer service, sales automation, document analysis), consumer applications (Humain Chat, embedded AI in Saudi consumer apps) and adjacent integration (cross-ministry workflows under SDAIA) cumulatively generate billions of tokens monthly across the Saudi LLM ecosystem.

The token-volume aggregate is a leading indicator of LLM ecosystem maturity. Through 2026, the cumulative Saudi LLM token volume should grow substantially as Year of AI 2026 ministry deployments scale citizen-facing AI applications. Watch the published deployment metrics from SDAIA, Humain and the major Saudi banks for cumulative token-volume signals.

For the Arabic-specific subset, see the Saudi Arabic LLMs ranking. For the broader AI startup ecosystem, see the Saudi AI startups ranking. For the agentic-AI subset, see the Saudi agentic AI ranking.

For deeper reading: