A Sovereign Arabic Foundation-Model Stack
Saudi Arabia is the only country to have made Arabic foundation-model capability a top-tier strategic priority and to have funded it with the kind of resources that allow it to be competitive with frontier English-language efforts. The Allam family of foundation models, developed under SDAIA’s leadership and announced in 2023 with successive generations released since, is the principal artifact of this strategy. Allam is positioned not just as an Arabic model but as a sovereign substrate on which Arabic-speaking enterprises and governments — first in Saudi Arabia, then across the wider Arab world — can build production AI without depending on US-controlled foundation models.
The strategic logic has three layers. First, language sovereignty: Arabic is the native language of more than 400 million people, and the structural quality of foundation-model performance in Arabic has historically lagged English by enough that consumer and enterprise applications have been measurably degraded. Second, content sovereignty: a foundation model trained primarily on Western internet content embeds a worldview that is not aligned with Saudi cultural and religious values, and remediation through fine-tuning and prompting is incomplete. Third, infrastructure sovereignty: foundation-model training is the workload that most concentrates strategic AI capability, and a country that cannot train its own foundation models at frontier scale is structurally dependent on those that can.
Allam Infrastructure
Allam’s training infrastructure has scaled rapidly. The first-generation Allam models were trained on relatively modest accelerator clusters by frontier standards, but the SDAIA-led ecosystem has progressively secured access to substantially larger clusters as the Humain partnerships with NVIDIA, AMD, and others have come online. The current generation training runs at the Hexagon data center outside Dammam, with peering to additional capacity at SDAIA-operated facilities and at partner sites operated by STC, Mobily, and Aramco Digital. Training-cluster sizes for the most ambitious runs are now in the same order of magnitude as second-tier frontier labs, and they are projected to reach frontier-comparable scale by 2027.
The compute envelope is one constraint. The other, equally important, is the data envelope. Arabic training data at the scale required for frontier models is substantially scarcer than English data, and the quality of available Arabic web content is uneven. SDAIA has been investing heavily in data curation — licensing high-quality Arabic content from publishers, digitizing the holdings of the King Fahd National Library and the King Abdulaziz Public Library, partnering with the Ministry of Culture and the Saudi Press Agency for archive access, and running large-scale synthetic data programs that use existing Allam generations to bootstrap higher-quality data for subsequent generations.
Hexagon Data Center Training Capacity
Hexagon is the principal operating asset for Saudi foundation-model training. The site is a hyperscale-class facility with multi-hundred-megawatt capacity, designed from the ground up for AI training workloads with the requisite power density, cooling, and high-bandwidth interconnect. Capacity at Hexagon is allocated under the Humain governance umbrella across a portfolio of strategic workloads, with Allam training, defense AI training, and enterprise foundation-model training as the principal claimants. The site’s compute roadmap is paced against the Kingdom’s 1.9 GW AI compute target and is expected to absorb a meaningful share of incremental capacity through 2030.
The architectural posture at Hexagon is deliberately heterogeneous. NVIDIA Hopper and Blackwell capacity is the dominant tier for frontier-style training, but AMD MI300-class and AI accelerator capacity from other vendors is also operated for workloads that benefit from cost or supply-resilience reasons. The orchestration layer is built on a combination of Kubernetes-based workload management and NVIDIA’s training-stack tooling, with Saudi-developed extensions that enforce sovereignty controls on data flows and on weight access.
Arabic Data Corpus Challenges
The Arabic data corpus problem is more than a scale problem. Arabic exists in a diglossic continuum that runs from the classical Arabic of the Quran and the medieval scholarly tradition through Modern Standard Arabic to a wide variety of regional spoken dialects — Saudi (and within Saudi, a substantial sub-dialect mix), Egyptian, Levantine, Maghrebi, Iraqi, Gulf, Yemeni, and others. A foundation model that handles only Modern Standard Arabic is operationally useless for many consumer applications, while a model that handles dialect mix without preserving classical-Arabic fidelity fails on cultural and religious content.
SDAIA’s data strategy has converged on a layered corpus design that explicitly weights classical, Modern Standard, and dialectal Arabic against the intended downstream applications. The training corpus is also explicitly curated against Saudi cultural and religious sensibilities — content that conflicts with the consensus interpretation of religious scholars, content that violates legal constraints on certain topics, and content that is structurally unreliable is filtered or down-weighted. This is a more aggressive curation posture than is conventional in Western foundation models, and it produces a model with a measurably different output distribution.
Fine-Tuning Pipelines
Allam is positioned as a foundation that downstream enterprises and government entities fine-tune for their specific applications. SDAIA operates a fine-tuning pipeline that is exposed to authorized integrators — primarily the major Saudi systems integrators, the cloud regions operated by hyperscalers under sovereign-cloud agreements, and a small set of strategic partners. The fine-tuning pipeline supports both supervised fine-tuning on labeled task data and reinforcement learning from human feedback for alignment refinement.
The fine-tuning posture distinguishes between three tiers of customization. The first is prompt-engineering and retrieval augmentation, which is available to any authorized user. The second is parameter-efficient fine-tuning (LoRA-class techniques) on top of the base Allam weights, which is available to enterprise customers under standard licensing. The third is full fine-tuning with weight modification, which requires explicit SDAIA authorization and is reserved for strategic deployments such as defense, healthcare, and major government services.
Evaluation Infrastructure
Evaluation is the under-appreciated half of foundation-model development, and the Saudi ecosystem has been building Arabic-specific evaluation infrastructure that goes substantially beyond what is available in Western Arabic NLP communities. The evaluation suites span Modern Standard Arabic comprehension, dialect handling across the major regional dialects, classical-Arabic and religious-text fidelity, code generation in Arabic-language contexts, mathematical reasoning, multilingual transfer, and a growing set of domain-specific evaluations for healthcare, legal, financial, and government workflows.
The evaluation infrastructure is operated as a shared resource through SDAIA, with academic partners at KAUST, KFUPM, King Saud University, and Princess Nourah bint Abdulrahman University contributing benchmarks. Allam’s reported evaluation performance has improved generation-over-generation on the Arabic benchmarks, with the most recent generation positioned competitively against the Arabic performance of the leading Western foundation models. The strategic objective through 2030 is for Allam to be the dominant Arabic foundation model not only in Saudi Arabia but across the Arabic-speaking world.
The SDAIA-Led Ecosystem
SDAIA, the Saudi Data and Artificial Intelligence Authority, was established in 2019 to consolidate the Kingdom’s AI strategy under a single entity. It includes the National Center for AI as its principal R&D arm and the National Data Management Office as its data-governance arm. SDAIA reports to the Council of Ministers and operates with substantial operational autonomy across a budget that has scaled significantly since its founding. Its mandate spans regulatory authority for the PDPL, foundation-model development, talent development, ecosystem cultivation, and international cooperation.
The ecosystem that SDAIA has cultivated includes the major Saudi universities, the principal hospital systems, the leading systems integrators, and a growing set of sovereign-AI-aligned international partners. The collaboration model differs from the conventional academic-industry-government model in that SDAIA itself operates infrastructure — including a substantial share of the Kingdom’s training compute — and acts as both a regulator and an operator. This dual role is unusual in the global AI ecosystem and reflects the Saudi posture that foundation-model capability is a sovereign function rather than a market function.
Vendor Selection Criteria and Compute Procurement
Vendors selling into the Arabic LLM training ecosystem are evaluated against criteria that emphasize sovereignty alignment, infrastructure scale, and willingness to localize. Compute vendors must commit to Kingdom-resident capacity at the scale required, with appropriate licenses and end-use commitments to navigate US export-control constraints. Software-stack vendors must support the heterogeneous accelerator portfolio that SDAIA and Humain have assembled. Talent and consulting partners must commit to Saudi-resident engineering footprints rather than treating the engagement as a remote-delivery opportunity.
The principal compute procurements have been the Humain partnerships with NVIDIA (a multi-year, multi-billion-dollar agreement for advanced accelerator capacity), with AMD (a comparable agreement for MI-class capacity), and with Qualcomm (a partnership for inference and edge capacity). Cisco, Supermicro, Dell, HPE, and Lenovo are the principal server-class hardware partners. The cooling and power infrastructure providers — Vertiv, Schneider, ABB, Stulz — are participating at scale on the data-center build-out that supports the training capacity.
Common Pitfalls
The two principal pitfalls in Arabic LLM training are the data pitfall and the alignment pitfall. The data pitfall is treating Arabic as a single homogeneous language and either over-fitting to Modern Standard Arabic or under-weighting classical and religious content. The alignment pitfall is applying Western alignment recipes — particularly Western RLHF preference data — to a model that is intended for a Saudi audience, which produces outputs that are subtly misaligned with the cultural and religious norms the model is meant to serve. Vendors that build with explicit awareness of these pitfalls, and that incorporate Saudi review into the development loop, produce outputs that are noticeably better fit for purpose.
Talent Pipeline and the Academic-Industry Interface
The talent pipeline for Arabic LLM training is one of the harder strategic problems SDAIA faces, and the response has been a multi-pronged investment that combines aggressive international recruitment, substantial returnee outreach to Saudi technical leaders working in international AI hubs, and a deepened academic-industry interface that channels students from KAUST, KFUPM, King Saud University, Princess Nourah bint Abdulrahman University, and Prince Sultan University into the SDAIA ecosystem. The Tuwaiq Academy, a SDAIA-affiliated technical training institution, operates intensive bootcamp-style programs that re-skill mid-career engineers into AI-engineering roles at the pace required by the program.
The KAUST AI Initiative, anchored by the new generation of AI-research leadership at KAUST, has been positioned as the principal academic anchor for frontier AI research in the Kingdom, with substantial investment in compute, in faculty recruitment, and in the graduate-program pipeline. The KFUPM AI program has a more applied orientation, with strong ties to Aramco and SABIC and a focus on industrial-AI applications. Together with the broader Saudi university system, these institutions form the academic substrate for the long-term capability development that the foundation-model program requires.
Open-Source Posture and the Broader Ecosystem Strategy
SDAIA has progressively moved toward a more open-source-aligned posture for the Allam family, with selective release of model weights, training data, and evaluation infrastructure under licenses that permit broader academic and commercial use. The strategic logic is that an open-source posture accelerates ecosystem development, attracts external contributions and scrutiny, and positions Allam as the default Arabic foundation-model substrate for the broader Arab world. The release decisions are made carefully, with attention to security, sovereignty, and commercial considerations, but the trajectory is toward greater openness rather than greater restriction.
The broader ecosystem strategy includes partnerships with the leading open-source-aligned AI vendors (Hugging Face, Mistral, and the broader open-source LLM community), with the principal cloud providers operating Saudi sovereign regions, and with a growing set of regional AI startups that build on the Allam substrate. The combined ecosystem effort is one of the more sophisticated national AI ecosystem-development programs anywhere in the world, and it is being explicitly positioned as a model for the broader sovereign-AI movement that has been emerging across multiple national contexts since 2023.
Inference Infrastructure and the Production-Serving Layer
Foundation-model training is one half of the Arabic LLM agenda; the production inference-serving layer is the other half, and it has its own substantial infrastructure requirements. The inference layer must serve a wide range of latency, throughput, and cost profiles — from low-latency consumer chat interactions to high-throughput batch document processing to ultra-low-latency edge applications in mobility and IoT contexts. The architecture being assembled by SDAIA, by Humain, and by the principal cloud providers operating Saudi sovereign regions combines centralized GPU-accelerated inference for the most capable model variants with distilled smaller models served across a more distributed footprint for cost-and-latency optimization.
The Qualcomm partnership announced under the Humain umbrella is particularly relevant to the inference layer, given Qualcomm’s positioning in efficient-inference accelerator capacity. The combined inference infrastructure is being designed to serve a domestic Allam-based application ecosystem at substantial scale by 2027-2028, with sufficient capacity to absorb the consumer-and-enterprise application load that the Saudi market is expected to generate. The economics of the inference layer — cost per token, cost per query, latency under peak load — are being actively monitored against the broader hyperscaler-provided alternatives, and the SDAIA-led ecosystem has been progressively achieving competitive economics for the workload categories where sovereignty matters.
For deeper reading: see SDAIA and the Allam program, Hexagon data center, Humain compute partnerships, and Arabic foundation-model evaluation.