The Data Layer Under the AI Stack

Databricks occupies a position in Saudi Arabia’s AI buildout that is qualitatively different from the hyperscalers. Microsoft, AWS, and Google Cloud are competing to be the cloud infrastructure layer — the compute, storage, and network fabric on which everything runs. Databricks is competing to be the data platform layer — the software stack that ingests, organizes, governs, and makes available for AI training and inference the data that ultimately determines what AI can do.

This distinction matters for how Databricks should be understood in the Saudi context. The hyperscalers are selling infrastructure; Databricks is selling the capability to make that infrastructure useful. A Saudi government agency with 10 petabytes of raw administrative data running on AWS Riyadh does not automatically have an AI-ready data asset — it has a data management problem. Databricks’ value proposition is solving that data management problem, making the raw data actionable for the machine learning and AI workloads that ultimately justify the infrastructure investment.

In markets where AI ambitions are high and data infrastructure is less mature — which characterizes much of Saudi Arabia’s government and enterprise data landscape in 2025 — the data platform layer may be the single most important technology investment. You can buy NVIDIA GPUs in bulk and deploy them in hyperscaler cloud regions, but if your underlying data is disorganized, ungoverned, inconsistently formatted, and inaccessible to your AI engineers, the compute investment is largely wasted. Databricks’ message to Saudi decision-makers is that the data platform investment is the prerequisite that makes the AI ambition achievable.

The SDAIA Partnership: Saudi Arabia’s AI Authority as Strategic Customer

Databricks’ partnership with SDAIA (Saudi Data and AI Authority) is the most significant indicator of its Saudi position. SDAIA operates the National Data Bank — over 430 integrated government systems containing the Saudi government’s most strategically valuable data assets — and is the developer of Allam, Saudi Arabia’s national Arabic large language model. The Databricks-SDAIA engagement is reported to involve a SR1.88 billion (~$500 million) contract, which would make it one of the largest enterprise software deals in Saudi Arabia and one of Databricks’ largest single-customer engagements globally.

To understand what a $500 million Databricks engagement with SDAIA covers, consider what SDAIA actually needs from a data platform. The National Data Bank’s 430+ integrated systems produce data in different formats, from different ministerial IT environments, with different quality levels, different metadata schemas, and different access control requirements. Before any of that data can be used for Allam model training or fine-tuning, it needs to be harmonized: standardized into consistent formats, deduplicated, validated for quality, catalogued with appropriate metadata, and made accessible to AI training pipelines with appropriate access controls that reflect the sensitivity of different government data categories.

Databricks’ lakehouse architecture — combining the low-cost scalable storage of a data lake with the transactional query capabilities and governance features of a data warehouse — is designed precisely for this challenge. The lakehouse can ingest raw data from heterogeneous sources, apply Databricks’ data transformation capabilities (built on the open-source Apache Spark engine that Databricks created) to harmonize it, and surface it through a governed catalog layer that enforces access controls and data lineage tracking.

Unity Catalog: The Data Governance Layer for Saudi Enterprise

Databricks Unity Catalog is the governance and metadata management layer that makes Databricks’ lakehouse architecture enterprise-ready for regulated environments. Unity Catalog provides a unified governance framework across all data assets in a Databricks environment: who can access what data, under what conditions, with full audit logging of every access event, and with data lineage tracking that shows where data came from, how it was transformed, and where it flowed.

In the Saudi context, Unity Catalog’s governance capabilities are directly relevant to the PDPL compliance requirements that govern personal data handling, and to the KSA-RoD data residency requirements that restrict where government data can be stored and processed. A Saudi government agency using Databricks with Unity Catalog can demonstrate compliance with both frameworks: Unity Catalog’s audit logs provide the evidentiary record needed for PDPL compliance audits, and the deployment on Saudi-located cloud infrastructure (AWS Riyadh, Azure Saudi Arabia, or Google Cloud Saudi) satisfies the data residency requirements.

For SDAIA’s National Data Bank specifically, Unity Catalog addresses a governance challenge of unusual complexity. The National Data Bank contains data from 430+ systems across dozens of ministries, each with different classification levels, different permitted use cases, and different data sharing agreements between ministries. Managing this governance complexity manually is not practical at scale; Unity Catalog’s programmatic governance framework — where access policies are defined in code and enforced automatically — is the practical solution for managing a government data asset of this scale and sensitivity.

Mosaic AI: Fine-Tuning Saudi Models at Scale

Databricks acquired MosaicML in 2023 for approximately $1.3 billion, integrating MosaicML’s capabilities for efficient large language model training into the Databricks platform under the Mosaic AI brand. MosaicML’s core intellectual property was in efficient model training: the algorithmic techniques, training infrastructure optimization, and tooling that allow organizations to train large models at lower cost and with greater efficiency than standard approaches.

For Saudi AI development — and specifically for Allam development and fine-tuning — Mosaic AI’s capabilities are directly relevant. Allam is a 34 billion parameter model trained on 8 petabytes of Arabic-weighted data. Training and fine-tuning a model of that scale requires both the raw compute (handled by SDAIA’s Blackwell GPU cluster and Humain’s infrastructure) and the training orchestration software that manages distributed training across thousands of GPUs efficiently. Mosaic AI provides the training orchestration layer: job scheduling, gradient checkpointing, fault tolerance, and the hyperparameter optimization tools that make large-scale training runs tractable.

The fine-tuning use case is perhaps even more commercially significant than initial training. Once Allam’s base model is trained, Saudi government agencies and enterprises will want domain-specific versions of Allam fine-tuned on their specific data: a fine-tuned Allam for legal document analysis, a fine-tuned version for medical records processing, an Allam variant specialized for oil and gas technical documentation. Each of these fine-tuning projects requires the same infrastructure — compute, training orchestration, data pipelines — that Mosaic AI provides. Databricks’ ability to offer Mosaic AI as the fine-tuning platform for Allam variants makes it the natural tool for the ecosystem of Allam-derived specialized models that will likely develop over the coming years.

Databricks’ Open-Source Heritage and Saudi Engineering Adoption

Apache Spark — the distributed computing framework that Databricks commercialized — is one of the most widely deployed open-source data processing technologies in the world. Saudi enterprises that have built data infrastructure over the past decade have in many cases built it on Spark: it is the standard choice for large-scale batch data processing in enterprises globally, and the Saudi engineering talent pool that has worked at international technology companies or studied abroad is familiar with the Spark ecosystem.

This open-source familiarity creates a lower adoption friction for Databricks relative to competing platforms that require learning new paradigms. A Saudi data engineering team that has worked with Spark on AWS EMR or Azure HDInsight can typically migrate to Databricks’ managed Spark environment with minimal retraining, gaining the governance and collaboration features that Databricks adds on top of Spark without abandoning the technical skills they have already developed.

The open-source strategy also affects the build-vs-buy decision for Saudi organizations. Databricks customers can extract their data from Databricks in open formats (Delta Lake, Parquet) at any time — there is no vendor lock-in through proprietary data formats. For Saudi government customers who are sensitive to technology dependency on foreign vendors, this portability is a significant factor. Databricks’ position is that the value it provides is in the software, tooling, and services, not in lock-in — and that position aligns well with Saudi data sovereignty concerns.

Competing with Snowflake: The Saudi Data Platform Battle

Databricks’ primary direct competitor in the enterprise data platform market is Snowflake. Both companies offer cloud-native data platforms for analytics and AI, both have significant enterprise customer bases, and both are actively pursuing Saudi Arabia as a high-priority market. The Databricks-Snowflake competition in Saudi Arabia parallels their global competition, with a few Saudi-specific dimensions.

Databricks’ strength is in the machine learning and AI workload dimension: organizations that want to train models, run experiments, and build ML pipelines are better served by Databricks’ data science and ML engineering tooling than by Snowflake’s. Snowflake’s strength is in the SQL analytics dimension: organizations that primarily need business intelligence, reporting, and data warehousing are well-served by Snowflake’s SQL-first architecture.

Saudi Arabia’s AI ambitions skew the competition toward Databricks. The National Data Bank, Allam training, fine-tuning pipelines, enterprise AI applications — these are all ML workloads where Databricks’ data science and ML platform capabilities are more relevant than Snowflake’s SQL analytics strengths. The SDAIA-Databricks partnership implicitly positions Databricks as the preferred data platform for Saudi sovereign AI development, which is a significant reference that Databricks can use in conversations with Saudi private sector enterprises considering data platform investments.

Snowflake has its own Saudi presence and partnerships, but the SDAIA relationship and the Allam connection give Databricks a differentiated narrative in the Saudi market: Databricks is not just a data platform company in Saudi Arabia, it is the platform behind Saudi Arabia’s national AI model. That narrative has significant sales and marketing value in a market where buyers want to align their technology choices with the national AI strategy.

The Commercial Model: Layered on Hyperscalers

Databricks does not operate its own cloud infrastructure. It runs on top of AWS, Azure, and Google Cloud — Databricks clusters consume compute and storage from the underlying hyperscaler, and Databricks charges for the software layer on top. This multi-cloud positioning is a commercial advantage in the Saudi market: Databricks is not competing with the hyperscalers for infrastructure revenue, so the hyperscalers have an incentive to partner with Databricks and resell Databricks services to their own customers rather than competing against them.

In practice, AWS, Azure, and Google Cloud all have Databricks listings in their cloud marketplaces, allowing Saudi enterprise customers to procure Databricks through their existing hyperscaler billing relationships and enterprise discount agreements. For a Saudi company that has a significant AWS commitment and wants to add Databricks, the procurement pathway through AWS Marketplace is simpler and potentially more cost-effective than a direct Databricks contract. This distribution model gives Databricks access to hyperscaler enterprise sales motions that it could not build independently.

The multi-cloud positioning also aligns with Saudi enterprise infrastructure strategies. Many large Saudi organizations — Saudi Aramco, the major banks, government agencies — are pursuing multi-cloud strategies that use different hyperscalers for different workloads based on cost, compliance, and capability considerations. A data platform that runs consistently across AWS, Azure, and GCP is more valuable to these multi-cloud organizations than one locked to a single hyperscaler.

Databricks as a Humain-Tier Saudi Partner

The significance of Databricks being a named partner in the Humain-adjacent Saudi AI ecosystem — alongside hyperscalers that are orders of magnitude larger in revenue and market capitalization — reflects the specific recognition in Saudi AI strategy that the data platform layer is foundational, not peripheral.

Humain’s infrastructure is meaningless without data to train models on and to build AI applications around. SDAIA’s National Data Bank is meaningless without the software stack to govern, transform, and make it usable. NVIDIA’s GPUs produce no value without training data pipelines flowing into them. Databricks is the connective tissue between the data assets, the governance frameworks, and the compute infrastructure that together constitute the Saudi AI stack. Its position as a named partner at the strategic level — rather than a background technology component — reflects that this is understood by Saudi AI decision-makers at the highest levels.

The $500 million SDAIA engagement, if it proceeds as reported, would represent one of the largest enterprise software engagements in Saudi history and a transformative contract for Databricks’ Middle East revenue. It positions Databricks not merely as a technology vendor but as a strategic infrastructure partner in Saudi Arabia’s AI sovereignty project — a positioning that carries relationship capital, reference power, and follow-on revenue potential well beyond the initial contract value.