The Source Corpus
The source layer is the foundation under every analytical claim on saudicompute.com. The platform’s editorial standard is that every fact has a source, every source is publicly inspectable, and every reader who wants to audit a claim can trace it back to the underlying disclosure. The source corpus — currently 100+ tracked domains, 3,751+ primary files, 307 PDFs, and a continuously growing newsroom-mirror archive — is the material that the analytical layer rests on.
The source layer is more than a citation list. It is a structured corpus with classification taxonomy, provenance metadata, capture timestamps, and version control. When a claim’s underlying source updates, the platform’s downstream content surfaces the change. When a source is retracted or amended, the platform’s records reflect the amendment. When a new source class becomes available — a newly published SDAIA decree, a fresh PIF annual report, a new BIS license action — the platform’s ingestion pipeline picks it up on the next refresh cycle.
Source Taxonomy
The platform classifies sources into seven primary categories.
Government primary sources include SDAIA, PIF, MCIT, NEOM, the Saudi Press Agency, the Saudi gazette, Vision 2030 publications, the National Information Center, KACST, KAUST, KFUPM, the Ministry of Energy, the Ministry of Industry and Mineral Resources, the Ministry of Investment, and the broader official Saudi domain set. These are first-class sources for any claim about Saudi policy, strategy, or institutional architecture.
Regulatory and supervisory sources include the Communications, Space and Technology Commission (CST), the Saudi Central Bank (SAMA), the Capital Market Authority (CMA), the Ministry of Health (MoH), the Personal Data Protection Authority components inside SDAIA, and the Economic Cities and Special Zones Authority (ECZA) that administers the Cloud SEZ.
US-side regulatory sources include the Bureau of Industry and Security (BIS) license publications, the Federal Register notices, the SEC filings of US-listed counterparties, the Treasury OFAC publications where they intersect Saudi or counterparty-relevant jurisdictions, and the broader US Commerce, State, and Defense Department publications that surface bilateral commercial and security architecture.
Partner-corporate primary sources include NVIDIA, AMD, Qualcomm, Groq, Intel, SambaNova, AWS (Amazon), Microsoft Azure, Google Cloud (Alphabet), OCI (Oracle), IBM, Cisco, Schneider Electric, ABB, Vertiv, ACWA Power, Saudi Electricity Company, Aramco, SABIC, STC, Center3, and the broader operating-partner roster. The platform mirrors press releases, earnings transcripts, investor-day materials, and SEC filings from this set.
Conference and industry sources include LEAP, FII, the US-Saudi Investment Forum, the World Economic Forum, the GITEX/GISEC events in the broader Gulf cycle, and the academic and industry-conference publication channels (NeurIPS, ICML, MLSys, OCP) where Saudi-relevant disclosures occasionally surface.
Verified-press sources include the FT, Bloomberg, Reuters, the Wall Street Journal, the Information, Semafor, the Saudi-domiciled English-language press (Arab News, Saudi Gazette), and the regional press (The National, Al Arabiya English) that cover the buildout. The platform’s editorial standard is that press claims are weighted by the reporting depth of the underlying article and are cross-referenced to primary sources where possible.
Academic and analytical sources include peer-reviewed publications, sovereign-wealth-fund research notes, sell-side equity research where it surfaces Saudi-relevant claims, and the broader analytical literature on AI infrastructure, sovereign AI, and Gulf economic policy.
Source Vetting
The platform applies a four-step vetting process to every source.
First, primary-versus-secondary classification. Primary sources (a SDAIA decree, a PIF annual report, a NVIDIA earnings transcript) are weighted higher than secondary sources (a press article reporting on a primary disclosure). The platform’s editorial preference is to surface the primary source directly where possible.
Second, provenance verification. Every source URL is verified at capture time and re-verified on a rolling cadence. URL drift, link rot, and source retraction are tracked at the source-record level. Where a primary source disappears from its original domain, the platform’s archive layer preserves the original capture.
Third, cross-referencing. Material claims are cross-referenced to at least two independent sources where possible. A single-source claim that cannot be cross-referenced is flagged as such in the underlying data layer; the analytical layer surfaces the flag implicitly through phrasing that makes the single-source posture clear.
Fourth, recency and version control. Sources have publication dates and, where applicable, version numbers (PIF annual report 2023, SDAIA strategy v2). The platform tracks the prevailing version and surfaces version transitions when they update.
High-Volume Source Domains
A small number of source domains dominate the corpus by volume. SDAIA’s primary domain, PIF’s investor-relations site, NVIDIA’s newsroom, the AWS-Saudi-region announcements page, the MCIT publication channel, the NEOM corporate site, and the broader official-Saudi domain set together account for the largest share of primary-source records. The partner-newsroom mirrors (NVIDIA, AMD, Qualcomm, Microsoft, Google, AWS) account for the largest share of capture volume on the corporate side.
The platform’s ingestion pipeline is calibrated to those high-volume domains: it polls them at a higher cadence, applies more sophisticated parsing logic to handle their site architectures, and ingests their structured outputs (RSS, sitemap, JSON-LD) at the schema level where they expose them.
How the Source Layer Surfaces
Every analytical page on saudicompute.com surfaces the source layer in three ways. First, in-page citation: claims are attributed inline where the source is operationally specific. Second, the per-record source URL: the underlying data layer carries source URLs at the record level, surfaced through the entity profile, deal flow, and capacity tracker pages. Third, the sources directory itself, which is this section, and which lists every tracked domain with extraction stats and topical coverage.
Subscribers at the institutional and enterprise tiers of the Sovereign Compute Terminal can query the source layer programmatically, retrieving the underlying URLs and metadata for any record set they specify.
Independence and Source Bias
The platform applies the same editorial discipline to every source category. Saudi government sources are not weighted higher because they are official; US government sources are not weighted higher because they are American; partner-corporate sources are not weighted higher because they are from public-listed counterparties. The vetting process is uniform across categories, and the analytical layer applies the same skepticism — particularly the announced-versus-delivered analytical lens — to claims from every source class.
The independence posture matters because the source layer is the foundation under every analytical claim. A platform whose source weighting was biased toward any single category would produce analytical outputs that institutional readers could not rely on. The platform’s commitment is uniform editorial discipline across categories.
Submitting Sources
Readers can submit sources through the platform’s contact layer. Submissions are evaluated against the vetting process; primary sources that are not yet in the corpus, or secondary sources that surface analytical angles the platform has not yet covered, are integrated into the next refresh cycle. The platform welcomes correction submissions — including from tracked entities — and integrates corrections into the underlying data layer where they meet the editorial standard.
Provenance Metadata
Every source in the corpus carries a provenance metadata record. The record includes the source URL at the time of capture, the capture timestamp, the source classification (which of the seven primary categories the source falls into), the disclosing entity (where applicable), the publication date, the version number (where applicable), and the language. The metadata flows through to the analytical layer, where each in-page citation is traceable to a specific provenance record.
The provenance discipline is particularly important for primary sources that are subject to revision. SDAIA strategy publications, PDPL implementing regulations, BIS license actions, and PIF annual reports are examples of source classes where the version-and-revision history matters analytically. The provenance metadata captures the prevailing version at any point in time, so that historical analytical claims remain traceable to the source as it existed at the time of the claim.
High-Risk Source Classes
Several source classes carry structural risk that the platform’s editorial discipline addresses through specific weighting and cross-referencing rules. Single-source corporate announcements that have not been cross-referenced to a primary regulatory filing carry a higher confidence-discount than dual-sourced disclosures. Conference materials that surface deal claims without a corresponding regulatory or partner-corporate disclosure carry a higher confidence-discount than corresponding regulatory filings. Press claims attributed to anonymous sources carry the highest confidence-discount and are typically held until they are corroborated by primary disclosure.
The platform’s editorial layer surfaces the confidence posture for each analytical claim. Claims with primary-source corroboration are surfaced with definitive phrasing; claims with single-source posture or anonymous attribution are surfaced with phrasing that makes the source-class explicit (“according to a partner-corporate disclosure,” “as reported by a single press source”). The discipline is uniform across topics and source classes.
The Documents and Tables Subsets
The sources layer is the parent corpus from which the documents (PDFs) and tables (structured-data extracts) layers are derived. Documents are the PDF subset of the source corpus, with their own indexing and analytical layer at /documents/. Tables are the structured-data subset extracted from documents and HTML sources, with their own analytical layer at /tables/. The sources layer is the broadest of the three; documents and tables are the more-structured subsets surfaced for specific analytical purposes.
The relationship between the three layers matters for how readers navigate the platform. A reader looking for the prose context of a claim navigates through the sources layer; a reader looking for the structured-numeric context navigates through the tables layer; a reader looking for the canonical-document context navigates through the documents layer. All three converge at the per-record provenance metadata that ties them back to the underlying source.
Source Domain Statistics
The 100+ tracked source domains are not equally weighted in the corpus. The high-volume domains — SDAIA’s primary site, the PIF investor-relations channel, the partner-newsroom set, the Saudi Press Agency, the broader official-Saudi domain set — together account for the largest share of primary-source records. The mid-volume domains include the broader regulatory-publication channels, the major partner-corporate sites, and the conference-publication channels. The long-tail of low-volume domains includes specialized regulatory channels, smaller partner-corporate disclosures, and the broader supporting-source set.
The platform’s source-statistics layer surfaces the per-domain extraction count, the per-domain capture cadence, and the per-domain content-class distribution. Subscribers and readers can audit the source-corpus composition through the statistics layer to understand which domains are driving the analytical layer for any given topic.
Source Discovery and Onboarding
New source domains are discovered through three primary channels. The first is the continuous monitoring of cross-references inside the existing corpus: when a tracked source cites or refers to a domain that is not yet in the corpus, the new domain is evaluated for inclusion. The second is reader and subscriber submissions: external parties surface domains that the platform’s editorial layer has not yet captured. The third is the editorial team’s continuous environmental scanning: regulatory publication channels, partner-corporate disclosures, and emerging analytical-output channels that are released into the broader information ecosystem.
Source onboarding follows a structured process: domain evaluation against the source-classification taxonomy; capture-pattern engineering to handle the domain’s specific publishing architecture; ingestion-pipeline integration so that the new domain’s content flows into the underlying corpus on the standard cadence; and per-domain quality validation to ensure the captured content meets the platform’s editorial standard.
Source Lifecycle and Retirement
Sources have lifecycles. New sources are added as they emerge; existing sources are maintained as long as they continue to publish content relevant to the Saudi compute story; sources are retired from active capture when they cease publishing or when their content drifts away from the platform’s editorial scope. Retired sources retain their historical capture in the corpus; the historical content remains searchable and citable even after active capture ends.
The lifecycle discipline matters because the platform’s source corpus is itself a historical artifact. Twenty years from now, the corpus will include domains that have stopped publishing, organizations that have ceased to exist, and analytical channels that are no longer operationally active. The platform’s commitment is to preserve those historical captures so that the corpus remains useful for retrospective analytical work even as the prevailing source landscape evolves.
For deeper reading:
- Documents — the PDF subset of the source corpus
- Tables — the structured-data subset extracted from sources
- About — platform mission and editorial standards
- Methodology — how sources flow into the SCS framework