The Structured Data Layer

The tables index is the structured-data subset of the saudicompute.com source corpus. It comprises 5,194 extracted tables drawn from PDFs and web sources, normalized into a unified schema, and surfaced through the platform’s analytical layer. The tables layer is the quantitative spine of the platform: where the analytical writing carries the qualitative narrative, the tables carry the numbers that anchor the narrative, and every claim that involves a quantitative datum is ultimately traceable to a row in the tables layer.

Tables are different from prose source content in three ways. First, they are inherently structured: rows have schema, columns have types, and the relationships between cells are explicit rather than implied. Second, they are computationally tractable: tables can be queried, joined, filtered, and aggregated, which the prose layer cannot. Third, they carry density that prose cannot match: a single financial-statement table can contain the equivalent quantitative content of dozens of prose paragraphs.

The platform’s tables layer therefore does double duty. It is both a directly-consumable analytical product (subscribers at the institutional and enterprise tiers of the Sovereign Compute Terminal query the tables layer programmatically) and the underlying numeric foundation under the platform’s analytical writing.

Table Categories

The platform’s 5,194 tables organize around six primary categories.

Financial statement tables include income statements, balance sheets, cash-flow statements, and segment-disclosure tables drawn from PIF annual reports, Aramco filings, SABIC filings, and the publicly-listed counterparty filings (NVIDIA, AMD, Qualcomm, Microsoft, Alphabet, Amazon, Oracle). These are the quantitative source for the capital section’s analytical claims, the per-entity AUM and capex disclosures, and the announced-to-deployed ratio analytics.

Capacity tables include the data-center IT-load tables, the per-campus capacity disclosures, the GPU-equivalent count tables, and the power-architecture tables (megawatts of generation contracted, megawatts of cooling capacity, PUE disclosures). These feed the infrastructure section’s analytical layer and the SCS Capacity component.

Partnership and counterparty matrices include the deal-flow tables, the counterparty-pair matrices that surface in the intersections section, and the silicon-vendor allocation tables (e.g., the GB300 allocation under the May 2025 NVIDIA framework). These are the quantitative spine of the deal flow tracker and the relationship-mapping analytical layer.

Regulatory schedule tables include the SDAIA decree publication schedule, the BIS license action tables, the AI Diffusion Framework’s tiered-country structure, the PDPL implementation schedule, and the Cloud SEZ regulatory tables. These feed the policy and silicon sections.

Macroeconomic and demographic tables include the Saudi GDP composition tables, the non-oil-GDP trajectory, the population and demographic tables, the energy-mix tables (gas, solar, nuclear, hydrogen), and the broader Vision 2030 KPI tables. These feed the geopolitics section and the contextual layer beneath the more compute-specific analytical pages.

Conference and disclosure tables include the LEAP and FII deal-disclosure summary tables, the US-Saudi Investment Forum communique tables, the partner-newsroom announcement tables, and the broader event-driven structured-data subset.

How Tables Flow Through the Platform

Tables enter the platform through the documents and web ingestion pipelines. Tables embedded in PDFs are extracted through layout-aware OCR and table-recognition processing; tables embedded in HTML are parsed through HTML-aware extraction. Each extracted table is normalized into a common schema (table identifier, source document, source page, table title, column headers, row data, source URL, capture timestamp) and ingested into the platform’s underlying database.

The normalized tables flow into three analytical paths. First, the tables index — this section — surfaces the corpus to readers and subscribers as a queryable resource. Second, the inline-surfacing layer pulls relevant tables into the per-entity, per-section, and per-deal pages where the table content is operationally relevant. Third, the SCS pipeline pulls quantitative inputs from the tables layer (capacity, capex, ownership classification) into the seven-component scoring framework.

The platform’s editorial layer is designed to make the tables layer discoverable rather than to replace it. When the analytical writing carries a quantitative claim, the underlying table is one click away through the in-page link or the entity-profile data layer.

Notable Tables

A small number of tables are referenced repeatedly across the analytical layer.

The PIF AUM allocation table — drawn from successive PIF annual reports — gives the per-asset-class composition of the fund’s $930 billion AUM and is the quantitative anchor for the capital section’s analytical claims about the AI-attributable subset of the fund.

The Saudi data-center capacity table aggregates IT load across the 6.6 GW of disclosed Saudi-resident campus capacity, with per-campus operational status, accelerator footprint, and sovereignty posture. It is the quantitative anchor of the infrastructure section.

The May 2025 NVIDIA framework allocation table — drawn from the US-Saudi Investment Forum communique materials and partner disclosures — gives the per-tranche GB300 allocation across the framework’s window and is the quantitative anchor of the silicon section.

The SDAIA decree publication schedule, drawn from the Saudi gazette and SDAIA primary publications, gives the per-decree publication and effective dates and is the quantitative anchor of the policy section’s chronology.

The AI Diffusion Framework tier table, drawn from the BIS Federal Register publications, gives the per-country tier classification and is the quantitative anchor for the geopolitics and silicon sections.

Programmatic Access

Subscribers at the institutional and enterprise tiers of the Sovereign Compute Terminal can query the tables layer programmatically. The schema is documented; the data model is stable; and the access pattern is structured query against the platform’s underlying database, with rate-limit and snapshotting controls.

The programmatic-access layer is the platform’s single most-requested institutional feature. Sovereign-wealth allocators, hyperscaler business-development desks, semiconductor strategy teams, and policy-research functions integrate the tables layer into their proprietary analytical workflows; the platform’s role is to be the highest-quality public-information data layer that those workflows consume.

Editorial Standards in the Tables Layer

The tables layer applies the same editorial discipline as the prose layer. Every table has a source URL. Every cell has a provenance trail back to the underlying document or web page. Updates and amendments at the source propagate into the tables layer through the next ingestion cycle. Where a table’s underlying source has been retracted or amended, the platform’s records preserve the original alongside the amendment.

The platform does not synthesize or fabricate tables. Tables are extracted from disclosed sources or are computed transparently from extracted inputs (the SCS aggregate scores, the capacity-aggregation tables, the announced-to-deployed ratio tables). Where a table is computed rather than extracted, the computation is documented and the underlying inputs are traceable.

Use Patterns

The tables layer supports four primary use patterns.

Reference look-up: a reader who wants to verify a specific quantitative claim on the platform can trace it through the in-page citation to the underlying table.

Comparative quantitative analysis: a reader running comparative work — say, comparing PIF and Mubadala AUM trajectories, or comparing Saudi and UAE data-center capacity — can pull the relevant tables and run the comparison directly.

Time-series analysis: a reader doing temporal work can pull the table series across publication dates and run the trajectory analysis directly.

Programmatic integration: institutional subscribers integrate the tables layer into proprietary analytical workflows through the Terminal’s programmatic-access tier.

Cadence

The tables layer is updated continuously through the ingestion pipeline. Material refreshes (a new PIF annual report, a new SDAIA strategy publication, a new BIS license action surfacing in the Federal Register) trigger immediate downstream propagation into the tables index and into the analytical layer that consumes it. The platform’s commitment is that the tables layer is always within a small number of days of the prevailing public-disclosure frontier.

Schema and Normalization

The tables layer applies a consistent schema across categories. Each table record carries a stable identifier, a source document or web page reference, a source page or section reference, the table title (from the original or, where the original lacks a title, an extraction-time heuristic title), the column headers, the typed row data, the source URL, the capture timestamp, and the provenance-metadata reference. The schema is documented in the API layer and surfaces in the per-table inspector view.

Normalization is the most analytically consequential aspect of the tables layer. Original tables in source documents use inconsistent units (USD, SAR, billions, millions), inconsistent column-headers (capacity, IT load, MW, gigawatts), and inconsistent row-organization conventions. The platform’s ingestion pipeline applies normalization rules at extraction time: monetary values are converted to USD with the prevailing exchange rate (the SAR is dollar-pegged at 3.75); capacity values are converted to megawatts of IT load; column headers are mapped to a canonical taxonomy; row data is typed.

The normalized records flow into the analytical layer with consistent units and consistent typing. Where the normalization is non-trivial (e.g., converting a partner-corporate revenue disclosure into the Saudi-attributable subset), the normalization rule is documented and the original record is retained alongside the normalized record. The principle is that the original disclosure is always available to subscribers who need to audit the normalization.

The Computed-Tables Layer

A subset of the tables layer is computed rather than extracted. The SCS aggregate-and-component score table is computed from the underlying entity-input attributes through the methodology framework. The capacity-aggregation table is computed from the per-campus extracted records. The announced-to-deployed ratio table is computed from the operational-status records. The intersection-counts table is computed from the entity-mention corpus.

Computed tables are tagged as such in the schema and the computation rules are documented at the methodology level. The principle is that any computed table is reproducible from the underlying inputs through the documented computation rule. Subscribers who want to audit a computed table can pull the inputs and run the computation independently; the platform’s computation matches the independent reproduction.

Edge Cases

Several recurring edge cases occur in the tables layer. Tables with merged cells in the original document are handled by row-and-column expansion, with the merged-cell content propagated to each constituent cell in the expanded form. Tables that span multiple pages in the original PDF are stitched at extraction time into a single record, with the page-spanning metadata preserved for traceability. Tables with footnoted cells carry the footnote content as a separate metadata field attached to the cell.

The platform’s editorial layer surfaces the edge-case treatments where they affect interpretation. Where a table’s interpretation depends on a footnote, the analytical writing surfaces the footnote alongside the headline number. Where a table’s normalization is non-trivial, the writing surfaces the normalization rule.

How Subscribers Build with the Tables Layer

The most common subscriber pattern is to integrate the tables layer into a proprietary analytical pipeline. A sovereign-wealth-fund allocator may pull the SCS aggregate, the per-entity capex, and the deal-flow records into an internal analytical workflow that pairs the platform’s data with the fund’s own portfolio-level analytics. A hyperscaler business-development team may pull the capacity tracker and the per-region operational-status records into an internal go-to-market planning tool. A semiconductor-strategy desk may pull the silicon-vendor allocation tables into a competitive-intelligence dashboard.

The platform’s role is to be the highest-quality public-information data layer that those proprietary pipelines consume. The role does not require the platform to know what the proprietary pipelines do; the access pattern is structured query against the documented schema, with the proprietary work happening on the subscriber’s side.

Visualization and Embed Use

A subset of the tables layer is available for visualization and embed use. The high-cardinality tables — the SCS aggregate scores, the per-campus capacity disclosures, the deal-flow records, the per-year aggregations — render natively as charts, sortable tables, and timeline views inside the platform’s analytical pages. The visualizations are integrated into the in-page analytical layer rather than living as standalone dashboards.

For institutional and enterprise subscribers, the embed layer supports inclusion of platform-rendered visualizations into subscriber-side internal documents and publications, with the appropriate citation-and-attribution discipline. The embed pattern preserves the platform’s source-vetting and methodology-disclosure framing inside the subscriber-side context.

Quality Audits and Spot Checks

The tables layer is subject to continuous quality audits. The platform’s editorial team runs spot checks on randomly-selected extracted tables, verifying the extracted values against the source documents and flagging extraction errors for correction. Systematic error patterns identified through the spot checks drive improvements to the extraction pipeline; idiosyncratic errors are corrected at the per-record level.

The audit cadence is calibrated to the corpus size and the extraction-error rate. The platform’s commitment is that the tables layer maintains a low extraction-error rate at the per-cell level, with the residual error rate documented and disclosed to subscribers who run audit-sensitive analytical workflows. Subscribers can also flag suspected extraction errors through the platform’s contact layer; flagged errors are evaluated against the source document and corrected where the flag is validated.

For deeper reading:

  • Sources — the broader source corpus
  • Documents — the PDF subset from which most tables are extracted
  • Methodology — how tables flow into the SCS framework
  • Terminal — the subscriber product where tables are programmatically accessible