July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 30

report.pdf

Page 30 · 705 words

30
Center for Research on Foundation Models (CRFM)
of still or moving objects, and include tasks of depth estimation, structure-from-motion, surface
normal detection, curvature line and keypoint estimation, to name a few [e.g., Laina et al. 2016;
Agarwal et al. 2011; Wang et al. 2015a; Zamir et al. 2018; Ullman 1979]. (3) multimodal integration
tasks, combining semantic and geometric understanding with other modalities such as natural
language; these include, for instance, visual question answering, image captioning, and instruction
following [e.g., Antol et al. 2015; Chen et al. 2015b; Anderson et al. 2018; Goyal et al. 2017b; Hudson
and Manning 2019b; Johnson et al. 2017; Luo et al. 2020; Akbari et al. 2021; Huang et al. 2021c;
Tsimpoukelli et al. 2021]. We highlight a subset of traditional core tasks in Figure 7.
The predominant paradigm for addressing these tasks, driven by the emergence of ImageNet
[Deng et al. 2009] during the early 2010s, tends to center around a familiar core idea: First, pretrain
a model on a large collection of carefully annotated data [Russakovsky et al. 2015] with a fully
supervised training task, like image classification. Then, adapt the model downstream on task-
specific datasets and domains [Lin et al. 2014; Chen et al. 2015b; Antol et al. 2015] by fine-tuning
to reach state-of-the-art performance [Krizhevsky et al. 2012; Simonyan and Zisserman 2015; He
et al. 2016a; Xu and Saenko 2016]. This notion of pretraining followed by adaptation persists
in the definitions we consider now for foundation models (§1: introduction). The limitations
of this fully supervised paradigm motivate the transition to foundation models: the reliance on
external supervised annotations constrains the upper bound capability of previous approaches to
capture the diverse spectrum of visual inputs in a scalable, robust and generalizable manner. Recent
developments in the domain of visual synthesis and unsupervised learning offer a compelling
alternative. GANs, for instance, learn to generate visual content of high fidelity, realism and diversity,
by featuring two competing networks of a generator and a discriminator that can supervise one
another from image collections alone [e.g., Goodfellow et al. 2014; Hudson and Zitnick 2021].
Other neural models infer the visual properties of objects and scenes without explicitly annotated
supervision, by employing variational auto-encoding, contrastive learning or other self-supervised
techniques [e.g., Kingma and Welling 2014; Chen et al. 2020c; He et al. 2020]. For instance, He et al.
[2021] build upon prior work on representation learning with masked image encoding [e.g., Pathak
et al. 2016; Vincent et al. 2008] by, in part, combining recent advancements in flexible architectures
(e.g., vision transformers [Dosovitskiy et al. 2021; Zhai et al. 2021]) with increased scaling.
With foundation models, the development of such self-supervision techniques has enabled train-
ing at greater scales of visual data [Changpinyo et al. 2021], both in terms of its scope as well
as its potential diversity. Accordingly, we have seen early indicators of progress on traditional
vision tasks in terms of both standard accuracy metrics and few-shot generalization. For image
classification and object detection, self-supervised techniques have reported competitive perfor-
mance to prior fully-supervised approaches [He et al. 2019; Chen et al. 2020c; Radford et al. 2021;
Hénaff et al. 2021], without explicit annot
→ report.pdf page 30