July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 29

report.pdf

Page 29 · 662 words

On the Opportunities and Risks of Foundation Models
29
The field of computer vision and the challenges we define draw inspiration in many ways from
human perception capabilities. Several classical theories [e.g., Biederman 1972; McClelland and
Rumelhart 1981; Marr 1982] suggested that humans may perceive real world scenes by contextual-
izing parts as a larger whole, and pointed the way for computer vision techniques to progressively
model the physical world with growing levels of abstractions [Lowe 1992; Girshick et al. 2014].
Gibson [1979] suggested that human vision is inherently embodied and interactive ecological
environments may play a key role in its development. These ideas continue to motivate the ongoing
development of computer vision systems, iterating towards a contextual, interactive, and embodied
perception of the world.
In the context of computer vision, foundation models translate raw perceptual information
from diverse sources and sensors into visual knowledge that may be adapted to a multitude of
downstream settings (Figure 7). To a large extent, this effort is a natural evolution of the key ideas
that have emerged from the field over the last decade. The introduction of ImageNet [Deng et al.
2009] and the advent of supervised pretraining led to a deep learning paradigm shift in computer
vision. This transition marked a new era, where we moved beyond the classic approaches and
task-specific feature engineering of earlier days [Lowe 2004; Bay et al. 2006; Rosten and Drummond
2006] towards models that could be trained once over large amounts of data, and then adapted
for a broad variety of tasks, such as image recognition, object detection, and image segmentation
[Krizhevsky et al. 2012; Szegedy et al. 2015; He et al. 2016a; Simonyan and Zisserman 2015]. This
idea remains at the core of foundation models.
The bridge to foundation models comes from the limitations of the previous paradigm. Traditional
supervised techniques rely on expensive and carefully-collected labels and annotations, limiting
their robustness, generalization and applicability; in contrast, recent advances in self-supervised
learning [Chen et al. 2020c; He et al. 2020] suggest an alternative route for the development
of foundation models that could make use of large quantities of raw data to attain a contextual
understanding of the visual world. Relative to the broader aims of the field, the current capabilities of
vision foundation models are currently early-stage (§2.2.1: vision-capabilities): we have observed
improvements in traditional computer vision tasks (particularly with respect to generalization
capability) [Radford et al. 2021; Ramesh et al. 2021] and anticipate that the near-term progress
will continue this trend. However, in the longer-term, the potential for foundation models to
reduce dependence on explicit annotations may lead to progress on essential cognitive skills
(e.g., commonsense reasoning) which have proven difficult in the current, fully-supervised paradigm
[Zellers et al. 2019a; Martin-Martin et al. 2021]. In turn, we discuss the potential implications of
foundation models for downstream applications, and the central challenges and frontiers that must
be addressed moving forward (§2.2.2: vision-challenges).
2.2.1
Key capabilities and approaches.
At a high-level, computer vision is the core sub-field of artificial intelligence that explores ways to
endow machines with the capacity to interpret and understand the visual world. I
→ report.pdf page 29