July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 32

report.pdf

Page 32 · 675 words

32
Center for Research on Foundation Models (CRFM)
geometric understanding in perception models may provide guidance for ongoing foundation model
development [Yi et al. 2019; Bakhtin et al. 2019; Li et al. 2020b]. Indeed, the continued incorporation
of multiple modalities (e.g., audio) in foundation models may prove beneficial towards these aims
[Zhang et al. 2017; Gao et al. 2020b; Jaegle et al. 2021a]. However, the specific techniques to enable
generalizing the initial observed capabilities robustly to a wide range of natural scenes and objects
at the level of humans remains an open research challenge for foundation models.
Computational efficiency and dynamics modeling. Humans are surprisingly efficient at pro-
cessing the continuous visual stream of objects, scenes, and events necessary to support an un-
derstanding of event dynamics [Zacks et al. 2001; Tversky and Zacks 2013]. Foundation models in
language (§2.1: language) have shown initial steps towards modeling longer-term coherence of
events; the analogous ability to capture long-range temporal correlations and causal coherence in
visual input would stand to benefit downstream settings like robotics [Dai et al. 2019; Alyamkin
et al. 2019; Goel et al. 2020b; Feng et al. 2019, §2.3: robotics]. However, relative to word token-level
inputs in language, low-level computer vision inputs are extremely high-dimensional: a single
1080p frame contains over 2 million pixels. In this context, modeling the richer event dynamics in
long-range video sequences seems like a daunting endeavor, especially with additional modalities
(e.g., speech, optical flow, etc.) and increasing resolutions. Understandably, a naïve approach to
fully processing every individual pixel is likely prohibitive. Current vision models [e.g., Radford
et al. 2021; Sun et al. 2019a; Tan and Bansal 2019; Kim et al. 2021a] often address this by processing
embeddings that summarize image patches or even groups of frames altogether, but this has the
potential drawback of losing fine-grained details [Ramesh et al. 2021]. In addition to considerations
of the raw input space, foundation models for vision may need to revisit the design of fundamental
architecture primitives (§4.1: modeling) for efficient and effective modeling: alternatives to 3D
convolutions may better address its cubic complexity [Fan et al. 2020; Sitzmann et al. 2019], while
particle-based representations may prove more effective for modeling physical dynamics [Bear
et al. 2021]. Further, deployment of these vision models to downstream application settings will
also necessitate advancements in systems design (§4.5: systems). Taken together, the bottleneck of
efficient and effective modeling for larger-scale, dynamic vision inputs remains a multi-faceted
research direction that must be addressed going forward.
Training, environments, and evaluation. Equally critical to realizing the potential of founda-
tion models are the supporting elements for training and evaluating them. Current foundation
models for vision have largely focused on a small subset of modalities shown in Figure 7 (e.g., datasets
of RGB images and text), since these are perhaps the most readily accessible [Changpinyo et al.
2021; Radford et al. 2021]. This motivates the development and use of additional large-scale training
datasets which contain a diverse collection of inputs across a broad spectrum of modalities. While
additional annotations may not strictly be necessary, the input quality i
→ report.pdf page 32