July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 25

report.pdf

Page 25 · 670 words

On the Opportunities and Risks of Foundation Models
25
2020], and whether their apparent multilingual performance relies more on assimilation [Lauscher
et al. 2020; Virtanen et al. 2019; Artetxe et al. 2020]. Multilingual models show better performance in
languages that are similar to the highest-resource languages in their training data, and it has been
shown that languages in multilingual models compete for model parameters, making it unclear
how much variation can fit in a single model [Wang et al. 2020d]. A salient issue stems from the
data that we use to train multilingual foundation models: in many multilingual corpora, English
data is not only orders of magnitude more abundant than that of lower-resource languages, but it
is often cleaner, broader, and contains examples showcasing more linguistic depth and complexity
[Caswell et al. 2021] (see Nekoto et al. [2020] on building participatory and robust multilingual
datasets). However, the answer does not simply lie in creating more balanced corpora: there are so
many axes of language variation that it would be infeasible to create a corpus that is balanced and
representative in all regards. The future, versatility, and equity of foundation models all depend on
robustly handling language variation despite unbalanced data [e.g., Oren et al. 2019].
Current multilingual foundation models in their raw form, and naive unsupervised multilingual
training as a method, may not model the subtleties of languages and language varieties to their full
extent. Nevertheless, they remain useful for some multilingual applications, for example through
adapting multilingual models for low-resource languages not in their original training set [Wang
et al. 2020b]. Moreover, the results for the (non-public) GShard neural machine translation model
show the largest gains over monolingual baselines for the lowest resource languages, with the
gains increasing with model size [Lepikhin et al. 2021]. The research community should critically
examine how foundation models deal with language variation, understand the limits of foundation
models in bringing equity and representation to NLP, and not settle on promoting foundation
models that erase language variation and mostly conform to the linguistic majority in their training
data.
2.1.4
Inspiration from human language acquisition.
Though foundation models have constituted a huge source of progress in creating NLP systems
that act more like humans, there are still significant ways in which the linguistic system that they
acquire, as well as the learning process, differ from human language. Understanding the implications
of this gap between machine and human language learning is a necessary part of developing a
research community informed about the linguistic limits and possibilities of foundation models.
Human language acquisition is very efficient: foundation models like GPT-3 are trained on around
three to four orders of magnitude more language data than most humans will ever hear or read, and
certainly much more than children have been exposed to by the time they are mostly linguistically
competent. One salient difference between foundation models and human language acquisition is
that human language is grounded to the real world [Saxton 2017]. For example babies and caretakers
point to objects during language development [Colonnesi et al. 2010], and babies learn the grounded
meanings of words that refer to common objects before they learn a lot of the other a
→ report.pdf page 25