July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 51

report.pdf

Page 51 · 699 words

On the Opportunities and Risks of Foundation Models
51
referentialism, there is still a further question of how these proxies relate to the actual world, but
the same question arises for human language users as well.
Bender and Koller [2020] give an interesting argument that combines referentialism with prag-
matism. They imagine an agent O that intercepts communications between two humans speaking
a natural language L. O inhabits a very different world from the humans and so does not have
the sort of experiences needed to ground the humans’ utterances in the ways that referentialism
demands. Nonetheless, O learns from the patterns in the humans’ utterances, to the point where O
can even successfully pretend to be one of the humans. Bender and Koller then seek to motivate the
intuition that we can easily imagine situations in which O’s inability to ground L in the humans’
world will reveal itself, and that this will in turn reveal that O does not understand L. The guiding
assumption seems to be that the complexity of the world is so great that no amount of textual
exchange can fully cover it, and the gaps will eventually reveal themselves. In the terms we have
defined, the inability to refer is taken to entail that the agent is not in the right dispositional state
for understanding.
Fundamentally, the scenario Bender and Koller describe is one in which some crucial information
for understanding is taken to be missing, and a simple behavioral test reveals this. We can agree
with this assessment without concluding that foundation models are in general incapable of
understanding. This again brings us back to the details of the training data involved. If we modify
Bender and Koller’s scenario so that the transmissions include digitally encoded images, audio, and
sensor readings from the humans’ world, and O is capable of learning associations between these
digital traces and linguistic units, then we might be more optimistic – there might be a practical
issue concerning O’s ability to get enough data to generalize, but perhaps not an in principle
limitation on what O can achieve.29
We tentatively conclude that there is no easy a priori reason to think that varieties of under-
standing falling under any of our three positions could not be learned in the relevant way. With
this possibility thus still open, we face the difficult epistemological challenge of clarifying how we
could hope to evaluate potential success.
Epistemology of understanding. A positive feature of pragmatism is that, by identifying success
with the manifestation of concrete behaviors, there is no great conceptual puzzle about how to test
for it. We simply have to convince ourselves that our limited observations of the system’s behavior
so far indicate a reliable disposition toward the more general class of behaviors that we took as our
target. Of course, agreeing on appropriate targets is very difficult. When concrete proposals are
made, they are invariably met with objections, often after putative success is demonstrated.
The history of the Turing Test is instructive here: although numerous artificial agents have passed
actual Turing Tests, none of them has been widely accepted as intelligent as a result. Similarly, in
recent years, a number of benchmark tasks within NLP have been proposed to evaluate specific
aspects of understanding (e.g., answering simple questions, performing commonsense reasoning).
When systems surpass our estimates of human performance, the communit
→ report.pdf page 51