July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 134

report.pdf

Page 134 · 670 words

134
Center for Research on Foundation Models (CRFM)
of harm requires upstream propagation of both feedback and accountability to foundation model
providers.
Intervention. General principles that govern intervention on technological systems apply to
the foundation model setting: identifying which sources are most responsible for bias or harm
provides the evidence required for targeted action. For example, the urgency of calls for improved
diversity in the teams that design, produce, and control technology (e.g., foundation models) and
their applications [Longino 1990; Harding 2015; Nielsen et al. 2017; O’Connor et al. 2019; Hofstra
et al. 2020; Katell et al. 2020] is further intensified if the lack of diversity is shown to relate to harm
[Caswell et al. 2021]. In addition, transparent documentation [e.g., Gebru et al. 2018; Bender and
Friedman 2018; Mitchell et al. 2019] and auditing [e.g., Raji and Buolamwini 2019] are similarly
critical in providing the impetus for intervention and change [Burrell 2016; Lipton 2018; Creel 2020;
Raji et al. 2020; Wilson et al. 2021]. The scale of foundation models, as well as the specifics of their
accessibility, introduce new challenges for existing protocols for documentation and auditing that
we discuss further in §5.6: ethics.
To date, many of the interventions considered for reducing the inequitable impact of technology,
including in the foundation model regime, are methods for technical mitigation that center the
data (to obviate reflecting inequities or biases) and modelling decisions (to avoid amplifying data
biases) involved. Of specific importance in the foundation model regime is recognizing that these
mitigation approaches may target different steps in the pipeline such as the training data [e.g., Lu
et al. 2020], modelling objectives [e.g., Zhao et al. 2018]), and adaptation methods and test-time
use [e.g., Park et al. 2018; Zhao et al. 2019]. As a result, different approaches may not only be
more or less effective, but require action from different entities (e.g., foundation model providers vs.
application developers) and more or less intensively affect the expensive training process for these
models (e.g., changing the process of creating a foundation model vs. altering it post hoc). Technical
intervention of this form may also target different goals: some interventions, such as changing
the training data, aims to reduce intrinsic bias. On the other hand, most work on mitigation in
algorithmic/ML fairness instead considers reducing outcome disparities in terms of model behavior,
i.e., the outputs of downstream systems that more directly relate to extrinsic harm. Technical
mitigation of all forms at present is severely limited: methods that measure or combat intrinsic
bias are brittle or ineffectual [Gonen and Goldberg 2019; Ethayarajh et al. 2019; Bommasani et al.
2020; Zhou et al. 2021b; Antoniak and Mimno 2021], methods that measure or combat extrinsic
outcome disparities may not align with stakeholder goals [Saha et al. 2020], and there is some
evidence to suggest certain types of technical intervention may be simultaneously unsatisfiable
[Corbett-Davies and Goel 2018; Kleinberg et al. 2017], impossible [Lechner et al. 2021], or may even
exacerbate inequity [Xu et al. 2021]. In spite of this state of affairs, we continue to believe technical
methods will still play an instrumental role in addressing the harms that arise in the foundation
model regime; in general, we advocate for transp
→ report.pdf page 134