July 10, 2026
Year of AI 2026 · Updated July 2026
SAUDI COMPUTE
The Kingdom's Compute Buildout, Tracked.
Sovereign AI Infrastructure · Capital Flows · Geopolitical Intelligence

Report · Page 114

report.pdf

Page 114 · 662 words

114
Center for Research on Foundation Models (CRFM)
4.9
AI safety and alignment
Authors: Alex Tamkin, Geoff Keeling, Jack Ryan, Sydney von Arx
The field of Artificial Intelligence (AI) Safety concerns itself with potential accidents, hazards,
and risks of advanced AI models, especially larger-scale risks to communities or societies. Current
foundation models may be far from posing such risks; however, the breadth of their capabilities
and potential applications is striking, and a clear shift from previous ML paradigms. While AI
safety has historically occupied a more marginal position within AI research, the current transition
towards foundation models and their corresponding generality offers an opportunity for AI safety
researchers to revisit the core questions of the field in a new light and reassess their immediate or
near-future relevance.80
4.9.1
Traditional problems in AI safety.
A major branch of AI safety research concerns the implications of advanced AI systems, including
those that might match or exceed human performance across a broad class of cognitive tasks
[Everitt et al. 2018].81 A central goal of safety research in this context is to mitigate large-scale risks
posed by the development of advanced AI.82 These risks may be significantly more speculative
than those considered in §5.2: misuse, §4.8: robustness, and §4.7: security; however, they are of
far greater magnitude, and could at least in principle result from future, highly-capable systems. Of
particular concern are global catastrophic risks: roughly, risks that are global or trans-generational
in scope—causing death or otherwise significantly reducing the welfare of those affected (e.g., a
nuclear war or rapid ecological collapse) [Bostrom and Cirkovic 2011]. What AI safety research
amounts to, then, is a family of projects which aim to characterize what (if any) catastrophic risks are
posed by the development of advanced AI, and develop plausible technical solutions for mitigating
the probability or the severity of these risks. The best-case scenario from the point of view of AI
safety is a solution to the control problem: how to develop an advanced AI system that enables us
to reap the computational benefits of that system while at the same time leaving us with sufficient
control such that the deployment of the system does not result in a global catastrophe [Bostrom
and Cirkovic 2011]. However technical solutions are not sufficient to ensure safety: ensuring that
safe algorithms are actually those implemented into real-world systems and that unsafe systems
are not deployed may require additional sociotechnical measures and institutions.
Reinforcement Learning (RL), which studies decision-making agents optimized towards rewards,
has been a dominant focus in AI safety for the past decade. What is at issue here is the difficulty of
specifying and instantiating a reward function for the AI that aligns with human values, in the
minimal sense of not posing a global catastrophic threat.83 While this problem, known as value
alignment [Gabriel 2020; Yudkowsky 2016], may seem trivial at first glance, human values are
diverse,84 amorphous, and challenging to capture quantitatively. Due to this, a salient concern is
reward hacking, where the AI finds an unforeseen policy that maximizes a proxy reward for human
wellbeing, but whose misspecification results in a significant harm.85 Many efforts to combat the
80See Amodei et al. [2016] and Hendrycks et al. [2021d] for broader p
→ report.pdf page 114