Report · Page 114
report.pdf
Page Content
114 Center for Research on Foundation Models (CRFM) 4.9 AI safety and alignment Authors: Alex Tamkin, Geoff Keeling, Jack Ryan, Sydney von Arx The field of Artificial Intelligence (AI) Safety concerns itself with potential accidents, hazards, and risks of advanced AI models, especially larger-scale risks to communities or societies. Current foundation models may be far from posing such risks; however, the breadth of their capabilities and potential applications is striking, and a clear shift from previous ML paradigms. While AI safety has historically occupied a more marginal position within AI research, the current transition towards foundation models and their corresponding generality offers an opportunity for AI safety researchers to revisit the core questions of the field in a new light and reassess their immediate or near-future relevance.80 4.9.1 Traditional problems in AI safety. A major branch of AI safety research concerns the implications of advanced AI systems, including those that might match or exceed human performance across a broad class of cognitive tasks [Everitt et al. 2018].81 A central goal of safety research in this context is to mitigate large-scale risks posed by the development of advanced AI.82 These risks may be significantly more speculative than those considered in §5.2: misuse, §4.8: robustness, and §4.7: security; however, they are of far greater magnitude, and could at least in principle result from future, highly-capable systems. Of particular concern are global catastrophic risks: roughly, risks that are global or trans-generational in scope—causing death or otherwise significantly reducing the welfare of those affected (e.g., a nuclear war or rapid ecological collapse) [Bostrom and Cirkovic 2011]. What AI safety research amounts to, then, is a family of projects which aim to characterize what (if any) catastrophic risks are posed by the development of advanced AI, and develop plausible technical solutions for mitigating the probability or the severity of these risks. The best-case scenario from the point of view of AI safety is a solution to the control problem: how to develop an advanced AI system that enables us to reap the computational benefits of that system while at the same time leaving us with sufficient control such that the deployment of the system does not result in a global catastrophe [Bostrom and Cirkovic 2011]. However technical solutions are not sufficient to ensure safety: ensuring that safe algorithms are actually those implemented into real-world systems and that unsafe systems are not deployed may require additional sociotechnical measures and institutions. Reinforcement Learning (RL), which studies decision-making agents optimized towards rewards, has been a dominant focus in AI safety for the past decade. What is at issue here is the difficulty of specifying and instantiating a reward function for the AI that aligns with human values, in the minimal sense of not posing a global catastrophic threat.83 While this problem, known as value alignment [Gabriel 2020; Yudkowsky 2016], may seem trivial at first glance, human values are diverse,84 amorphous, and challenging to capture quantitatively. Due to this, a salient concern is reward hacking, where the AI finds an unforeseen policy that maximizes a proxy reward for human wellbeing, but whose misspecification results in a significant harm.85 Many efforts to combat the 80See Amodei et al. [2016] and Hendrycks et al. [2021d] for broader p
Source Document
→ report.pdf
page 114