Report · Page 48
report.pdf
Page Content
48 Center for Research on Foundation Models (CRFM) 2.6 Philosophy of understanding Authors: Christopher Potts, Thomas Icard, Eva Portelance, Dallas Card, Kaitlyn Zhou, John Etchemendy What could a foundation model come to understand about the data it is trained on? An answer to this question would be extremely informative about the overall capacity of foundation models to contribute to intelligent systems. In this section, we focus on the case of natural language, since language use is a hallmark of human intelligence and central to the human experience. The best foundation models at present can consume and produce language with striking fluency, but they invariably lapse into the sort of incoherence that suggests they are merely “stochastic parrots” [Bender et al. 2021]. Are these lapses evidence of inherent limitations, or might future foundation models truly come to understand the symbols they process? Our aim in this section is to clarify these questions, and to help structure debates around them. We begin by explaining what we mean by foundation model, paying special attention to how foundation models are trained, since the training regime delimits what information the model gets about the world. We then address why it is important to clarify these questions for the further development of such models. Finally, we seek to clarify what we mean by understanding, addressing both what understanding is (metaphysics) and how we might come to reliably determine whether a model has achieved understanding (epistemology). Ultimately, we conclude that skepticism about the capacity of future models to understand natural language may be premature. It is by no means obvious that foundation models alone could ever achieve understanding, but neither do we know of definitive reasons to think they could not. 2.6.1 What is a foundation model? There is not a precise technical definition of foundation model. Rather, this is an informal label for a large family of models, and this family of models is likely to grow and change over time in response to new research. This poses challenges to reasoning about their fundamental properties. However, there is arguably one defining characteristic shared by all foundation models: they are self-supervised. Our focus is on the case where self-supervision is the model’s only formal objective. In self-supervision, the model’s sole objective is to learn abstract co-occurrence patterns in the sequences of symbols it was trained on. This task enables many of these models to generate plausible strings of symbols as well. For example, many foundation models are structured so that one can prompt them with a sequence like “The sandwich contains peanut” and ask them to generate a continuation – say, “butter and jelly”. Other models are structured so that they are better at filling in gaps; you might prompt a model with “The sandwich contains __ and jelly” and expect it to fill in “peanut butter”. Both capabilities derive from these models’ ability to extract co-occurrence patterns from their training data. There is no obvious sense in which this kind of self-supervision tells the model anything about what the symbols mean. The only information it is given directly is information about which words tend to co-occur with which other words. On the face of it, knowing that “The sandwich contains peanut” is likely to be continued with “butter and jelly” says nothing about what sandwiches are, what jelly is, how these objects will b
Source Document
→ report.pdf
page 48