Report · Page 29
report.pdf
Page Content
On the Opportunities and Risks of Foundation Models 29 The field of computer vision and the challenges we define draw inspiration in many ways from human perception capabilities. Several classical theories [e.g., Biederman 1972; McClelland and Rumelhart 1981; Marr 1982] suggested that humans may perceive real world scenes by contextual- izing parts as a larger whole, and pointed the way for computer vision techniques to progressively model the physical world with growing levels of abstractions [Lowe 1992; Girshick et al. 2014]. Gibson [1979] suggested that human vision is inherently embodied and interactive ecological environments may play a key role in its development. These ideas continue to motivate the ongoing development of computer vision systems, iterating towards a contextual, interactive, and embodied perception of the world. In the context of computer vision, foundation models translate raw perceptual information from diverse sources and sensors into visual knowledge that may be adapted to a multitude of downstream settings (Figure 7). To a large extent, this effort is a natural evolution of the key ideas that have emerged from the field over the last decade. The introduction of ImageNet [Deng et al. 2009] and the advent of supervised pretraining led to a deep learning paradigm shift in computer vision. This transition marked a new era, where we moved beyond the classic approaches and task-specific feature engineering of earlier days [Lowe 2004; Bay et al. 2006; Rosten and Drummond 2006] towards models that could be trained once over large amounts of data, and then adapted for a broad variety of tasks, such as image recognition, object detection, and image segmentation [Krizhevsky et al. 2012; Szegedy et al. 2015; He et al. 2016a; Simonyan and Zisserman 2015]. This idea remains at the core of foundation models. The bridge to foundation models comes from the limitations of the previous paradigm. Traditional supervised techniques rely on expensive and carefully-collected labels and annotations, limiting their robustness, generalization and applicability; in contrast, recent advances in self-supervised learning [Chen et al. 2020c; He et al. 2020] suggest an alternative route for the development of foundation models that could make use of large quantities of raw data to attain a contextual understanding of the visual world. Relative to the broader aims of the field, the current capabilities of vision foundation models are currently early-stage (§2.2.1: vision-capabilities): we have observed improvements in traditional computer vision tasks (particularly with respect to generalization capability) [Radford et al. 2021; Ramesh et al. 2021] and anticipate that the near-term progress will continue this trend. However, in the longer-term, the potential for foundation models to reduce dependence on explicit annotations may lead to progress on essential cognitive skills (e.g., commonsense reasoning) which have proven difficult in the current, fully-supervised paradigm [Zellers et al. 2019a; Martin-Martin et al. 2021]. In turn, we discuss the potential implications of foundation models for downstream applications, and the central challenges and frontiers that must be addressed moving forward (§2.2.2: vision-challenges). 2.2.1 Key capabilities and approaches. At a high-level, computer vision is the core sub-field of artificial intelligence that explores ways to endow machines with the capacity to interpret and understand the visual world. I
Source Document
→ report.pdf
page 29