Report · Page 30
report.pdf
Page Content
30 Center for Research on Foundation Models (CRFM) of still or moving objects, and include tasks of depth estimation, structure-from-motion, surface normal detection, curvature line and keypoint estimation, to name a few [e.g., Laina et al. 2016; Agarwal et al. 2011; Wang et al. 2015a; Zamir et al. 2018; Ullman 1979]. (3) multimodal integration tasks, combining semantic and geometric understanding with other modalities such as natural language; these include, for instance, visual question answering, image captioning, and instruction following [e.g., Antol et al. 2015; Chen et al. 2015b; Anderson et al. 2018; Goyal et al. 2017b; Hudson and Manning 2019b; Johnson et al. 2017; Luo et al. 2020; Akbari et al. 2021; Huang et al. 2021c; Tsimpoukelli et al. 2021]. We highlight a subset of traditional core tasks in Figure 7. The predominant paradigm for addressing these tasks, driven by the emergence of ImageNet [Deng et al. 2009] during the early 2010s, tends to center around a familiar core idea: First, pretrain a model on a large collection of carefully annotated data [Russakovsky et al. 2015] with a fully supervised training task, like image classification. Then, adapt the model downstream on task- specific datasets and domains [Lin et al. 2014; Chen et al. 2015b; Antol et al. 2015] by fine-tuning to reach state-of-the-art performance [Krizhevsky et al. 2012; Simonyan and Zisserman 2015; He et al. 2016a; Xu and Saenko 2016]. This notion of pretraining followed by adaptation persists in the definitions we consider now for foundation models (§1: introduction). The limitations of this fully supervised paradigm motivate the transition to foundation models: the reliance on external supervised annotations constrains the upper bound capability of previous approaches to capture the diverse spectrum of visual inputs in a scalable, robust and generalizable manner. Recent developments in the domain of visual synthesis and unsupervised learning offer a compelling alternative. GANs, for instance, learn to generate visual content of high fidelity, realism and diversity, by featuring two competing networks of a generator and a discriminator that can supervise one another from image collections alone [e.g., Goodfellow et al. 2014; Hudson and Zitnick 2021]. Other neural models infer the visual properties of objects and scenes without explicitly annotated supervision, by employing variational auto-encoding, contrastive learning or other self-supervised techniques [e.g., Kingma and Welling 2014; Chen et al. 2020c; He et al. 2020]. For instance, He et al. [2021] build upon prior work on representation learning with masked image encoding [e.g., Pathak et al. 2016; Vincent et al. 2008] by, in part, combining recent advancements in flexible architectures (e.g., vision transformers [Dosovitskiy et al. 2021; Zhai et al. 2021]) with increased scaling. With foundation models, the development of such self-supervision techniques has enabled train- ing at greater scales of visual data [Changpinyo et al. 2021], both in terms of its scope as well as its potential diversity. Accordingly, we have seen early indicators of progress on traditional vision tasks in terms of both standard accuracy metrics and few-shot generalization. For image classification and object detection, self-supervised techniques have reported competitive perfor- mance to prior fully-supervised approaches [He et al. 2019; Chen et al. 2020c; Radford et al. 2021; Hénaff et al. 2021], without explicit annot
Source Document
→ report.pdf
page 30