Report · Page 32
report.pdf
Page Content
32 Center for Research on Foundation Models (CRFM) geometric understanding in perception models may provide guidance for ongoing foundation model development [Yi et al. 2019; Bakhtin et al. 2019; Li et al. 2020b]. Indeed, the continued incorporation of multiple modalities (e.g., audio) in foundation models may prove beneficial towards these aims [Zhang et al. 2017; Gao et al. 2020b; Jaegle et al. 2021a]. However, the specific techniques to enable generalizing the initial observed capabilities robustly to a wide range of natural scenes and objects at the level of humans remains an open research challenge for foundation models. Computational efficiency and dynamics modeling. Humans are surprisingly efficient at pro- cessing the continuous visual stream of objects, scenes, and events necessary to support an un- derstanding of event dynamics [Zacks et al. 2001; Tversky and Zacks 2013]. Foundation models in language (§2.1: language) have shown initial steps towards modeling longer-term coherence of events; the analogous ability to capture long-range temporal correlations and causal coherence in visual input would stand to benefit downstream settings like robotics [Dai et al. 2019; Alyamkin et al. 2019; Goel et al. 2020b; Feng et al. 2019, §2.3: robotics]. However, relative to word token-level inputs in language, low-level computer vision inputs are extremely high-dimensional: a single 1080p frame contains over 2 million pixels. In this context, modeling the richer event dynamics in long-range video sequences seems like a daunting endeavor, especially with additional modalities (e.g., speech, optical flow, etc.) and increasing resolutions. Understandably, a naïve approach to fully processing every individual pixel is likely prohibitive. Current vision models [e.g., Radford et al. 2021; Sun et al. 2019a; Tan and Bansal 2019; Kim et al. 2021a] often address this by processing embeddings that summarize image patches or even groups of frames altogether, but this has the potential drawback of losing fine-grained details [Ramesh et al. 2021]. In addition to considerations of the raw input space, foundation models for vision may need to revisit the design of fundamental architecture primitives (§4.1: modeling) for efficient and effective modeling: alternatives to 3D convolutions may better address its cubic complexity [Fan et al. 2020; Sitzmann et al. 2019], while particle-based representations may prove more effective for modeling physical dynamics [Bear et al. 2021]. Further, deployment of these vision models to downstream application settings will also necessitate advancements in systems design (§4.5: systems). Taken together, the bottleneck of efficient and effective modeling for larger-scale, dynamic vision inputs remains a multi-faceted research direction that must be addressed going forward. Training, environments, and evaluation. Equally critical to realizing the potential of founda- tion models are the supporting elements for training and evaluating them. Current foundation models for vision have largely focused on a small subset of modalities shown in Figure 7 (e.g., datasets of RGB images and text), since these are perhaps the most readily accessible [Changpinyo et al. 2021; Radford et al. 2021]. This motivates the development and use of additional large-scale training datasets which contain a diverse collection of inputs across a broad spectrum of modalities. While additional annotations may not strictly be necessary, the input quality i
Source Document
→ report.pdf
page 32