Report · Page 131
report.pdf
Page Content
On the Opportunities and Risks of Foundation Models 131 Intrinsic biases. Properties of the foundation model can lead to harm in downstream systems. As a result, these intrinsic biases can be measured directly within the foundation model, though the harm itself is only realized when the foundation model is adapted, and thereafter applied, i.e., these are latent biases or harms [DeCamp and Lindvall 2020]. We focus on the most widely studied form of intrinsic bias, representational bias, specifically considering misrepresentation, underrepresentation and overrepresentation. People can be misrepresented by pernicious stereo- types [Bolukbasi et al. 2016; Caliskan et al. 2017; Abid et al. 2021; Nadeem et al. 2021; Gehman et al. 2020] or negative attitudes [Hutchinson et al. 2020], which can propagate through downstream models to reinforce this misrepresentation in society [Noble 2018; Benjamin 2019]. People can be underrepresented or entirely erased, e.g., when LGBTQ+ identity terms [Strengers et al. 2020; Oliva et al. 2021; Tomasev et al. 2021] or data describing African Americans [Buolamwini and Gebru 2018; Koenecke et al. 2020; Blodgett and O’Connor 2017] is excluded in training data, downstream models will struggle with similar data at test-time. People can be overrepresented, e.g., BERT appears to encode an Anglocentric perspective [Zhou et al. 2021a] by default, which can amplify majority voices and contribute to homogenization of perspectives [Creel and Hellman 2021] or monoculture [Kleinberg and Raghavan 2021] (§5.6: ethics). These representational biases pertain to all AI systems, but their significance is greatly heightened in the foundation model paradigm. Since the same foundation model serves as the basis for myriad applications, biases in the representation of people propagate to many applications and settings. Further, since the foundation model does much of the heavy-lifting (compared to adaptation, which is generally intended to be lightweight), we anticipate that many of the experienced harms will be significantly determined by the internal properties of the foundation model. Extrinsic harms. Users can experience specific harms from the downstream applications that are created by adapting a foundation model. These harms can be representational [Barocas et al. 2017; Crawford 2017; Blodgett et al. 2020], such as the sexualized depictions of black women produced by information retrieval systems [Noble 2018], the misgendering of persons by machine translation systems that default to male pronouns [Schiebinger 2013, 2014], or the generation of pernicious stereotypes [Nozza et al. 2021; Sheng et al. 2019; Abid et al. 2021]. They can consist of abuse, such as when dialogue agents based on foundation models attack users with toxic content [Dinan et al. 2021; Gehman et al. 2020] or microaggressions [Breitfeller et al. 2019; Jurgens et al. 2019]. All of these user-facing behaviors can lead to psychological harms or the reinforcement of pernicious stereotypes [Spencer et al. 2016; Williams 2020]. In addition to harms experienced by individuals, groups or sub-populations may also be subject to harms such as group-level performance disparities. For example, systems may perform poorly on text or speech in African American English [Blodgett and O’Connor 2017; Koenecke et al. 2020], incorrectly detect medical conditions from clinical notes for racial, gender, and insurance-status minority groups [Zhang et al. 2020b], or fail to detect the
Source Document
→ report.pdf
page 131