Report · Page 126
report.pdf
Page Content
126 Center for Research on Foundation Models (CRFM) these approaches are separate from the model whose behavior is analyzed, which by itself is not interpretable. This separation can be problematic, as the provided explanations can lack faithfulness [Jacovi and Goldberg 2020], by being unreliable and misleading about the causes of a behavior [cf. Rudin 2019]. Even further, unsound explanations can entice humans into trusting unsound models more than they otherwise would (for a detailed discussion of trust in artificial intelligence, see Jacovi et al. [2021]). These types of concerns grow as we transition from task-specific models towards the wide adoption of foundation models, as their behavior is vastly more complex. Current explanatory approaches can largely be divided into either providing local or global explanations of model behavior [Doshi-Velez and Kim 2017]. Local explanations seek to explain a model’s response to a specific input, e.g., by attributing a relevance to each input feature for the behavior or by identifying the training samples most relevant for the behavior [Simonyan et al. 2013; Bach et al. 2015; Sundararajan et al. 2017; Shrikumar et al. 2017; Springenberg et al. 2014; Zeiler and Fergus 2014; Lundberg and Lee 2017; Zintgraf et al. 2017; Fong and Vedaldi 2017; Koh and Liang 2017]. Global explanations, in contrast, are not tied to a specific input and instead aim to uncover qualities of the data at large that affect model behaviors, e.g., by synthesizing the input that the model associates most strongly with a behavior [Simonyan et al. 2013; Nguyen et al. 2016]. Local and global explanations have provided useful insights into the behavior of task-specific models [e.g., Li et al. 2015; Wang et al. 2015b; Lapuschkin et al. 2019; Thomas et al. 2019; Poplin et al. 2018]. Here, the resulting explanations are often taken to be a heuristic of the model mechanisms that gave rise to a behavior; for example, seeing that an explanation attributes high importance to horizontal lines when the model reads a handwritten digit ‘7’ easily creates the impression that horizontal lines are a generally important feature that the model uses to identify all sevens or perhaps to distinguish all digits. Given the one model–many models nature of foundation models, however, we should be careful not to jump from specific explanations of a behavior to general assumptions about the model’s behavior. While current explanatory approaches may shed light on specific behaviors, for example, by identifying aspects of the data that strongly effected these behaviors, the resulting explanations do not necessarily provide insights into the model’s behaviors for other (even seemingly similar) inputs, let alone other tasks and domains. Another approach could be to sidestep these types of post-hoc explanations altogether by leveraging the generative abilities of foundation models in the form of self-explanations [cf. Elton 2020; Chen et al. 2018], that is, by training these models to generate not only the response to an input, but to jointly generate a human-understandable explanation of that response. While it is unclear whether this approach will be fruitful in the future, there are reasons to be skeptical: language models, and now foundation models, are exceptional at producing fluent, seemingly plausible content without any grounding in truth. Simple self-generated “explanations” could follow suit. It is thus important to be discerning of the difference
Source Document
→ report.pdf
page 126