Report · Page 198
report.pdf
Page Content
198 Center for Research on Foundation Models (CRFM) PyTorch. 2021. PyTorch JIT. https://pytorch.org/docs/stable/jit.html. Guanghui Qin and Jason Eisner. 2021. Learning How To Ask: Querying LMs with Mixtures of Soft Prompts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT). Online, 5203–5212. http://cs.jhu.edu/~jason/papers/#qin-eisner-2021 Marc Queudot, Éric Charton, and Marie-Jean Meurs. 2020. Improving Access to Justice with Legal Chatbots. Stats 3, 3 (2020), 356–375. Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence. 2009. When Training and Test Sets Are Different: Characterizing Learning Transfer. In Dataset Shift in Machine Learning. 3–28. Markus N. Rabe, Dennis Lee, Kshitij Bansal, and Christian Szegedy. 2021. Mathematical reasoning via self-supervised skip-tree training. ICLR (2021). https://openreview.net/forum?id=YmqAnY0CMEy Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020 (2021). Alec Radford and Karthik Narasimhan. 2018. Improving Language Understanding by Generative Pre-Training. Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. Technical Report. OpenAI. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019). Kira Radinsky. 2015. Data monopolists like Google are threatening the economy. Harvard Business Review 2 (2015). Evani Radiya-Dixit and Florian Tramèr. 2021. Data Poisoning Won’t Save You From Facial Recognition. arXiv preprint arXiv:2106.14851 (2021). Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv preprint arXiv:2112.11446 (2021). Colin Raffel. 2021. A Call to Build Models Like We Build Open-Source Software. https://colinraffel.com/blog/a-call-to-build- models-like-we-build-open-source-software.html. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683 (2019). Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein. 2017. On the expressive power of deep neural networks. In international conference on machine learning. PMLR, 2847–2854. Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. 2019. Transfusion: Understanding Transfer Learning for Medical Imaging. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc. https://proceedings.neurips.cc/paper/2019/ file/eb1e78328c46506b46a4ac4a1e378b91-Paper.pdf Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations toward Training Trillion Parameter Models. In SC20: International Conference for High Performance Computing, Networking, Storage and Ana
Source Document
→ report.pdf
page 198