Mathematics

Handbook of Bayesian Deep Learning

Published on

Authors: Claudio Agostinelli, Laurence Aitchison, Emanuel Aldea, Richard Allmendinger, Sebastian Ament, Ben Anson, Julyan Arbel, Eytan Bakshy, Max Balandat, Nicola Bariletto, Freddie Bickford Smith, Battista Biggio, Michele Caprio, Neil Chada, Wenlin Chen, Wenlong Chen, Nathaël da Costa, Nico Daheim, Andreas Damianou, Michael Deistler, David Dunson, Nikita Durasov, Maurizio Filippone, Lucia Filippozzi, Gergely Flamich, Vincent Fortuin, Gianni Franchi, Giorgio Fumera, Yarin Gal, Wenbo Gong, Yuqi Gu, Cheng Guo, Nick Hauptvogel, Jiajun He, Philipp Hennig, José Miguel Hernández-Lobato, Nhat Ho, Clara Hoffmann, Aliaksandr Hubin, Christian Igel, Ajay Jasra, Matt Jones, Nadja Klein, Olivier Laurent, Kody Law, Emanuele Ledda, Bolian Li, Yingzhen Li, Xinzhu Liang, Jakob Macke, Stephan Mandt, Luckeciano Carvalho Melo, Benjamin Kurt Miller, Edward Milsom, Bruno Mlodozeniec, Bálint Mucsányi, Samuel Müller, Kevin Murphy, Eric Nalisnick, Marcos Negre Saura, Christopher Nemeth, Huy Nguyen, Khai Nguyen, Shreyas Padhy, Konstantina Palla, Wei Pan, Theodore Papamarkou, Dhruvesh Patel, Alina Peluso, Simon Pepin Lehalleur, Konstantinos Pitas, Patrick Pynadath, Tom Rainforth, Maxime Robeyns, Fabio Roli, Breeshey Roskams-Hieter, Simone Rossi, Tim Rudner, David Rügamer, Dino Sejdinovic, Yevgeny Seldin, Torben Sell, Louis Sharrock, Sumeetpal Singh, Joanna Sliwa, Emanuel Sommer, Qifan Song, Ba-Hien Tran, Richard Turner, Tycho van der Ouderaa, Mark van der Wilk, Mariia Vladimirova, Sara Wade, Xi Wang, Tobias Weber, Susan Wei, Christopher Wikle, Andrew Gordon Wilson, Lisa Wimmer, Adam Yang, Yibo Yang, Myungsoo Yoo, Cheng Zhang, Likun Zhang, Ruqi Zhang, Yuren Zhou

Bayesian deep learning (BDL) asks how the machinery of Bayesian inference—priors, posteriors, predictive distributions, marginal likelihoods, and the decisions they license—can be brought to bear on models with millions or billions of parameters, for which none of the classical guarantees or algorithms apply unmodified. This book covers that question in ten parts. Five of them develop the major families of approximate inference used in practice: Monte Carlo sampling (Part I), Laplace approximations (Part II), variational inference (Part III), ensembles (Part IV), and kernel and Gaussian-process methods (Part V). Two cover the modelling questions that are assumed as prerequisites by the aforementioned algorithms: what a prior over a neural network actually means (Part VI), and what the symmetries and non-identifiability of neural parameterisations do to inference and interpretation (Part VII). One is devoted to making all of this run at modern scale (Part VIII), one to worked applications and software (Part IX), and the last to topics at the current research frontier, from diffusion models and singular learning theory to causal, credal, and reinforcement-learning extensions (Part X). The parts are largely parallel rather than sequential: they are alternative and complementary answers to the same problem, and comparing them is a large part of what the book is for.