Simulation-based inference with deep ensembles: evaluating calibration uncertainty and detecting model misspecification
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
IOP Publishing
https://doi.org/10.1088/2632-2153/ae3103
https://doi.org/10.1088/2632-2153/ae3103
Abstract
Description
Acknowledgements: We thank Jonas Elias El Gammal, Michele Mancarella, Michael Williams, and Christoph Weniger for their comments on a nearly final version of this draft. M P would like to thank Lorenzo Cevolani for very useful comments on a draft of this work. M P and J A acknowledge the hospitality of Imperial College London, which provided office space during some parts of this project. J A is supported by a fellowship from the Kavli Foundation. The work of MP is supported by the Comunidad de Madrid under the Programa de Atracción de Talento Investigador with number 2024-T1TEC-3134.
Funder: Kavli Foundation; doi: http://dx.doi.org/10.13039/100001201
Simulation-based inference (SBI) offers a principled and flexible framework for conducting Bayesian inference in any situation where forward simulations are feasible. However, validating the accuracy and reliability of the inferred posteriors remains a persistent challenge. In this work, we point out a simple diagnostic approach rooted in ensemble learning methods to assess the internal consistency of SBI outputs that does not require access to the true posterior. By training multiple neural estimators under identical conditions and evaluating their pairwise Kullback–Leibler (KL) divergences, we define a consistency criterion that quantifies agreement across the ensemble. We highlight two core use cases for this framework: (a) for generating a robust estimate of the systematic uncertainty in parameter reconstruction associated with the training procedure, and (b) for detecting possible model misspecification when using trained estimators on real data. We also demonstrate the relationship between significant KL divergences and issues such as insufficient convergence due to, e.g. too low a simulation budget, or intrinsic variance in the training process. Overall, this ensemble-based diagnostic framework provides a lightweight, scalable, and model-agnostic tool for enhancing the trustworthiness of SBI in scientific applications.
Funder: Kavli Foundation; doi: http://dx.doi.org/10.13039/100001201
Simulation-based inference (SBI) offers a principled and flexible framework for conducting Bayesian inference in any situation where forward simulations are feasible. However, validating the accuracy and reliability of the inferred posteriors remains a persistent challenge. In this work, we point out a simple diagnostic approach rooted in ensemble learning methods to assess the internal consistency of SBI outputs that does not require access to the true posterior. By training multiple neural estimators under identical conditions and evaluating their pairwise Kullback–Leibler (KL) divergences, we define a consistency criterion that quantifies agreement across the ensemble. We highlight two core use cases for this framework: (a) for generating a robust estimate of the systematic uncertainty in parameter reconstruction associated with the training procedure, and (b) for detecting possible model misspecification when using trained estimators on real data. We also demonstrate the relationship between significant KL divergences and issues such as insufficient convergence due to, e.g. too low a simulation budget, or intrinsic variance in the training process. Overall, this ensemble-based diagnostic framework provides a lightweight, scalable, and model-agnostic tool for enhancing the trustworthiness of SBI in scientific applications.