Real-data validation of Bayesian optimal experimental design for closed-box scientific simulators
0score · Sign in to vote
Problem statement
Bayesian optimal experimental design (BOED) exists to spend real experimental resources better. The last five years produced a line of methods that make BOED feasible for simulators whose likelihood is implicit and whose code is not differentiable: mutual-information neural estimation (MINEBED), amortized design policies (iDAD), reinforcement-learning design agents (RL-BOED), and joint surrogate-plus-design training (SBI-BOED), with flow-based variational variants covering the differentiable regime. Every one of these papers justifies itself by the cost of physical experiments, with high-throughput screening in drug development as the canonical example, and every one of them is evaluated without performing a physical experiment. The "experiment" performed at the optimized design is one more draw from the same simulator the method trains on, , under the simulator's own data-generating process. That evaluation loop cannot detect the failure mode that matters in deployment: the simulator is misspecified with respect to the physical system, and a design chosen to be maximally informative under the simulator concentrates measurements exactly where the simulator's assumptions bind hardest. Whether optimized designs retain their advantage under real misspecification is unmeasured in this literature. The gap is compounded by a documented divergence between the objective being optimized and the inference quality practitioners need. Expected information gain (EIG) is estimated through variational mutual-information bounds, and papers now report cases where the method with the best EIG produces the worst-calibrated posterior: on a BMP signaling model, MINEBED-BO reaches EIG against SBI-BOED's , while its calibration statistic (L-C2ST, lower is better) is against . The EIG winner is the calibration loser. Inside a simulator this is a caveat; on a physical system, where the posterior drives a decision, it is the difference between the method helping and the method misleading. There is no accepted criterion for which quantity a real deployment should optimize and report. A resolution is a big win because it converts a decade of proxy-metric methodology into an instrument scientists can trust with scarce experiments. A demonstrated, matched-budget advantage on a physical system, or a well-powered demonstration of no advantage, would settle whether simulator-trained design optimization survives contact with reality, and would fix the reporting standard for every subsequent method paper in this line.
Current state
Design optimization for implicit-likelihood models is technically mature. MINEBED (Kleinegesse and Gutmann, 2020) optimizes mutual-information neural estimators with Bayesian optimization or simulator gradients. iDAD (Ivanova et al., 2021) amortizes a design policy but requires a differentiable simulator and on the order of simulator calls. RL-BOED (Blau et al., 2022) removes the differentiability requirement at the cost of tens of millions of simulated experiments. SBI-BOED (Zaballa and Hui, UAI 2026) trains a normalizing-flow likelihood and the designs jointly through an InfoNCE- bound, needs no simulator gradients, and runs at roughly calls. Flow-based variational BOED (Dong et al., 2025; Shen et al., 2025) covers the differentiable and reinforcement-learning regimes. Benchmarks are standardized: linear design problems, the SIR epidemic model, BMP pathway signaling, with calibration checked by simulation-based calibration and L-C2ST (Linhart et al., 2023). The problem stays open for two reasons. First, every result above is simulator-internal. Even the biological BMP study performs its "experiments" as convex-optimization solves of the model, not as measurements; and the SBI community has separately shown that posterior approximations can be confidently wrong (Hermans et al., 2022), which is precisely the failure a misspecified real system would amplify. Second, the field's own results disagree on which metric certifies success: EIG rankings and calibration rankings invert across tasks (Zaballa and Hui, Table 2; Rainforth et al., 2024 survey the estimator landscape), so a real-data trial has no settled target to pre-register against. Closing the loop needs wet-lab rounds that cost days and money per design iteration, and no group has yet paid that cost with a matched-budget baseline arm in place.
Resolution criteria
A paper solves this Grand Challenge when all four of the following hold: 1. **Physical head-to-head.** Designs are selected for a physical experiment (wet-lab, hardware, or field measurement) by a BOED method that treats the simulator as closed-box (implicit likelihood, no gradients through the simulator), and are executed alongside a matched-budget baseline arm (random, Sobol or grid, or the domain's standard protocol) with the same number of physical measurements per arm. Design choices are committed, pre-registered or time-stamped, before the outcomes are observed. 2. **Downstream inference measured on real data.** The comparison reports posterior quality from the real measurements: held-out predictive accuracy and at least one calibration diagnostic (simulation-based calibration, L-C2ST, or empirical coverage of credible sets). The optimized arm shows an advantage with uncertainty quantified over at least 3 independent replicates. A well-powered null result, showing no advantage under the same protocol, also resolves the challenge, in the negative direction. 3. **Proxy versus outcome measured.** The paper reports the information-gain estimate used for design selection alongside the downstream metrics, so the EIG-versus-calibration divergence documented in simulators is measured on a physical system at least once. 4. **Recomputable.** Code, the simulator, and the raw measurement data are released, sufficient to recompute the design selection and the posterior comparisons.
Citations
- Zaballa, V. D. and Hui, E. E. Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design. UAI 2026. arXiv:2502.08004.
- Kleinegesse, S. and Gutmann, M. U. Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation. ICML 2020. arXiv:2002.08129.
- Ivanova, D. R., Foster, A., Kleinegesse, S., Gutmann, M. U., and Rainforth, T. Implicit Deep Adaptive Design: Policy-Based Experimental Design without Likelihoods. NeurIPS 2021. arXiv:2111.02329.
- Blau, T., Bonilla, E. V., Chades, I., and Dezfouli, A. Optimizing Sequential Experimental Design with Deep Reinforcement Learning. ICML 2022. arXiv:2202.00821.
- Foster, A., Jankowiak, M., Bingham, E., Horsfall, P., Teh, Y. W., Rainforth, T., and Goodman, N. Variational Bayesian Optimal Experimental Design. NeurIPS 2019. arXiv:1903.05480.
- Dong, J., Jacobsen, C., Khalloufi, M., Akram, M., Liu, W., Duraisamy, K., and Huan, X. Variational Bayesian optimal experimental design with normalizing flows. Computer Methods in Applied Mechanics and Engineering 433, 117457, 2025.
- Shen, W., Dong, J., and Huan, X. Variational sequential optimal experimental design using reinforcement learning. Computer Methods in Applied Mechanics and Engineering 444, 118068, 2025.
- Rainforth, T., Foster, A., Ivanova, D. R., and Bickford Smith, F. Modern Bayesian Experimental Design. Statistical Science 39(1), 100-114, 2024.
- Hermans, J., Delaunoy, A., Rozet, F., Wehenkel, A., Begy, V., and Louppe, G. A Trust Crisis In Simulation-Based Inference? Your Posterior Approximations Can Be Unfaithful. TMLR 2022. arXiv:2110.06581.
- Linhart, J., Gramfort, A., and Rodrigues, P. L. C. L-C2ST: Local Diagnostics for Posterior Approximations in Simulation-Based Inference. NeurIPS 2023. arXiv:2306.03580.
- Paul, S. M., Mytelka, D. S., Dunwiddie, C. T., Persinger, C. C., Munos, B. H., Lindborg, S. R., and Schacht, A. L. How to improve R&D productivity: the pharmaceutical industry's grand challenge. Nature Reviews Drug Discovery 9(3), 203-214, 2010.