| Inferred field | Initial condition | Conductivity | Sheer modulus | |
| Training Data | Rectangular | MNIST | Circles | Circles |
| $ N_x $ | 784 | 784 | 4096 | 3136 |
| $ N_z$ | 3 | 100 | 50 | 50 |
| $ \Upsilon=\left\lfloor N_x / N_z\right\rfloor $ | 261 | 7 | 81 | 62 |
In this work, we train conditional Wasserstein generative adversarial networks to effectively sample from the posterior of physics-based Bayesian inference problems. The generator is constructed using a U-Net architecture, with the latent information injected using conditional instance normalization. The former facilitates a multiscale inverse map, while the latter enables the decoupling of the latent space dimension from the dimension of the measurement, and introduces stochasticity at all scales of the U-Net. We solve PDE-based inverse problems to demonstrate the performance of our approach in quantifying the uncertainty in the inferred field. Further, we show the generator can learn inverse maps which are local in nature, which in turn promotes generalizability when testing with out-of-distribution samples.
| Citation: |
Table 1. Dimension reduction with cGANs
| Inferred field | Initial condition | Conductivity | Sheer modulus | |
| Training Data | Rectangular | MNIST | Circles | Circles |
| $ N_x $ | 784 | 784 | 4096 | 3136 |
| $ N_z$ | 3 | 100 | 50 | 50 |
| $ \Upsilon=\left\lfloor N_x / N_z\right\rfloor $ | 261 | 7 | 81 | 62 |
Table 2. Hyper-parameters for cWGAN
| Inferred field | Initial condition | Conductivity | Sheer modulus | |
| Training Data | Rectangular | MNIST | Circles | Circles |
| Training samples | 10,000 | 10,000 | 8000 | 8000 |
| $ N_x $ | $ 28\times28 $ | $ 28\times28 $ | $ 64\times64 $ | $ 56\times56 $ |
| $ N_z$ | Multiple | 100 | 50 | 50 |
| Batch size | 50 | 50 | 64 | 64 |
| Activation param. | 0.1 | 0.1 | 0.2 | 0.2 |
| $ n_\text{critic}/n_\text{gen} $ | 4 | 4 | 5 | 5 |
| [1] |
J. Adler and O. Öktem, Deep Bayesian inversion, arXiv: 1811.05910.
|
| [2] |
J. Adler and Ozan Öktem, Solving ill-posed inverse problems using iterative deep neural networks, Inverse Problems, 33 (2017), 124007.
doi: 10.1088/1361-6420/aa9581.
|
| [3] |
A. Almahairi, S. Rajeshwar, A. Sordoni, P. Bachman and A. Courville, Augmented cyclegan: Learning many-to-many mappings from unpaired data, in International Conference on Machine Learning, (2018), 195-204.
|
| [4] |
M. S. Alnæs, J. Blechta, J. E. Hake, A. Johansson, B. Kehlet, A. Logg, C. N. Richardson, J. Ring, M. E. Rognes and G. N. Wells, The FEniCS project version 1.5., Archive of Numerical Software, 3 (2015).
|
| [5] |
P. E. Barbone and A. A. Oberai, Elastic modulus imaging: some exact solutions of the compressible elastography inverse problem, Physics in Medicine & Biology, 52 (2007), 1577.
doi: 10.1088/0266-5611/20/1/017.
|
| [6] |
P. E. Barbone and A. A. Oberai, A review of the mathematical and computational foundations of biomechanical imaging, Computational Modeling in Biomechanics, (2010), 375-408.
doi: 10.1007/978-90-481-3575-2_13.
|
| [7] |
T. F. Chan and P. C. Hansen, Some applications of the rank revealing qr factorization, SIAM Journal on Scientific and Statistical Computing, 13 (1992), 727-741.
doi: 10.1137/0913043.
|
| [8] |
I. J. D. Craig and J. C. Brown, Inverse Problems in Astronomy, A Guide To Inversion Strategies for Remotely Sensed Data, Boston: A. Hilger, 1986.
|
| [9] |
M. Dashti and A. M. Stuart, The Bayesian approach to inverse problems, in Handbook of Uncertainty Quantification, Springer International Publishing, Cham, (2017), 311-428.
doi: 10.1007/978-3-319-12385-1_7.
|
| [10] |
J. F. Dord, S. Goenezen, A. A. Oberai, P. E. Barbone, J. F. Jiang, T. J. Hall and T. Pavan, Validation of quantitative linear and nonlinear compression elastography, Ultrasound Elastography for Biomedical Applications and Medicine, (2018), 129-142.
|
| [11] |
V. Dumoulin, J. Shlens and M. Kudlur, A learned representation for artistic styl, arXiv: 1610.07629.
|
| [12] |
H. W. Engl, M. Hanke and A. Neubauer, Regularization of Inverse Problems, Mathematics and Its Applications, Vol. 375, Springer Science & Business Media, 1996.
|
| [13] |
H. Goh, S. Sheriffdeen, J. Wittmer and T. Bui-Thanh, Solving Bayesian inverse problems via variational autoencoders, Proceedings of the 2nd Mathematical and Scientific Machine Learning Conference, 145 (2022), 386-425.
|
| [14] |
W. P. Gouveia and J. A. Scales, Resolution of seismic waveform inversion: Bayes versus Occam, Inverse Problems, 13 (1997), 323-349.
doi: 10.1088/0266-5611/13/2/009.
|
| [15] |
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin and A. C. Courville, Improved training of wasserstein gans, in Advances in Neural Information Processing Systems, (2017), 5767–5777.
|
| [16] |
J. Hadamard, Sur les problèmes aux dérivées partielles et leur signification physique, Princeton University Bulletin, (1902), 49-52.
|
| [17] |
S. Huang, J. Xiang, H. Du and X. Cao, Inverse problems in atmospheric science and their application, Journal of Physics: Conference Series, 12 (2005), 45-57.
doi: 10.1088/1742-6596/12/1/005.
|
| [18] |
C. Jackson, M. K. Sen and P. L. Stoffa, An efficient stochastic Bayesian approach to optimal parameter and uncertainty estimation for climate model predictions, Journal of Climate, 17 (2004), 2828-2841.
|
| [19] |
T. Kadeethum, D. O'Malley, J. N. Fuhg, Y. Choi, J. Lee, H. S. Viswanathan and N. Bouklas, A framework for data-driven solution and parameter estimation of PDEs using conditional generative adversarial networks, Nature Computational Science, 1 (2021).
|
| [20] |
D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, arXiv: 1412.6980.
|
| [21] |
Y. LeCun, C. Cortes and C. J. Burges, Mnist handwritten digit database, ATT Labs, 2010.
|
| [22] |
F. Natterer, The Mathematics of Computerized Tomography, Society for Industrial and Applied Mathematics, 2001.
doi: 10.1137/1.9780898719284.
|
| [23] |
G. Ongie, A. Jalal, C. A. Metzler, R. G. Baraniuk, A. G Dimakis and R. Willett, Deep learning techniques for inverse problems in imaging, IEEE Journal on Selected Areas in Information Theory, 1 (2020), 39-56.
|
| [24] |
D. V. Patel and A. A. Oberai, Gan-based priors for quantifying uncertainty in supervised learning, SIAM/ASA Journal on Uncertainty Quantification, 9 (2021), 1314-1343.
doi: 10.1137/20M1354210.
|
| [25] |
D. V. Patel, D. Ray and A. A. Oberai, Solution of physics-based Bayesian inverse problems with deep generative priors, Computer Methods in Applied Mechanics and Engineering, 400 (2022).
doi: 10.1016/j.cma.2022.115428.
|
| [26] |
T. Z. Pavan, E. L. Madsen, G. R. Frank, J. F. Jiang, A. A. O. Carneiro and T. J. Hall, A nonlinear elasticity phantom containing spherical inclusions, Physics in Medicine & Biology, 57 (2012), 4787.
|
| [27] |
Y. Qian, M. Forghani, J. H. Lee, M. Farthing, T. H. P. Kitanidis and E. Darve, Application of deep learning-based interpolation methods to nearshore bathymetry, arXiv: 2011.09707.
|
| [28] |
G. Rizzuti, A. Siahkoohi, P. A. Witte and F. J. Herrmann, Parameterizing uncertainty by deep invertible networks: An application to reservoir characterization, in SEG Technical Program Expanded Abstracts 2020, Society of Exploration Geophysicists, (2020), 1541-1545.
|
| [29] |
O. Ronneberger, P. Fischer and T. Brox, U-net: Convolutional networks for biomedical image segmentation, in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, Springer International Publishing, (2015), 234-241.
|
| [30] |
C. Villani, Optimal Transport: Old and New, Grundlehren der mathematischen Wissenschaften, Springer Berlin Heidelberg, 2008.
doi: 10.1007/978-3-540-71050-9.
|
| [31] |
J. Whang, E. Lindgren and A. Dimakis, Composing normalizing flows for inverse problems, In Proceedings of the 38th International Conference on Machine Learning, (2021), 11158-11169.
|
| [32] |
L. Yang, D. Zhang and G. E. M. Karniadakis, Physics-informed generative adversarial networks for stochastic differential equations, SIAM Journal on Scientific Computing, 42 (2020), A292–A317.
doi: 10.1137/18M1225409.
|
| [33] |
Öz Yilmaz, Seismic Data Analysis: Processing, Inversion, and Interpretation of Seismic Data, Society of Exploration Geophysicists, 2001.
doi: 10.1190/1.9781560801580.
|
Architecture of generator and critic used in the conditional GAN. The spatial dimension
A schematic representation of the subsets
Samples from rectangular dataset used to train the cWGAN. The clean measurements are also shown to contextualize the amount of noise added
Test sample with reference mean and SD
Mean and SD computed with 800 samples of
Most important samples ranked left to right using RRQR algorithm on 800 samples for
Samples from MNIST dataset used to train the cWGAN. The clean measurements are also shown to contextualize the amount of noise added
Inferring initial condition for test samples chosen from the same distribution as the training set (MNIST prior)
Comparing cWGAN and cWGAN-stacked for inferring initial condition (MNIST prior)
Inferring initial condition for OOD test samples (notMNIST prior)
The profiles of
The value of
Average component-wise gradient of
Samples from the dataset that was used to train the network for inferring the conductivity. Each x sample consists of a circular inclusion with a randomly chosen contrast value
Inferring conductivity for test samples generated from circular priors (same distribution as the training set)
Inferring conductivity for OOD samples generated with elliptical priors
Inferring conductivity for OOD samples involving two circles
Average component-wise gradient data from the network used for conductivity inference. The red marker denotes the component/location (of
Samples from the dataset (circular priors) that was used to train the network for inferring shear modulus
Inference on an experimentally measured displacement field
Key components used to build the generator and critic networks