\`x^2+y_1+z_12^34\`
Advanced Search
Article Contents
Article Contents

Efficient learning methods for large-scale optimal inversion design

This paper is handled by Andreas Mang as the guest editor.

The first author is supported by NSF grants DMS-1654175 and DMS-1723005. The second author is supported by NSF grant DMS-1723005 and DMS-215266. The third author is supported by EPSRC grant EP/T001593/1. The fourth author is supported by NSF grant DMS-1502640.

Abstract / Introduction Full Text(HTML) Figure(10) / Table(1) Related Papers Cited by
  • In this work, we investigate various approaches that use learning from training data to solve inverse problems, following a bi-level learning approach. We consider a general framework for optimal inversion design, where training data can be used to learn optimal regularization parameters, data fidelity terms, and regularizers, thereby resulting in superior variational regularization methods. In particular, we describe methods to learn optimal $ p $ and $ q $ norms for $ {\rm L}^p-{\rm L}^q $ regularization and methods to learn optimal parameters for regularization matrices defined by covariance kernels. We exploit efficient algorithms based on Krylov projection methods for solving the regularized problems, both at training and validation stages, making these methods well-suited for large-scale problems. Our experiments show that the learned regularization methods perform well even when there is some inexactness in the forward operator, resulting in a mixture of model and measurement error.

    Mathematics Subject Classification: Primary: 65F22, 65K10; Secondary: 62F15.

    Citation:

    \begin{equation} \\ \end{equation}
  • 加载中
  • Figure 1.  Four prototype true images used for generating the training set (top row) and validation set (bottom row) in the OID experiment with $ \mathit{\boldsymbol{\theta }} = [\lambda, p, q]^{ \top }$

    Figure 2.  For the image deblurring problem with model error and impulse noise, we provide scatter plots of RRE norms for OID$ _{\lambda, p, q} $, OID$ _{\lambda, 2, 2} $, OID$ _{\lambda, 1, 2} $, and OID$ _{\lambda, 2, 1} $ (where the inner problem is solving using MM-GKS). Results for OID$ _{\lambda, 2, 1}^{\rm vpal} $ correspond to using a variable projection augmented Lagrangian method to solve the inner problem. Each column of dots corresponds to one sample from the validation set, where the indices have been sorted based on the RRE norms for OID$ _{\lambda, p, q} $. As a further comparison, $ \lambda_{\rm opt}^j $ corresponds to RRE norms for (21), where the optimal regularization parameter is selected for each image using the learned $ \widehat p $ and $ \widehat q $

    Figure 3.  For one sample of the validation data set, we provide in the top row the true image and the observed image. In the second row, we provide the OID$ _{\lambda, p, q} $ reconstruction, the OID$ _{\lambda, 1, 2} $ reconstruction, and the reconstruction computed using the optimal regularization parameter for this image, which is provided for comparison purposes only. In the bottom row are reconstructions for OID$ _{\lambda, 2, 2} $ and OID$ _{\lambda, 2, 1} $, and OID$ _{\lambda, 2, 1}^{\rm vpal} $. RRE values are provided in the titles

    Figure 4.  Investigating the impact of model error on the overall noise statistics. From left to right in the top row, we provide pixel-wise values of the additive Gaussian noise (of level 1%), the model error, and the sum of these two errors. The density plot for the combined error is provided, along with the density function for $ p = 2 $ (corresponding to Gaussian noise) and the best density fit to the true errors $ {\bf A}(5.6, 5.6){\bf x}_{\rm true}- {\bf b} $, i.e., $ p = 1.3835 $

    Figure 5.  Seismic image reconstruction example. The true image (left) contains $ 256 \times 256 $ pixels and represents a smooth medium. The noisy sinogram image (right) represents projection data from a setup with $ 256 $ sources and $ 512 $ receivers

    Figure 6.  Seismic example - random samples from the training set

    Figure 7.  For the seismic example, we provide scatter plots of RRE norms for OID, OID-wgcv, SC-wgcv, and HyBR opt. Each column of dots corresponds to one sample from the validation set, where the indices have been sorted based on the RRE norms for OID

    Figure 8.  For the validation set of the seismic example, we provide histograms of the RRE norms for OID, OID-wgcv, and SC-wgcv

    Figure 9.  For the validation image in Figure 5, we provide reconstructions obtained with HyBR opt and OID reconstructions for the Mat$ \acute{\rm{e}} $rn and squared exponential kernels

    Figure 10.  Design objective for OID with the squared exponential kernel for the seismic example. The filled contour corresponds to OID, and the white point denotes the OID computed values

    Table 1.  Computed values of the hyperparameters for OID, along with the mean reconstruction errors for the validation set. OID with $ \lambda $ computed using WGCV corresponds to using OID-wgcv for estimating $ \mathit{\boldsymbol{\beta }} $ only. 'SC' corresponds to estimating $ \mathit{\boldsymbol{\beta }} $ directly from the sample covariance matrix as described in [20] and then using genHyBR with WGCV

    Mat$ \acute{\rm{e}} $rn $ \lambda $ $ \mathit{\boldsymbol{\beta }} $ $ {\cal P} $, validation
    OID 18.8313 5.0312, 0.3344 6.0833
    OID wgcv 12.2812, 0.3344 29.9412
    SC wgcv 123.1735, 0.2011 49.3447
    sq. exp. $ \lambda $ $ \beta $ $ {\cal P} $, validation
    OID 50.0500 0.2550 3.4468
    OID wgcv 0.3163 24.7547
    SC wgcv 0.2010 47.6977
     | Show Table
    DownLoad: CSV
  • [1] Global optimization toolbox, https://www.mathworks.com/help/gads/index.html?s_tid=CRUX_lftnav, Accessed: 2021-09-20.
    [2] Nasa, http:www.nasa.gov.
    [3] A. Alexanderian, P. J. Gloor, O. Ghattas et al., On Bayesian A-and D-optimal experimental designs in infinite dimensions, Bayesian Analysis, 11 (2016), 671-695. doi: 10.1214/15-BA969.
    [4] H. AntilZ. W. Di and R. Khatri, Bilevel optimization, deep learning and fractional Laplacian regularization with applications in tomography, Inverse Problems, 36 (2020), 064001.  doi: 10.1088/1361-6420/ab80d7.
    [5] F. Archetti and A. Candelieri, The acquisition function, in Bayesian Optimization and Data Science, Springer, (2019), 57-72.
    [6] S. ArridgeP. MaassO. Öktem and C.-B. Schönlieb, Solving inverse problems using data-driven models, Acta Numerica, 28 (2019), 1-174.  doi: 10.1017/s0962492919000059.
    [7] A. C. Atkinson and A. N. Donev, Optimum Experimental Designs, 1992.
    [8] J. F. Bard, Practical Bilevel Optimization: Algorithms and Applications, Vol. 30, Springer, Berlin, 2013. doi: 10.1007/978-1-4757-2836-1.
    [9] J. M. Bardsley and J. G. Nagy, Covariance-preconditioned iterative methods for nonnegatively constrained astronomical imaging, SIAM Journal on Matrix Analysis and Applications, 27 (2006), 1184-1197.  doi: 10.1137/040615043.
    [10] A. Björk, Numerical Methods for Least Squares Problems, SIAM, 1996.
    [11] R. D. BrownJ. M. Bardsley and T. Cui, Semivariogram methods for modeling Whittle–Matérn priors in Bayesian inverse problems, Inverse Problems, 36 (2020), 055006.  doi: 10.1088/1361-6420/ab762e.
    [12] A. BucciniM. Pasha and L. Reichel, Modulus-based iterative methods for constrained $\ell_p-\ell_q$ minimization, Inverse Problems, 36 (2020), 084001.  doi: 10.1088/1361-6420/ab9f86.
    [13] A. BucciniY. Park and L. Reichel, Numerical aspects of the nonstationary modified linearized Bregman algorithm, Applied Mathematics and Computation, 337 (2018), 386-398.  doi: 10.1016/j.amc.2018.05.044.
    [14] A. BucciniM. Pasha and L. Reichel, Projected Bregman in Krylov subspaces., Mathematics of Computation, 78 (2009), 1515-1536. 
    [15] J.-F. CaiS. Osher and Z. Shen, Linearized Bregman iterations for frame-based image deblurring, SIAM Journal on Imaging Sciences, 2 (2009), 226-252.  doi: 10.1137/080733371.
    [16] L. CalatroniC. CaoJ. C. De Los ReyesC.-B. Schönlieb and T. Valkonen, Bilevel approaches for learning of variational imaging models, Variational Methods: In Imaging and Geometric Control, 18 (2017), 2. 
    [17] L. CalatroniJ. C. De Los Reyes and C.-B. Schönlieb, Infimal convolution of data discrepancies for mixed noise removal, SIAM Journal on Imaging Sciences, 10 (2017), 1196-1233.  doi: 10.1137/16M1101684.
    [18] D. Calvetti and E. Somersalo, An Introduction to Bayesian Scientific Computing: Ten Lectures on Subjective Computing, Vol. 2, Springer Science & Business Media, 2007.
    [19] A. Chambolle and T. Pock, A first-order primal-dual algorithm for convex problems with applications to imaging, Journal of Mathematical Imaging and Vision, 40 (2011), 120-145.  doi: 10.1007/s10851-010-0251-1.
    [20] T. ChoJ. Chung and J. Jiang, Hybrid projection methods for large-scale inverse problems with mixed Gaussian priors, Inverse Problems, 37 (2021), 044002.  doi: 10.1088/1361-6420/abd29d.
    [21] J. Chung and S. Gazzola, Flexible Krylov methods for $\ell_p$ regularization, SIAM Journal on Scientific Computing, 41 (2019), S149-S171.  doi: 10.1137/18M1194456.
    [22] J. ChungM. Chung and D. P. O'Leary, Designing optimal spectral filters for inverse problems, SIAM Journal on Scientific Computing, 33 (2011), 3132-3152.  doi: 10.1137/100812938.
    [23] J. Chung and M. I. Español, Learning regularization parameters for general-form Tikhonov, Inverse Problems, 33 (2017), 074004.  doi: 10.1088/1361-6420/33/7/074004.
    [24] J. Chung and J. G. Nagy, An efficient iterative approach for large-scale separable nonlinear inverse problems, SIAM Journal on Scientific Computing, 31 (2010), 4654-4674.  doi: 10.1137/080732213.
    [25] J. Chung and A. K. Saibaba, Generalized hybrid iterative methods for large-scale Bayesian inverse problems, SIAM Journal on Scientific Computing, 39 (2017), S24-S46.  doi: 10.1137/16M1081968.
    [26] M. Chung and R. Renaut, The variable projected augmented Lagrangian method, arXiv: 2207.08216.
    [27] J. C. De los ReyesC.-B. Schönlieb and T. Valkonen, Bilevel parameter learning for higher-order total variation regularisation models, Journal of Mathematical Imaging and Vision, 57 (2017), 1-25.  doi: 10.1007/s10851-016-0662-8.
    [28] J. C. De los Reyes and C.-B. Schönlieb, Image denoising: learning the noise model via nonsmooth PDE-constrained optimization, Inverse Problems & Imaging, 7 (2013), 1183-1214.  doi: 10.3934/ipi.2013.7.1183.
    [29] J. C. De los Reyes and D. Villacís, Bilevel Optimization Methods in Imaging, Springer International Publishing, 2022.
    [30] S. Dempe, Foundations of Bilevel Programming, Kluwer Academic Publishers, New York, 2002.
    [31] S. Dempe, V. Kalashnikov, G. A. Pérez-Valdés and N. Kalashnykova, Bilevel Programming Problems, Springer, Berlin, 2015. doi: 10.1007/978-3-662-45827-3.
    [32] M. M. Dunlop, T. Helin and A. M. Stuart, Hyperparameter estimation in Bayesian MAP estimation: Parameterizations and consistency, arXiv: 1905.04365. doi: 10.5802/smai-jcm.62.
    [33] H. W. Engl, M. Hanke and A. Neubauer, Regularization of Inverse Problems, Vol. 375, Springer Science & Business Media, 1996.
    [34] E. EsserX. Zhang and T. F. Chan, A general framework for a class of first order primal-dual algorithms for convex optimization in imaging science, SIAM Journal on Imaging Sciences, 3 (2010), 1015-1046.  doi: 10.1137/09076934X.
    [35] S. GazzolaP. C. Hansen and J. G. Nagy, Ir tools: a matlab package of iterative regularization methods and large-scale test problems, Numerical Algorithms, 81 (2019), 773-811.  doi: 10.1007/s11075-018-0570-7.
    [36] R. B. Gramacy, Surrogates: Gaussian Process Modeling, Design, and Optimization for the Applied Sciences, Chapman and Hall/CRC, 2020. doi: 10.1201/9780367815493.
    [37] E. Haber and L. Tenorio, Learning regularization functionals-a supervised training approach, Inverse Problems, 19 (2003), 611.  doi: 10.1088/0266-5611/19/3/309.
    [38] E. HaberL. Horesh and L. Tenorio, Numerical methods for experimental design of large-scale linear ill-posed inverse problems, Inverse Problems, 24 (2008), 055012.  doi: 10.1088/0266-5611/24/5/055012.
    [39] E. HaberL. Horesh and L. Tenorio, Numerical methods for the design of large-scale nonlinear discrete ill-posed inverse problems, Inverse Problems, 26 (2010), 025002.  doi: 10.1088/0266-5611/26/2/025002.
    [40] M. Haltmeier and L. V. Nguyen, Regularization of inverse problems by neural networks, arXiv: 2006.03972.
    [41] K. HammernikT. KlatzerE. KoblerM. P. RechtD. K. SodicksonT. Pock and F. Knoll, Learning a variational network for reconstruction of accelerated MRI data, Magnetic Resonance in Medicine, 79 (2018), 3055-3071.  doi: 10.1007/978-3-319-66709-6.
    [42] P. C. Hansen, Discrete Inverse Problems: Insight and Algorithms, SIAM, 2010. doi: 10.1137/1.9780898718836.
    [43] P. C. Hansen and J. S. Jørgensen, Air tools ⅱ: algebraic iterative reconstruction methods, improved implementation, Numerical Algorithms, 79 (2018), 107-137.  doi: 10.1007/s11075-017-0430-x.
    [44] M. HintermüllerC. N. RautenbergT. Wu and A. Langer, Optimal selection of the regularization function in a weighted total variation model. part Ⅱ: Algorithm, its analysis and numerical tests, Journal of Mathematical Imaging and Vision, 59 (2017), 515-533.  doi: 10.1007/s10851-017-0736-2.
    [45] G. HollerK. Kunisch and R. C. Barnard, A bilevel approach for parameter learning in inverse problems, Inverse Problems, 34 (2018), 115012.  doi: 10.1088/1361-6420/aade77.
    [46] X. Huan and Y. Marzouk, Gradient-based stochastic optimization methods in bayesian experimental design, International Journal for Uncertainty Quantification, 4 (2014), 479-510.  doi: 10.1615/Int.J.UncertaintyQuantification.2014006730.
    [47] G. HuangA. LanzaS. MorigiL. Reichel and F. Sgallari, Majorization–minimization generalized krylov subspace methods for $\ell_p-\ell_q$ optimization applied to image restoration, BIT Numerical Mathematics, 57 (2017), 351-378.  doi: 10.1007/s10543-016-0643-8.
    [48] J. HuangM. Donatelli and R. H. Chan, Nonstationary iterated thresholding algorithms for image deblurring, Inverse Problems & Imaging, 7 (2013), 717-736.  doi: 10.3934/ipi.2013.7.717.
    [49] D. R. Hunter and K. Lange, A tutorial on MM algorithms, The American Statistician, 58 (2004), 30-37.  doi: 10.1198/0003130042836.
    [50] J. Kaipio and E. Somersalo, Statistical and Computational Inverse Problems, Vol. 160, Springer Science & Business Media, 2006.
    [51] J. P. KaipioV. KolehmainenM. Vauhkonen and E. Somersalo, Inverse problems with structural prior information, Inverse Problems, 15 (1999), 713.  doi: 10.1088/0266-5611/15/3/306.
    [52] M. Kubínová and J. G. Nagy, Robust regression for mixed Poisson–Gaussian model, Numerical Algorithms, 79 (2018), 825-851.  doi: 10.1007/s11075-017-0463-1.
    [53] K. Lange, MM Optimization Algorithms, SIAM, 2016. doi: 10.1137/1.9781611974409.ch1.
    [54] H. LiJ. SchwabS. Antholzer and M. Haltmeier, NETT: Solving inverse problems with deep neural networks, Inverse Problems, 36 (2020), 065005.  doi: 10.1088/1361-6420/ab6d57.
    [55] A. LucasM. IliadisR. Molina and A. K. Katsaggelos, Using deep neural networks for inverse problems in imaging: beyond analytical methods, IEEE Signal Processing Magazine, 35 (2018), 20-36. 
    [56] S. LunzA. HauptmannT. TarvainenC.-B. Schönlieb and S. Arridge, On learned operator correction in inverse problems, SIAM Journal on Imaging Sciences, 14 (2021), 92-127.  doi: 10.1137/20M1338460.
    [57] S. LunzO. Öktem and C.-B. Schönlieb, Adversarial regularizers in inverse problems, Advances in Neural Information Processing Systems, 31 (2018), 8507-8516. 
    [58] M. T. McCannK. H. Jin and M. Unser, Convolutional neural networks for inverse problems in imaging: A review, IEEE Signal Processing Magazine, 34 (2017), 85-95.  doi: 10.1109/TIP.2017.2713099.
    [59] G. OngieA. JalalC. A. M. R. G. BaraniukA. G. Dimakis and R. Willett, Deep learning techniques for inverse problems in imaging, IEEE Journal on Selected Areas in Information Theory, 1 (2020), 39-56.  doi: 10.1109/JSAIT.2020.2991563.
    [60] M. A. Osborne, R. Garnett and S. J. Roberts, Gaussian processes for global optimization, in $3^{rd}$ International Conference on Learning and Intelligent Optimization (LION3), (2009), 1-15.
    [61] J. Prost, A. Houdard, A. Almansa and N. Papadakis, Learning local regularization for variational image restoration, arXiv: 2102.06155.
    [62] F. Pukelsheim, Optimal Design of Experiments, SIAM, 2006. doi: 10.1137/1.9780898719109.
    [63] N. A. B. RiisY. Dong and P. C. Hansen, Computed tomography with view angle estimation using uncertainty quantification, Inverse Problems, 37 (2021), 065007.  doi: 10.1088/1361-6420/abf5ba.
    [64] P. Rodríguez and B. Wohlberg, Efficient minimization method for a generalized total variation functional, IEEE Transactions on Image Processing, 18 (2008), 322-332.  doi: 10.1109/TIP.2008.2008420.
    [65] L. RuthottoJ. Chung and M. Chung, Optimal experimental design for inverse problems with state constraints, SIAM Journal on Scientific Computing, 40 (2018), B1080-B1100.  doi: 10.1137/17M1143733.
    [66] F. SherryM. BenningJ. C. De los ReyesM. J. GravesG. MaierhoferG. WilliamsC.-B. Schönlieb and M. J. Ehrhardt, Learning the sampling pattern for MRI, IEEE Transactions on Medical Imaging, 39 (2020), 4310-4321.  doi: 10.1109/TMI.2020.3017353.
    [67] A. SinhaP. Malo and K. Deb, A review on bilevel optimization: from classical to evolutionary approaches and applications, IEEE Transactions on Evolutionary Computation, 22 (2017), 276-295. 
    [68] D. SmylT. N. TallmanJ. A. BlackA. Hauptmann and D. Liu, Learning and correcting non-Gaussian model errors, Journal of Computational Physics, 432 (2021), 110152.  doi: 10.1016/j.jcp.2021.110152.
    [69] M. UrquhartE. Ljungskog and S. Sebben, Surrogate-based optimisation using adaptively scaled radial basis functions, Applied Soft Computing, 88 (2020), 106050. 
    [70] T. Weise, Global optimization algorithms-theory and application, Self-Published Thomas Weise.
    [71] C. K. Williams, Gaussian Processes for Machine Learning, Vol. 2, MIT Press, 2006.
    [72] K. Zhang, W. Zuo, S. Gu and L. Zhang, Learning deep CNN denoiser prior for image restoration, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 3929-3938.
    [73] M. Zhu and T. Chan, An efficient primal-dual hybrid gradient algorithm for total variation image restoration, UCLA CAM Report, 34 (2008), 8-34.
  • 加载中

Figures(10)

Tables(1)

SHARE

Article Metrics

HTML views(8624) PDF downloads(652) Cited by(0)

Access History

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return