\`x^2+y_1+z_12^34\`
Advanced Search
Article Contents
Article Contents

A modification of adaptive moment estimation (adam) for machine learning

  • *Corresponding author: Qiang Long

    *Corresponding author: Qiang Long
Abstract / Introduction Full Text(HTML) Figure(7) / Table(8) Related Papers Cited by
  • In deep learning, the accuracy and generalization ability of the model largely depend on the optimization of the loss function. Up to now, dozens of optimization methods have been used in the deep learning models. Among them, stochastic gradient descent (SGD) is a popular and widely used method, and most of the other up-to-date optimization methods are variants or improvements of the original SGD. Among all the variations and improvements, the adaptive moment estimation (Adam) is one of the classics. However, Adam has also been pointed out to have non-convergence or error-convergence. Combining with the improvement points of the existing algorithms, this paper proposes an improved algorithm based on Adam, called NewAdam. NewAdam is modified from Adam in both search direction and learning rate. We perform a theoretical analysis on it and conduct numerical experiments on three data sets and two network architectures to illustrate the effectiveness of NewAdam.

    Mathematics Subject Classification: Primary: 68T07; Secondary: 90C26.

    Citation:

    \begin{equation} \\ \end{equation}
  • 加载中
  • Figure 1.  The optimization paths of algorithms under $ f(x) = 0.5x_1^2+2x_2^2 $

    Figure 2.  Accuracy (left) and loss function values (right) of MNIST data set based on AlexNet

    Figure 5.  Accuracy (left) and loss function values (right) of Fashion-MNIST data set based on VGG16

    Figure 3.  Accuracy (left) and loss function values (right) of MNIST data set based on VGG16

    Figure 4.  Accuracy (left) and loss function values (right) of Fashion-MNIST data set based on AlexNet

    Figure 6.  Accuracy (left) and loss function values (right) of Cifar-10 data set based on AlexNet

    Figure 7.  Accuracy (left) and loss function values (right) of Cifar-10 data set based on VGG16

    Table 1.  The optimization results of algorithms under $ f(x) = 0.5x_1^2+2x_2^2 $

    algorithm $ x_1 $ $ x_2 $ algorithm $ x_1 $ $ x_2 $
    SGD $ -1.20e-4 $ $ 0.00 $ Adam $ -2.34e-3 $ $ -5.41e-4 $
    Momentum $ 1.43e-3 $ $ -1.06e-3 $ Nadam $ 0.00 $ $ 1.9e-5 $
    NAG $ 8.82e-4 $ $ 3.45e-4 $ Amsgrad $ 2.06e-3 $ $ 8.71e-3 $
    Adagrad $ 0.00 $ $ -2.76e-4 $ NewAdam $ 0.00 $ $ 1.0e-5 $
    RMSProp $ -9.77e-2 $ $ 7.04e-2 $
     | Show Table
    DownLoad: CSV

    Table 2.  Experimental parameter settings of multivariate neural network model

    $ \eta $ $ \beta_{1} $ $ \beta_{2} $ $ \epsilon $
    0.001 0.99 0.999 $ 1e-8 $
     | Show Table
    DownLoad: CSV

    Table 3.  Results of the MNIST data set based on AlexNet

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 99.43% 0.019 99.06% 0.034 790
    Amsgrad 99.33% 0.023 99.01% 0.041 825
    Nadam 99.27% 0.025 99.00% 0.035 843
    NewAdam 99.68% 0.011 99.09% 0.030 853
     | Show Table
    DownLoad: CSV

    Table 4.  Results of the MNIST data set based on VGG16

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 99.39% 0.024 99.31% 0.034 1984
    Amsgrad 99.05% 0.034 99.05% 0.041 2120
    Nadam 99.64% 0.020 99.38% 0.029 2031
    NewAdam 99.05% 0.034 99.05% 0.038 1841
     | Show Table
    DownLoad: CSV

    Table 5.  Results of the Fashion-MNIST data set based on AlexNet

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 94.28% 0.156 92.26% 0.233 800
    Amsgrad 95.00% 0.137 92.18% 0.239 824
    Nadam 95.83% 0.117 92.70% 0.223 838
    NewAdam 95.33% 0.129 92.70% 0.230 940
     | Show Table
    DownLoad: CSV

    Table 6.  Results of the Fashion-MNIST data set based on VGG16

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 94.19% 0.163 92.52% 0.222 2638
    Amsgrad 97.08% 0.086 92.21% 0.269 2578
    Nadam 94.39% 0.158 92.77% 0.223 2599
    NewAdam 98.29% 0.054 92.59% 0.315 1827
     | Show Table
    DownLoad: CSV

    Table 7.  Results of the Cifar-10 data set based on AlexNet

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 73.37% 0.760 72.37% 0.836 1336
    Amsgrad 74.33% 0.739 73.48% 0.783 1367
    Nadam 73.91% 0.745 72.34% 0.811 1414
    NewAdam 75.78% 0.696 74.21% 0.770 1442
     | Show Table
    DownLoad: CSV

    Table 8.  Results of the Cifar-10 data set based on VGG16

    Algorithm Training set Testing set time(s)
    acc loss acc loss
    Adam 87.91% 0.389 85.24% 0.500 1717
    Amsgrad 87.73% 0.390 85.51% 0.480 1769
    Nadam 88.24% 0.369 86.26% 0.471 1850
    NewAdam 89.40% 0.338 86.74% 0.453 1804
     | Show Table
    DownLoad: CSV
  • [1] S. Bhattacharya, P. K. R. Maddikunta, Q. V. Pham, T. R. Gadekallu, C. L. Chowdhary, M. Alazab, M. J. Piran, et al., Deep learning and medical image processing for coronavirus (covid-19) pandemic: A survey, Sustainable Cities and Society, 65 (2021), 102589.
    [2] T. Dozat, Incorporating nesterov momentum into adam, ICLR 2016 Workshop, (2016), 1-4.
    [3] J. DuchiE. Hazan and Y. Singer, Adaptive subgradient methods for online learning and stochastic optimization, Journal of Machine Learning Research, 12 (2011), 2121-2159. 
    [4] I. GuellilH. SaâdaneF. AzouaouB. Gueni and D. Nouvel, Arabic natural language processing: An overview, Journal of King Saud University-Computer and Information Sciences, 33 (2021), 497-507. 
    [5] D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, arXiv preprint, arXiv: 1412.6980, (2014).
    [6] A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, Handbook of Systemic Autoimmune Diseases, 1 (2009).
    [7] A. Krizhevsky, I. Sutskever and G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in Neural Information Processing Systems, 25 (2012).
    [8] X. MaY. NiuL. GuY. WangY. ZhaoJ. Bailey and F. Lu, Understanding adversarial attacks on deep learning based medical image analysis systems, Pattern Recognition, 110 (2021), 107332. 
    [9] M. MasudG. MuhammadH. AlhumyaniS. S. AlshamraniO. CheikhrouhouS. Ibrahim and M. S. Hossain, Deep learning-based intelligent face recognition in iot-cloud environment, Computer Communications, 152 (2020), 215-222. 
    [10] Y. E. Nesterov, A method for solving a convex programming problem with convergence rate $o(1/k^2)$, Soviet Mathematics-Doklady, 27 (1983), 543-547. 
    [11] N. Qian, On the momentum term in gradient descent learning algorithms, Neural Networks, 12 (1999), 145-151. 
    [12] X. QiuT. SunY. XuY. ShaoN. Dai and X. Huang, Pre-trained models for natural language processing: A survey, Science China Technological Sciences, 63 (2020), 1872-1897. 
    [13] S. J. Reddi, S. Kale and S. Kumar, On the convergence of adam and beyond, arXiv preprint, arXiv: 1904.09237, (2019).
    [14] H. Robbins and S. Monro, A stochastic approximation method, The Annals of Mathematical Statistics, 22 (1951), 400-407.  doi: 10.1214/aoms/1177729586.
    [15] S. Ruder, An overview of gradient descent optimization algorithms, arXiv preprint, arXiv: 1609.04747, (2016).
    [16] S. Shalev-Shwartz, Y. Singer and N. Srebro, Pegasos: Primal estimated sub-gradient solver for svm, Proceedings of the 24th International Conference on Machine Learning, (2007), 807-814.
    [17] O. Shamir and T. Zhang, Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes, International Conference on Machine Learning, (2013), 71-79.
    [18] M. ShenH. YuL. ZhuK. XuQ. Li and J. Hu, Effective and robust physical-world attacks on deep learning face recognition systems, IEEE Transactions on Information Forensics and Security, 16 (2021), 4063-4077. 
    [19] K. Simonyan and A. Zisserman, Very deep convolutional networks for large-scale image recognition, International Conference on Learning Representations (ICLR), (2015), 1-14.
    [20] R. Stewart and S. Velupillai, Applied natural language processing in mental health big data, Neuropsychopharmacology, 46 (2021), 252. 
    [21] P. T. Tran and et al., On the convergence proof of amsgrad and a new version, IEEE Access, 7 (2019), 61706-61716. 
    [22] J. WangH. ZhuS. H. Wang and Y. D. Zhang, A review of deep learning on medical image analysis, Mobile Networks and Applications, 26 (2021), 351-380. 
    [23] H. Xiao, K. Rasul and R. Vollgraf, Fashion-mnist: A novel image dataset for benchmarking machine learning algorithms, arXiv preprint, arXiv: 1708.07747, (2017).
    [24] M. D. Zeiler, Adadelta: An adaptive learning rate method, arXiv preprint, arXiv: 1212.5701, (2012).
    [25] Y. Zhu and Y. Jiang, Optimization of face recognition algorithm based on deep learning multi feature fusion driven by big data, Image and Vision Computing, 104 (2020), 104032. 
    [26] M. Zinkevich, Online convex programming and generalized infinitesimal gradient ascent, International Conference on Machine Learning (ICML), (2003), 928-936.
  • 加载中

Figures(7)

Tables(8)

SHARE

Article Metrics

HTML views(12539) PDF downloads(2819) Cited by(0)

Access History

Other Articles By Authors

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return