Research article Special Issues

Data-driven full waveform inversion with vision transformers

  • Published: 17 August 2026
  • Full waveform inversion (FWI) is a powerful technique for reconstructing subsurface physical parameters. However, conventional FWI often suffers from cycle-skipping and local-minimum issues. With the rapid development of deep learning, data-driven methods have attracted increasing attention in FWI, where the inversion problem is reformulated as an end-to-end reconstruction task and neural networks are employed to directly estimate velocity models from seismic data. Most existing approaches rely on convolutional neural networks (CNNs) to extract local features from input data. However, compared with natural images, seismic data exhibit strong long-range dependencies that CNNs struggle to model effectively, since information associated with a single reflection interface is spatially distributed across the seismic data. Although CNNs perform well in computer vision, their local feature extraction mechanism limits reconstruction performance in FWI. To address this issue, we have proposed a novel FWI method based on the vision transformer (ViT). Unlike CNN-based approaches, the proposed method leverages the self-attention mechanism to model long-range dependencies in the input data. By stacking multiple transformer blocks, global physical information in seismic data is effectively encoded into the inversion process, resulting in improved reconstruction performance. Numerical results demonstrated that the proposed method outperforms CNN-based approaches in terms of reconstruction accuracy.

    Citation: Feng Li, Shuo Peng, Zongzeng Li, Hongsun Fu, Jinlong Yuan. Data-driven full waveform inversion with vision transformers[J]. Electronic Research Archive, 2026, 34(10): 7096-7117. doi: 10.3934/era.2026307

    Related Papers:

  • Full waveform inversion (FWI) is a powerful technique for reconstructing subsurface physical parameters. However, conventional FWI often suffers from cycle-skipping and local-minimum issues. With the rapid development of deep learning, data-driven methods have attracted increasing attention in FWI, where the inversion problem is reformulated as an end-to-end reconstruction task and neural networks are employed to directly estimate velocity models from seismic data. Most existing approaches rely on convolutional neural networks (CNNs) to extract local features from input data. However, compared with natural images, seismic data exhibit strong long-range dependencies that CNNs struggle to model effectively, since information associated with a single reflection interface is spatially distributed across the seismic data. Although CNNs perform well in computer vision, their local feature extraction mechanism limits reconstruction performance in FWI. To address this issue, we have proposed a novel FWI method based on the vision transformer (ViT). Unlike CNN-based approaches, the proposed method leverages the self-attention mechanism to model long-range dependencies in the input data. By stacking multiple transformer blocks, global physical information in seismic data is effectively encoded into the inversion process, resulting in improved reconstruction performance. Numerical results demonstrated that the proposed method outperforms CNN-based approaches in terms of reconstruction accuracy.



    加载中


    [1] J. Virieux, S. Operto, An overview of full-waveform inversion in exploration geophysics, Geophysics, 74 (2009), WCC1–WCC26. https://doi.org/10.1190/1.3238367 doi: 10.1190/1.3238367
    [2] D. C. Liu, J. Nocedal, On the limited memory BFGS method for large scale optimization, Math. Program., 45 (1989), 503–528. https://doi.org/10.1007/BF01589116 doi: 10.1007/BF01589116
    [3] H. S. Fu, Y. Zhang, X. L. Li, Adaptive overcomplete dictionary learning-based sparsity-promoting regularization for full-waveform inversion, Pure Appl. Geophys., 178 (2021), 1–12. https://doi.org/10.1007/s00024-021-02662-w doi: 10.1007/s00024-021-02662-w
    [4] R. Fletcher, C. M. Reeves, Function minimization by conjugate gradients, Comput. J., 7 (1964), 149–154. https://doi.org/10.1093/comjnl/7.2.149 doi: 10.1093/comjnl/7.2.149
    [5] P. Mora, Nonlinear two dimensional elastic inversion of multioffset seismic data, Geophysics, 52 (1987), 1211–1228. https://doi.org/10.1190/1.1442384 doi: 10.1190/1.1442384
    [6] A. N. Tikhonov, A. V. Goncharsky, V. V. Stepanov, A. G. Yagola, Regularization methods, in Numerical Methods for the Solution of Ill-Posed Problems, Springer Netherlands, (1995), 7–63. https://doi.org/10.1007/978-94-015-8480-7_2
    [7] A. Asnaashari, R. Brossier, S. Garambois, F. Audebert, P. Thore, J. Virieux, Regularized seismic full-waveform inversion with prior model information, Geophysics, 78 (2013), R25–R36. https://doi.org/10.1190/geo2012-0104.1 doi: 10.1190/geo2012-0104.1
    [8] E. Esser, L. Guasch, T. van Leeuwen, A. Y. Aravkin, F. J. Herrmann, Total variation regularization strategies in full-waveform inversion, SIAM J. Imaging Sci., 11 (2018), 376–406. https://doi.org/10.1137/17M111328X doi: 10.1137/17M111328X
    [9] W. Zhang, M. Z. Xu, H. X. Yang, X. Wang, S. S. Zheng, X. Li, Data-driven deep learning approach for thrust prediction of solid rocket motors, Measurement, 225 (2024), 114051. https://doi.org/10.1016/j.measurement.2023.114051 doi: 10.1016/j.measurement.2023.114051
    [10] X. W. Wang, Y. H. Wang, X. C. Su, L. Wang, C. Lu, H. J. Peng, et al., Deep reinforcement learning-based air combat maneuver decision-making: Literature review, implementation tutorial and future direction, Artif. Intell. Rev., 57 (2024). https://doi.org/10.1007/s10462-023-10620-2
    [11] M. Araya-Polo, J. Jennings, A. Adler, T. Dahlke, Deep-learning tomography, Lead. Edge, 37 (2018), 58–66. https://doi.org/10.1190/tle37010058.1 doi: 10.1190/tle37010058.1
    [12] K. Hornik, M. Stinchcombe, H. White, Multilayer feedforward networks are universal approximators, Neural Netw., 2 (1989), 359–366. https://doi.org/10.1016/0893-6080(89)90020-8 doi: 10.1016/0893-6080(89)90020-8
    [13] C. Song, T. Alkhalifah, Wavefield reconstruction inversion via physics-informed neural networks, IEEE Trans. Geosci. Remote Sens., 60 (2021), 1–12. https://doi.org/10.1109/TGRS.2021.3123122 doi: 10.1109/TGRS.2021.3123122
    [14] F. X. Wu, Y. Li, Z. W. Fu, Q. L. He, Y. Chen, B. Han, et al., Robust physics-informed full waveform inversion via a transformer-based autoencoder, IEEE Trans. Geosci. Remote Sens., 63 (2025), 1–16. https://doi.org/10.1109/TGRS.2025.3597592 doi: 10.1109/TGRS.2025.3597592
    [15] A. Adler, M. Araya-Polo, T. Poggio, Deep learning for seismic inverse problems: Toward the acceleration of geophysical analysis workflows, IEEE Signal Process. Mag., 38 (2021), 89–119. https://doi.org/10.1109/MSP.2020.3037429 doi: 10.1109/MSP.2020.3037429
    [16] Y. Wu, Y. Z. Lin, Inversionnet: An efficient and accurate data-driven full waveform inversion, IEEE Trans. Comput. Imaging, 6 (2019), 419–433. https://doi.org/10.1109/TCI.2019.2956866 doi: 10.1109/TCI.2019.2956866
    [17] S. H. Feng, Y. Z. Lin, B. Wohlberg, Multiscale data-driven seismic full-waveform inversion with field data study, IEEE Trans. Geosci. Remote Sens., 60 (2021), 1–14. https://doi.org/10.1109/TGRS.2021.3114101 doi: 10.1109/TGRS.2021.3114101
    [18] Z. P. Zhang, Y. Z. Lin, Data-driven seismic waveform inversion: A study on the robustness and generalization, IEEE Trans. Geosci. Remote Sens., 58 (2020), 6900–6913. https://doi.org/10.1109/TGRS.2020.2977635 doi: 10.1109/TGRS.2020.2977635
    [19] C. Y. Deng, S. H. Feng, H. C. Wang, X. T. Zhang, P. Jin, Y. N. Feng, et al., OpenFWI: Large-scale multi-structural benchmark datasets for full waveform inversion, in Advances in Neural Information Processing Systems, Curran Associates, Inc., (2022), 6007–6020. https://doi.org/10.52202/068431-0435
    [20] X. Y. Zhang, F. Min, S. L. Pan, Q. Xu, X. Y. Min, G. J. Song, et al., DD-Net: Dual decoder network with curriculum learning for full waveform inversion, IEEE Trans. Geosci. Remote Sens., 62 (2024), 1–17. https://doi.org/10.1109/TGRS.2024.3358492 doi: 10.1109/TGRS.2024.3358492
    [21] X. Yan, K. Wu, Z. Q. J. Xu, Z. Ma, A deep learning approach for solving the inverse problem of the wave equation, CSIAM Trans. Appl. Math., 7 (2026), 29–61. https://doi.org/10.4208/csiam-am.SO-2024-0027 doi: 10.4208/csiam-am.SO-2024-0027
    [22] W. Ding, K. Ren, L. Zhang, Coupling deep learning with full waveform inversion, preprint, arXiv: 2203.01799.
    [23] F. S. Yang, J. W. Ma, FWIGAN: Full-waveform inversion via a physics-informed generative adversarial network, J. Geophys. Res.: Solid Earth, 128 (2023), e2022JB025493. https://doi.org/10.1029/2022JB025493 doi: 10.1029/2022JB025493
    [24] S. L. Yang, T. Alkhalifah, Y. X. Ren, B. Liu, Y. Y. Li, P. Jiang, Well-log information-assisted high-resolution waveform inversion based on deep learning, IEEE Trans. Geosci. Remote Sens. Lett., 20 (2023), 1–5. https://doi.org/10.1109/LGRS.2023.3234211 doi: 10.1109/LGRS.2023.3234211
    [25] P. Jin, X. T. Zhang, Y. P. Chen, S. X. Huang, Z. C. Liu, Y. Z. Lin, Unsupervised learning of full-waveform inversion: Connecting cnn and partial differential equation in a loop, preprint, arXiv: 2110.07584.
    [26] J. Ren, J. Li, C. Liu, S. Chen, L. Liang, Y. Liu, Deep learning with physics-embedded neural network for full waveform ultrasonic brain imaging, IEEE Trans. Med. Imaging, 43 (2024), 2332–2346. https://doi.org/10.1109/TMI.2024.3363144 doi: 10.1109/TMI.2024.3363144
    [27] X. S. Peng, X. M. Zhang, Y. P. Li, B. Q. Liu, Research on image feature extraction and retrieval algorithms based on convolutional neural network, J. Vis. Commun. Image Represent., 69 (2020), 102705. https://doi.org/10.1016/j.jvcir.2019.102705 doi: 10.1016/j.jvcir.2019.102705
    [28] K. Han, Y. H. Wang, H. T. Chen, X. H. Chen, J. Y. Guo, Z. H. Liu, et al., A survey on vision transformer, IEEE Trans. Pattern Anal. Mach. Intell., 45 (2022), 87–110. https://doi.org/10.1109/TPAMI.2022.3152247 doi: 10.1109/TPAMI.2022.3152247
    [29] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al., Attention is all you need, in Advances in Neural Information Processing Systems, Curran Associates, Inc., (2017), 5998–6008.
    [30] S. K. Mondal, H. X. Zhang, H. D. Kabir, K. Ni, H. N. Dai, Machine translation and its evaluation: A study, Artif. Intell. Rev., 56 (2023), 10137–10226. https://doi.org/10.1007/s10462-023-10423-5 doi: 10.1007/s10462-023-10423-5
    [31] Y. Li, X. Wu, J. C. Wang, Y. M. Bo, F. Ni, C. H. Jiang, Differential attention vision transformer with adaptive spatial feature conditioning for remote sensing scene classification, Pattern Recogn., 178 (2026), 113461. https://doi.org/10.1016/j.patcog.2026.113461 doi: 10.1016/j.patcog.2026.113461
    [32] Q. Wang, S. T. Liu, J. Chanussot, X. L. Li, Scene classification with recurrent attention of VHR remote sensing images, IEEE Trans. Geosci. Remote Sens., 57 (2018), 1155–1167. https://doi.org/10.1109/TGRS.2018.2864987 doi: 10.1109/TGRS.2018.2864987
    [33] M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, et al., Generative pretraining from pixels, in Proceedings of the 37th International Conference on Machine Learning, PMLR, (2020), 1691–1703.
    [34] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, et al., An image is worth 16x16 words: Transformers for image recognition at scale, preprint, arXiv: 2010.11929.
    [35] R. Strudel, R. Garcia, I. Laptev, C. Schmid, Segmenter: Transformer for semantic segmentation, in 2021 IEEE/CVF International Conference on Computer Vision, IEEE, (2021), 7242–7252. https://doi.org/10.1109/ICCV48922.2021.00717
    [36] Z. Y. Wang, S. Wang, L. Y. Wu, D. Y. Liu, L. Gao, L. Qi, et al., Entroformer: An entropy-based sparse vision transformer for real-time semantic segmentation, Comput. Vis. Image Underst., 260 (2025), 104482. https://doi.org/10.1016/j.cviu.2025.104482 doi: 10.1016/j.cviu.2025.104482
    [37] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-end object detection with transformers, in European Conference on Computer Vision, Springer, (2020), 213–229. https://doi.org/10.1007/978-3-030-58452-8_13
    [38] C. Hao, Z. T. Yu, X. Liu, J. Xu, H. J. Yue, J. Y. Yang, A simple yet effective network based on vision transformer for camouflaged object and salient object detection, IEEE Trans. Image Process., 34 (2025), 608–622. https://doi.org/10.1109/TIP.2025.3528347 doi: 10.1109/TIP.2025.3528347
    [39] Y. J. Xiang, Z. L. Wang, Z. Song, R. Huang, G. J. Song, F. Min, Seismictransformer: An attention-based deep learning method for the simulation of seismic wavefields, Comput. Geosci., 190 (2024), 105629. https://doi.org/10.1016/j.cageo.2024.105629 doi: 10.1016/j.cageo.2024.105629
    [40] H. Z. Wang, J. Lin, Y. Li, X. T. Dong, X. Q. Tong, S. P. Lu, Self-supervised pretraining transformer for seismic data denoising, IEEE Trans. Geosci. Remote Sens., 62 (2024), 1–25. https://doi.org/10.1109/TGRS.2024.3368282 doi: 10.1109/TGRS.2024.3368282
    [41] C. Y. Ning, B. Y. Wu, B. H. Wu, Transformer and convolutional hybrid neural network for seismic impedance inversion, IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 17 (2024), 4436–4449. https://doi.org/10.1109/JSTARS.2024.3358610 doi: 10.1109/JSTARS.2024.3358610
    [42] B. Engquist, A. Majda, Absorbing boundary conditions for the numerical simulation of waves, Math. Comp., 31 (1977), 629–651. https://doi.org/10.1090/S0025-5718-1977-0436612-4 doi: 10.1090/S0025-5718-1977-0436612-4
    [43] A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, et al., Xcit: Cross-covariance image transformers, in Advances in Neural Information Processing Systems, Curran Associates, Inc., (2021), 20014–20027.
    [44] X. Zhang, X. Zhou, M. Lin, J. Sun, Shufflenet: An extremely efficient convolutional neural network for mobile devices, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, (2018), 6848–6856. https://doi.org/10.1109/CVPR.2018.00716
    [45] A. Dhara, M. K. Sen, Elastic full-waveform inversion using a physics-guided deep convolutional encoder-decoder, IEEE Trans. Geosci. Remote Sens., 61 (2023), 1–18. https://doi.org/10.1109/TGRS.2023.3294427 doi: 10.1109/TGRS.2023.3294427
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(587) PDF downloads(33) Cited by(0)

Article outline

Figures and Tables

Figures(6)  /  Tables(9)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog