With the rapid advancement of digital technologies and social media, news dissemination has become increasingly complex, often spanning multiple domains and encompassing diverse modalities (e.g., images and text). At the same time, the generation and spread of fake news have grown more covert and sophisticated, posing significant threats to social stability and public trust. In this context, traditional detection approaches that rely on a single modality or target a specific domain are no longer sufficient to handle the diversity and dynamic nature of real-world misinformation. To address this challenge, we propose an efficient multimodal, cross-domain fake news detection framework. Our framework incorporates a lightweight linear attention mechanism inspired by the Mamba model, along with a novel multimodal attention mechanism for efficient attention design and robust multimodal fusion. In addition, it integrates a masked linear attention module and a single-modality feature distiller to further capture modality-specific characteristics. Extensive experiments on two real-world datasets, together with comprehensive ablation studies, demonstrate that our model consistently outperforms existing methods and validates the necessity of each core component. Our source code is publicly available at: https://github.com/biubiu-cpp/zqf_Fake-News-Detection.
Citation: Qingfan Zhang, Tenghui Wang. A multistage multimodal fusion network for cross-domain fake news detection[J]. Big Data and Information Analytics, 2026, 10: 158-176. doi: 10.3934/bdia.2026008
With the rapid advancement of digital technologies and social media, news dissemination has become increasingly complex, often spanning multiple domains and encompassing diverse modalities (e.g., images and text). At the same time, the generation and spread of fake news have grown more covert and sophisticated, posing significant threats to social stability and public trust. In this context, traditional detection approaches that rely on a single modality or target a specific domain are no longer sufficient to handle the diversity and dynamic nature of real-world misinformation. To address this challenge, we propose an efficient multimodal, cross-domain fake news detection framework. Our framework incorporates a lightweight linear attention mechanism inspired by the Mamba model, along with a novel multimodal attention mechanism for efficient attention design and robust multimodal fusion. In addition, it integrates a masked linear attention module and a single-modality feature distiller to further capture modality-specific characteristics. Extensive experiments on two real-world datasets, together with comprehensive ablation studies, demonstrate that our model consistently outperforms existing methods and validates the necessity of each core component. Our source code is publicly available at: https://github.com/biubiu-cpp/zqf_Fake-News-Detection.
| [1] |
Jing J, Wu H, Sun J, Fang X, Zhang H, (2023) Multimodal fake news detection via progressive fusion networks. Inf Process Manag 60: 103120. https://doi.org/10.1016/j.ipm.2022.103120 doi: 10.1016/j.ipm.2022.103120
|
| [2] | Liu X, Nourbakhsh A, Li Q, Fang R, Shah S, (2015) Real-time rumor debunking on twitter. In: Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, 1867–1870. https://doi.org/10.1145/2806416.2806651 |
| [3] |
Yu F, Liu Q, Wu S, Wang L, Tan T, (2017) A convolutional approach for misinformation identification. In: IJCAI, 3901–3907. https://doi.org/10.24963/ijcai.2017/545 doi: 10.24963/ijcai.2017/545
|
| [4] | Nan Q, Cao J, Zhu Y, Wang Y, Li J, (2021) MDFEND: Multi-domain fake news detection. In: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 3343–3347. https://doi.org/10.1145/3459637.3482139 |
| [5] |
Zhang W, Deng L, Zhang L, Wu D, (2022) A survey on negative transfer. IEEE/CAA J Autom Sin 10: 305–329. https://doi.org/10.1109/JAS.2022.106004 doi: 10.1109/JAS.2022.106004
|
| [6] |
Han D, Wang Z, Xia Z, Han Y, Pu Y, Ge C, et al. (2024) Demystify mamba in vision: A linear attention perspective. Adv Neural Inf Process Syst 37: 127181–127203. https://doi.org/10.52202/079017-4039 doi: 10.52202/079017-4039
|
| [7] |
Zhu Y, Sheng Q, Cao J, Nan Q, Shu K, Wu M, et al. (2022) Memory-guided multi-view multi-domain fake news detection. IEEE Trans Know Data Eng 35: 7178–7191. https://doi.org/10.1109/TKDE.2022.3185151 doi: 10.1109/TKDE.2022.3185151
|
| [8] |
Wu D, Tan Z, Zhao H, Jiang T, Geng N, (2024) Domain-and category-style clustering for general fake news detection via contrastive learning. Inf Process Manag 61: 103725. https://doi.org/10.1016/j.ipm.2024.103725 doi: 10.1016/j.ipm.2024.103725
|
| [9] |
Lu W, Tong Y, Ye Z, (2025) Dammfnd: Domain-aware multimodal multi-view fake news detection. In: Proceedings of the AAAI Conference on Artificial Intelligence 39: 559–567. https://doi.org/10.1609/aaai.v39i1.32036 doi: 10.1609/aaai.v39i1.32036
|
| [10] |
Yang X, Wang Y, Zhang X, Wang S, Wang H, Lam K, (2025) A macro-and micro-hierarchical transfer learning framework for cross-domain fake news detection. In: Proceedings of the ACM on Web Conference 2025, 5297–5307. https://doi.org/10.1145/3696410.3714517 doi: 10.1145/3696410.3714517
|
| [11] | Wang Y, Ma F, Jin Z, Yuan Y, Xun G, Jha K, et al. (2018) Eann: Event adversarial neural networks for multi-modal fake news detection. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 849–857. https://doi.org/10.1145/3219819.3219903 |
| [12] | Wei Z, Pan H, Qiao L, Niu X, Dong P, Li D, (2022) Cross-modal knowledge distillation in multi-modal fake news detection. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing, 4733–4737. https://doi.org/10.1109/ICASSP43922.2022.9747280 |
| [13] |
Wu L, Liu P, Zhang Y, (2023) See how you read? multi-reading habits fusion reasoning for multi-modal fake news detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, 37: 1373613744. https://doi.org/10.1609/aaai.v37i11.26609 doi: 10.1609/aaai.v37i11.26609
|
| [14] |
Ying Q, Hu X, Zhou Y, Qian Z, Zeng D, Ge S, (2023) Bootstrapping multi-view representations for fake news detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, 37: 5384–5392. https://doi.org/10.1609/aaai.v37i4.25670 doi: 10.1609/aaai.v37i4.25670
|
| [15] | Nasir S, Wasim M, Rehman A, Ghani MU, (2025) FACT-CLIP: Fake news detection via CLIP-based cross-modal attention and transformer fusion. In: 2025 International Conference on Emerging Technologies in Electronics, Computing, and Communication (ICETECC), 1–6. https://doi.org/10.1109/ICETECC65365.2025.11070224 |
| [16] | Qi P, Yan Z, Hsu W, Lee ML, (2024) Sniffer: Multimodal large language model for explainable out-of-context misinformation detection. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13052–13062. https://doi.org/10.1109/CVPR52733.2024.01240 |
| [17] |
Song C, Ning N, Zhang Y, Wu B, (2021) Knowledge augmented transformer for adversarial multidomain multiclassification multimodal fake news detection. Neurocomputing 462: 88–100. https://doi.org/10.1016/j.neucom.2021.07.077 doi: 10.1016/j.neucom.2021.07.077
|
| [18] | Tong Y, Lu W, Zhao Z, Lai S, Shi T, (2024) Mmdfnd: Multi-modal multi-domain fake news detection. In: Proceedings of the 32nd ACM International Conference on Multimedia, 1178–1186. https://doi.org/10.1145/3664647.3681317 |
| [19] | Koroteev MV, (2021) BERT: A review of applications in natural language processing and understanding, preprint, arXiv: 2103.11943. https://doi.org/10.48550/arXiv.2103.11943 |
| [20] |
He K, Chen X, Xie S, Li Y, Dollár P, Girshick R, (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16000–16009. https://doi.org/10.1109/CVPR52688.2022.01553 doi: 10.1109/CVPR52688.2022.01553
|
| [21] | Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, 8748–8763. |
| [22] | Gu A, Dao T, (2024) Mamba: Linear-time sequence modeling with selective state spaces, In: First Conference on Language Modeling. |
| [23] |
Hu J, Shen L, Sun G, (2018) Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7132–7141. https://doi.org/10.1109/CVPR.2018.00745 doi: 10.1109/CVPR.2018.00745
|
| [24] | Ma J, Zhao Z, Yi X, Chen J, Hong L, Chi EH, (2018) Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1930–1939. https://doi.org/10.1145/3219819.3220007 |
| [25] | Qin Z, Cheng Y, Zhao Z, Chen Z, Metzler D, Qin J, (2020) Multitask mixture of sequential experts for user activity streams. In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 3083–3091. https://doi.org/10.1145/3394486.3403359 |
| [26] | Singhal S, Shah RR, Chakraborty T, Kumaraguru P, Satoh SI, (2019) Spotfake: A multi-modal framework for fake news detection. In: 2019 IEEE Fifth International Conference on Multimedia Big Data (BigMM), 39–47. https://doi.org/10.1109/BigMM.2019.00-44 |
| [27] |
Chen Y, Li D, Zhang P, Sui J, Lv Q, Tun L, Shang L, (2022) Cross-modal ambiguity learning for multimodal fake news detection. In: Proceedings of the ACM Web Conference 2022, 2897–2905. https://doi.org/10.1145/3485447.3511968 doi: 10.1145/3485447.3511968
|