Beyond Single-Modality Detection: A Systematic Review of Multimodal AI for Social Media Cybercrime

Authors

  • Ekta R. Kaneriya Ph.D. Research Scholar, Department of Computer Science, Atmiya University, Rajkot, Gujarat
  • Disha M. Ganatra Assistant Professor, Department of Computer Science, Atmiya University, Rajkot, Gujarat

DOI:

https://doi.org/10.69968/ijisem.2026v5i3505-509

Keywords:

Social Media Cybercrime, Multimodal AI, Cross-Platform Detec-tion

Abstract

Social media platforms have emerged as a key battleground for cy-bercriminals, which leverages the speed, scale and interconnectedness of social media to orchestrate financial crimes such as phishing and OTP Fraud, identity-based crime such as deep-fake and account takeover, and social harms such as cyberbullying and coordinated influence campaigns. The threats are increasingly multifaceted and cross-platform, making it difficult for traditional detection solutions to stay ahead of the curve, as they are typically limited to a specific data type or a single platform. We review recent publications on detecting cybercrime in social networks ranging from misinformation detection, to machine learning and deep learning based classifiers, fraud and intrusion detection pipelines, to novel multimodal architectures for AI. The reviewed literature can be divided into three the-matic groups: (1) detection methods based on content and propagation patterns, (2) methods based on user behaviour and network structure, and (3) multimodal fusion models coupled with explainable AI mechanisms. We evaluate the data sets used, algorithm selection, reported performance, and inherent limitations of the underlying studies for each cluster. The analysis shows that there is a gap. Although there has been significant progress in the detection of individual tasks, existing systems are still highly specialised and platform-dependent, or type-dependent, or modal-ity dependent. Very few integrate text, image and behaviour into a single model. It also identifies a lack of attention to interpretability and cross-platform real time operation. To address these gaps, we suggest the necessity of developing a multimodal detection system that is uniform, adaptive and explainable, which can work in heterogeneous environments on social media platforms, and we propose some research directions for achieving this goal.

References

[1] M. Nasser et al., “A systematic review of multimodal fake news detection on social media using deep learning models,” Results in Engineering, vol. 26, p. 104752, 2025, doi: https://doi.org/10.1016/j.rineng.2025.104752.

[2] X. Shen, M. Huang, Z. Hu, S. Cai, and T. Zhou, “Multimodal Fake News De-tection with Contrastive Learning and Optimal Transport,” Front. Comput. Sci., vol. Volume 6-2024, 2024, doi: 10.3389/fcomp.2024.1473457.

[3] L. Shen et al., “GAMED: Knowledge Adaptive Multi-Experts Decoupling for Multimodal Fake News Detection,” 2024.

[4] Mukherjee and S. Ghosh, “UNITE-FND: Reframing Multimodal Fake News Detection through Unimodal Scene Translation,” 2025.

[5] J. Lv, Y. Gao, L. Li, L. Shi, and S. Li, “Multi-modal fake news detection: A comprehensive survey on deep learning technology, advances, and challenges,” Journal of King Saud University Computer and Information Sciences, vol. 37, Jul. 2025, doi: 10.1007/s44443-025-00317-7.

[6] T. Wu, Z. Ma, Y. Cui, Z. Zhou, and E. Wang, “MSM-BD: Multimodal Social Media Bot Detection Using Heterogeneous Information,” 2025.

[7] T. Huang, Y. Wang, Q. Li, C. He, and J. Gao, “Can LLMs Find Fraudsters? Multi-level LLM Enhanced Graph Fraud Detection,” in Proceedings of the 33rd ACM International Conference on Multimedia (MM ’25), New York, NY, USA: ACM, 2025, pp. 1530–1538. doi: 10.1145/3746027.3755245.

[8] S. M. Qureshi, A. Saeed, S. H. Almotiri, F. Ahmad, and M. A. Al Ghamdi, “Deepfake forensics: a survey of digital forensic methods for multimodal deep-fake identification on social media,” PeerJ Comput. Sci., vol. 10, p. e2037, 2024, doi: 10.7717/peerj-cs.2037.

[9] N. A. Chandra et al., “Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024,” 2025.

[10] U. Sahin, I. E. Kucukkaya, O. Ozcelik, and C. Toraman, “ARC-NLP at Multi-modal Hate Speech Event Detection 2023: Multimodal Methods Boosted by Ensemble Learning, Syntactical and Entity Features,” 2023.

[11] Chhabra and D. K. Vishwakarma, “MHS-STMA: Multimodal Hate Speech Detection via Scalable Transformer-Based Multilevel Attention Framework,” 2024.

[12] D. Premkumar and S. K. Nachimuthu, “Privacy-preserving cyberthreat detec-tion in decentralized social media with federated cross-modal graph transform-ers,” Sci. Rep., vol. 16, no. 1, p. 3608, 2026, doi: 10.1038/s41598-025-33596-1.

[13] G. Joshi et al., “Explainable Misinformation Detection Across Multiple Social Media Platforms,” 2022.

[14] R. Jadhav, V. Meshram, A. Bhosle, K. Patil, S. Dash, and S. Jadhav, “Explain-able multilingual and multimodal fake-news detection: toward robust and trust-worthy AI for combating misinformation,” Front. Artif. Intell., vol. Volume 8-2025, 2025, doi: 10.3389/frai.2025.1690616.

Downloads

Published

01-09-2026

Issue

Section

Articles

How to Cite

[1]
Ekta R. Kaneriya and Disha M. Ganatra 2026. Beyond Single-Modality Detection: A Systematic Review of Multimodal AI for Social Media Cybercrime. International Journal of Innovations in Science, Engineering And Management. 5, 3 (Sep. 2026), 505–509. DOI:https://doi.org/10.69968/ijisem.2026v5i3505-509.