Towards Transparent Fake News Detection Systems Using Large Language Models
Keywords:
Large language models, fake news detection, transparency, explainable AI, BERT, GPT, natural language processingAbstract
The rapid spread of fake news across digital platforms poses growing risks to public health, electoral integrity, and trust in institutions. Classical machine-learning and deep-learning detectors rely on handcrafted features or large labelled datasets and generalize poorly to unfamiliar topics and evolving manipulation tactics. Transformer-based large language models (LLMs) such as BERT, GPT, and RoBERTa markedly improve detection through contextual understanding and transfer learning, yet their opacity, heavy computational cost, susceptibility to bias, and vulnerability to adversarial attacks hinder trustworthy deployment. Although several surveys examine LLM-based detection, the theme of transparency remains underexplored. This paper reviews LLM-based fake news detection through the lens of transparency. Specifically, it (i) traces the evolution from classical models to transformer-based LLMs; (ii) proposes a transparency framework spanning interpretability, explanation generation, confidence calibration, bias auditing, accountability, and human oversight; (iii) annotates an end-to-end detection pipeline to indicate where each transparency mechanism can be applied; and (iv) synthesizes recent explainable and trustworthy approaches, mapping them to these dimensions. The review finds that LLMs deliver clear accuracy gains over earlier methods, but that transparency, fairness, and robustness are the decisive obstacles to real-world adoption. It concludes with a research agenda directed at auditable and socially responsible detection systems.
References
Abdullah, M., Madain, A., & Jararweh, Y. (2022). ChatGPT: Fundamentals, applications and social impacts. In Proceedings of the 9th International Conference on Social Networks Analysis, Management and Security (pp. 1–8). IEEE. https://doi.org/10.1109/SNAMS58071.2022.10062688
Allcott, H., & Gentzkow, M. (2017). Social media and fake news in the 2016 election. Journal of Economic Perspectives, 31(2), 211–236. https://doi.org/10.1257/jep.31.2.211
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In Proceedings of the 34th Conference on Neural Information Processing Systems (pp. 1877–1901). Curran Associates.
Castillo, C., Mendoza, M., & Poblete, B. (2011). Information credibility on Twitter. In Proceedings of the 20th International Conference on World Wide Web (pp. 675–684). ACM. https://doi.org/10.1145/1963405.1963500
Chen, Y., Li, D., Zhang, P., Sui, J., Lv, Q., Tun, L., & Shang, L. (2022). Cross-modal ambiguity learning for multimodal fake news detection. In Proceedings of the ACM Web Conference 2022 (pp. 2897–2905). ACM. https://doi.org/10.1145/3485447.3511968
Clark, K., Luong, M. T., Le, Q. V., & Manning, C. D. (2020). ELECTRA: Pre-training text encoders as discriminators rather than generators. In Proceedings of the International Conference on Learning Representations. ICLR. https://openreview.net/forum?id=r1xMH1BtvB
Danesh, A., & Rezanejad, A. (2025). Truth-aware explainable fake news detection system using large language models. Preprints.org. https://doi.org/10.20944/preprints202501.0001.v1
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Gunning, D., Stefik, M., Choi, J., Miller, T., Stumpf, S., & Yang, G. Z. (2019). XAI—Explainable artificial intelligence. Science Robotics, 4(37), Article eaay7120. https://doi.org/10.1126/scirobotics.aay7120
Hassan, S., Mahmood, A., & Ahmad, S. (2021). Multimodal deep learning for fake news detection. IEEE Transactions on Affective Computing, 12(4), 892–904. https://doi.org/10.1109/TAFFC.2021.3097755
Hu, B., Liu, X., & Wang, Y. (2023). Exploring the role of large language models in fake news detection. arXiv. https://doi.org/10.48550/arXiv.2309.12247
Jadhav, R., Rane, D., & Shaikh, A. (2024). Explainable multilingual and multimodal fake news detection using large language models. Frontiers in Artificial Intelligence, 7, Article 1345839. https://doi.org/10.3389/frai.2024.1345839
Jin, Z., Cao, J., Zhang, Y., Zhou, J., & Tian, Q. (2017). Novel visual and statistical image features for microblogs news verification. IEEE Transactions on Multimedia, 19(3), 598–608. https://doi.org/10.1109/TMM.2016.2617078
Kumar, A., & Gupta, D. (2022). Transformer-based models for fake news classification: A comprehensive survey. IEEE Access, 10, 74567–74589. https://doi.org/10.1109/ACCESS.2022.3191850
Kwon, S., Cha, M., Jung, K., Chen, W., & Wang, Y. (2013). Prominent features of rumor propagation in online social media. In Proceedings of the IEEE 13th International Conference on Data Mining (pp. 1103–1108). IEEE. https://doi.org/10.1109/ICDM.2013.61
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2020). ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the International Conference on Learning Representations. ICLR. https://openreview.net/forum?id=H1eA7AEtvS
LekshmiAmmal, H. R., & Madasamy, A. (2025). Explainable large language models for misinformation detection in online social networks. Journal of Big Data, 12(1), Article 15. https://doi.org/10.1186/s40537-025-01034-6
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7871–7880). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.703
Li, S., Zhang, W., & Chen, H. (2023). Evaluating large language model performance on misinformation detection: Challenges and opportunities. ACM Computing Surveys, 55(9), 1–38. https://doi.org/10.1145/3571730
Liu, Y., & Wu, Y. F. B. (2018). Early detection of fake news on social media through propagation path classification with recurrent and convolutional networks. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence (pp. 354–361). AAAI Press.
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://doi.org/10.48550/arXiv.1907.11692
Liu, H., Wang, J., & Li, Y. (2021). Harnessing transformer models for misinformation detection. Journal of Information Processing, 29, 103–115. https://doi.org/10.2197/ipsjjip.29.103
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2021). GPT understands, too. arXiv. https://doi.org/10.48550/arXiv.2103.10385
Liu, H., Chen, R., & Wang, S. (2024). TELLER: A trustworthy framework for explainable large language model-based fake news detection system. arXiv. https://doi.org/10.48550/arXiv.2402.07776
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607
Papageorgiou, E., Koutkias, V., & Maglogiannis, I. (2024). A survey on large language models for fake news detection: Applications, challenges, and future perspectives. Future Internet, 16(8), Article 287. https://doi.org/10.3390/fi16080287
Pérez-Rosas, V., Kleinberg, B., Lefevre, A., & Mihalcea, R. (2018). Automatic detection of fake news. In Proceedings of the 27th International Conference on Computational Linguistics (pp. 3391–3401). Association for Computational Linguistics.
Qian, Y., Zhang, X., & Chen, L. (2021). Explainable fake news detection via deep neural networks. IEEE Transactions on Knowledge and Data Engineering, 33(6), 2325–2338. https://doi.org/10.1109/TKDE.2021.3073275
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI Technical Report. https://openai.com/research/language-unsupervised
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Technical Report. https://openai.com/research/language-unsupervised
Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv. https://doi.org/10.48550/arXiv.1910.01108
Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1), 22–36. https://doi.org/10.1145/3137597.3137600
Shu, K., Wang, S., & Liu, H. (2019). Beyond news contents: The role of social context for fake news detection. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining (pp. 312–320). ACM. https://doi.org/10.1145/3289600.3290994
Tachet des Combes, R., Dhillon, P., & Awerbuch, B. (2017). Domain adversarial training for classification of misleading news articles. In Proceedings of the International Workshop on News Recommendation and Analytics (pp. 1–6).
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (pp. 6000–6010). Curran Associates.
Vig, J. (2019). A multiscale visualization of attention in the transformer model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 37–42). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-3007
Vosoughi, S., Roy, D., & Aral, S. (2018). The spread of true and false news online. Science, 359(6380), 1146–1151. https://doi.org/10.1126/science.aap9559
Wang, W. Y. (2017). 'Liar, liar pants on fire': A new benchmark dataset for fake news detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 422–426). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-2067
Wang, Y., Ma, F., Jin, Z., Yuan, Y., Xun, G., Jha, K., Su, L., & Gao, J. (2018). EANN: Event adversarial neural networks for multi-modal fake news detection. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 849–857). ACM. https://doi.org/10.1145/3219819.3219903
Wang, Y., Zhang, L., & Chen, J. (2024). Large language model-based fake news detection: Adversarial robustness and defense mechanisms. IEEE Transactions on Neural Networks and Learning Systems, 35(3), 2156–2169. https://doi.org/10.1109/TNNLS.2024.3357891
Wiegreffe, S., & Pinter, Y. (2019). Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 11–20). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1002
Xiao, Y., Zhang, M., & Liu, W. (2024). Adversarial threats to large language model-based fake news detection systems: A comprehensive analysis. IEEE Transactions on Dependable and Secure Computing. Advance online publication. https://doi.org/10.1109/TDSC.2024.3382957
Yang, X., Chen, X., & Wang, X. (2023). Cross-domain few-shot fake news detection via graph neural networks. IEEE Transactions on Computational Social Systems, 10(4), 1756–1767. https://doi.org/10.1109/TCSS.2022.3218421
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending against neural fake news. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (pp. 9054–9065). Curran Associates. https://doi.org/10.5555/3454287.3455088
Zhang, H., Wang, D., & Xu, F. (2022). Cross-domain fake news detection with transformers. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (pp. 4892–4898). IJCAI. https://doi.org/10.24963/ijcai.2022/679
Zhao, Z., Zhao, J., Sano, Y., Levy, O., Takayasu, H., Takayasu, M., Li, D., Wu, J., & Havlin, S. (2020). Fake news propagates differently from real news even at early stages of spreading. EPJ Data Science, 9(1), Article 7. https://doi.org/10.1140/epjds/s13688-020-00224-z
Zhu, H., Wang, D., & Feng, Y. (2018). Combating fake news: An investigation of information verification behaviors on social networking sites. In Proceedings of the 51st Hawaii International Conference on System Sciences (pp. 3977–3986). University of Hawaii at Manoa. https://doi.org/10.24251/HICSS.2018.501
Zhu, H., Han, C., & Li, Y. (2020). Combating fake news with deep contextualized representations using BERT. IEEE Transactions on Computational Social Systems, 7(3), 667–678. https://doi.org/10.1109/TCSS.2020.2986035
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Usha Rani, Kavita Mittal

This work is licensed under a Creative Commons Attribution 4.0 International License.
