Towards Transparent Fake News Detection Systems Using Large Language Models

Authors

Keywords:

Large language models, fake news detection, transparency, explainable AI, BERT, GPT, natural language processing

Abstract

The rapid spread of fake news across digital platforms poses growing risks to public health, electoral integrity, and trust in institutions. Classical machine-learning and deep-learning detectors rely on handcrafted features or large labelled datasets and generalize poorly to unfamiliar topics and evolving manipulation tactics. Transformer-based large language models (LLMs) such as BERT, GPT, and RoBERTa markedly improve detection through contextual understanding and transfer learning, yet their opacity, heavy computational cost, susceptibility to bias, and vulnerability to adversarial attacks hinder trustworthy deployment. Although several surveys examine LLM-based detection, the theme of transparency remains underexplored. This paper reviews LLM-based fake news detection through the lens of transparency. Specifically, it (i) traces the evolution from classical models to transformer-based LLMs; (ii) proposes a transparency framework spanning interpretability, explanation generation, confidence calibration, bias auditing, accountability, and human oversight; (iii) annotates an end-to-end detection pipeline to indicate where each transparency mechanism can be applied; and (iv) synthesizes recent explainable and trustworthy approaches, mapping them to these dimensions. The review finds that LLMs deliver clear accuracy gains over earlier methods, but that transparency, fairness, and robustness are the decisive obstacles to real-world adoption. It concludes with a research agenda directed at auditable and socially responsible detection systems.

References

Abdullah, M., Madain, A., & Jararweh, Y. (2022). ChatGPT: Fundamentals, applications and social impacts. In Proceedings of the 9th International Conference on Social Networks Analysis, Management and Security (pp. 1–8). IEEE. https://doi.org/10.1109/SNAMS58071.2022.10062688

Allcott, H., & Gentzkow, M. (2017). Social media and fake news in the 2016 election. Journal of Economic Perspectives, 31(2), 211–236. https://doi.org/10.1257/jep.31.2.211

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In Proceedings of the 34th Conference on Neural Information Processing Systems (pp. 1877–1901). Curran Associates.

Castillo, C., Mendoza, M., & Poblete, B. (2011). Information credibility on Twitter. In Proceedings of the 20th International Conference on World Wide Web (pp. 675–684). ACM. https://doi.org/10.1145/1963405.1963500

Chen, Y., Li, D., Zhang, P., Sui, J., Lv, Q., Tun, L., & Shang, L. (2022). Cross-modal ambiguity learning for multimodal fake news detection. In Proceedings of the ACM Web Conference 2022 (pp. 2897–2905). ACM. https://doi.org/10.1145/3485447.3511968

Clark, K., Luong, M. T., Le, Q. V., & Manning, C. D. (2020). ELECTRA: Pre-training text encoders as discriminators rather than generators. In Proceedings of the International Conference on Learning Representations. ICLR. https://openreview.net/forum?id=r1xMH1BtvB

Danesh, A., & Rezanejad, A. (2025). Truth-aware explainable fake news detection system using large language models. Preprints.org. https://doi.org/10.20944/preprints202501.0001.v1

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Gunning, D., Stefik, M., Choi, J., Miller, T., Stumpf, S., & Yang, G. Z. (2019). XAI—Explainable artificial intelligence. Science Robotics, 4(37), Article eaay7120. https://doi.org/10.1126/scirobotics.aay7120

Hassan, S., Mahmood, A., & Ahmad, S. (2021). Multimodal deep learning for fake news detection. IEEE Transactions on Affective Computing, 12(4), 892–904. https://doi.org/10.1109/TAFFC.2021.3097755

Hu, B., Liu, X., & Wang, Y. (2023). Exploring the role of large language models in fake news detection. arXiv. https://doi.org/10.48550/arXiv.2309.12247

Jadhav, R., Rane, D., & Shaikh, A. (2024). Explainable multilingual and multimodal fake news detection using large language models. Frontiers in Artificial Intelligence, 7, Article 1345839. https://doi.org/10.3389/frai.2024.1345839

Jin, Z., Cao, J., Zhang, Y., Zhou, J., & Tian, Q. (2017). Novel visual and statistical image features for microblogs news verification. IEEE Transactions on Multimedia, 19(3), 598–608. https://doi.org/10.1109/TMM.2016.2617078

Kumar, A., & Gupta, D. (2022). Transformer-based models for fake news classification: A comprehensive survey. IEEE Access, 10, 74567–74589. https://doi.org/10.1109/ACCESS.2022.3191850

Kwon, S., Cha, M., Jung, K., Chen, W., & Wang, Y. (2013). Prominent features of rumor propagation in online social media. In Proceedings of the IEEE 13th International Conference on Data Mining (pp. 1103–1108). IEEE. https://doi.org/10.1109/ICDM.2013.61

Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2020). ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the International Conference on Learning Representations. ICLR. https://openreview.net/forum?id=H1eA7AEtvS

LekshmiAmmal, H. R., & Madasamy, A. (2025). Explainable large language models for misinformation detection in online social networks. Journal of Big Data, 12(1), Article 15. https://doi.org/10.1186/s40537-025-01034-6

Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7871–7880). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.703

Li, S., Zhang, W., & Chen, H. (2023). Evaluating large language model performance on misinformation detection: Challenges and opportunities. ACM Computing Surveys, 55(9), 1–38. https://doi.org/10.1145/3571730

Liu, Y., & Wu, Y. F. B. (2018). Early detection of fake news on social media through propagation path classification with recurrent and convolutional networks. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence (pp. 354–361). AAAI Press.

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://doi.org/10.48550/arXiv.1907.11692

Liu, H., Wang, J., & Li, Y. (2021). Harnessing transformer models for misinformation detection. Journal of Information Processing, 29, 103–115. https://doi.org/10.2197/ipsjjip.29.103

Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2021). GPT understands, too. arXiv. https://doi.org/10.48550/arXiv.2103.10385

Liu, H., Chen, R., & Wang, S. (2024). TELLER: A trustworthy framework for explainable large language model-based fake news detection system. arXiv. https://doi.org/10.48550/arXiv.2402.07776

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607

Papageorgiou, E., Koutkias, V., & Maglogiannis, I. (2024). A survey on large language models for fake news detection: Applications, challenges, and future perspectives. Future Internet, 16(8), Article 287. https://doi.org/10.3390/fi16080287

Pérez-Rosas, V., Kleinberg, B., Lefevre, A., & Mihalcea, R. (2018). Automatic detection of fake news. In Proceedings of the 27th International Conference on Computational Linguistics (pp. 3391–3401). Association for Computational Linguistics.

Qian, Y., Zhang, X., & Chen, L. (2021). Explainable fake news detection via deep neural networks. IEEE Transactions on Knowledge and Data Engineering, 33(6), 2325–2338. https://doi.org/10.1109/TKDE.2021.3073275

Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI Technical Report. https://openai.com/research/language-unsupervised

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Technical Report. https://openai.com/research/language-unsupervised

Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv. https://doi.org/10.48550/arXiv.1910.01108

Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1), 22–36. https://doi.org/10.1145/3137597.3137600

Shu, K., Wang, S., & Liu, H. (2019). Beyond news contents: The role of social context for fake news detection. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining (pp. 312–320). ACM. https://doi.org/10.1145/3289600.3290994

Tachet des Combes, R., Dhillon, P., & Awerbuch, B. (2017). Domain adversarial training for classification of misleading news articles. In Proceedings of the International Workshop on News Recommendation and Analytics (pp. 1–6).

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (pp. 6000–6010). Curran Associates.

Vig, J. (2019). A multiscale visualization of attention in the transformer model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 37–42). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-3007

Vosoughi, S., Roy, D., & Aral, S. (2018). The spread of true and false news online. Science, 359(6380), 1146–1151. https://doi.org/10.1126/science.aap9559

Wang, W. Y. (2017). 'Liar, liar pants on fire': A new benchmark dataset for fake news detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 422–426). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-2067

Wang, Y., Ma, F., Jin, Z., Yuan, Y., Xun, G., Jha, K., Su, L., & Gao, J. (2018). EANN: Event adversarial neural networks for multi-modal fake news detection. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 849–857). ACM. https://doi.org/10.1145/3219819.3219903

Wang, Y., Zhang, L., & Chen, J. (2024). Large language model-based fake news detection: Adversarial robustness and defense mechanisms. IEEE Transactions on Neural Networks and Learning Systems, 35(3), 2156–2169. https://doi.org/10.1109/TNNLS.2024.3357891

Wiegreffe, S., & Pinter, Y. (2019). Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 11–20). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1002

Xiao, Y., Zhang, M., & Liu, W. (2024). Adversarial threats to large language model-based fake news detection systems: A comprehensive analysis. IEEE Transactions on Dependable and Secure Computing. Advance online publication. https://doi.org/10.1109/TDSC.2024.3382957

Yang, X., Chen, X., & Wang, X. (2023). Cross-domain few-shot fake news detection via graph neural networks. IEEE Transactions on Computational Social Systems, 10(4), 1756–1767. https://doi.org/10.1109/TCSS.2022.3218421

Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending against neural fake news. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (pp. 9054–9065). Curran Associates. https://doi.org/10.5555/3454287.3455088

Zhang, H., Wang, D., & Xu, F. (2022). Cross-domain fake news detection with transformers. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (pp. 4892–4898). IJCAI. https://doi.org/10.24963/ijcai.2022/679

Zhao, Z., Zhao, J., Sano, Y., Levy, O., Takayasu, H., Takayasu, M., Li, D., Wu, J., & Havlin, S. (2020). Fake news propagates differently from real news even at early stages of spreading. EPJ Data Science, 9(1), Article 7. https://doi.org/10.1140/epjds/s13688-020-00224-z

Zhu, H., Wang, D., & Feng, Y. (2018). Combating fake news: An investigation of information verification behaviors on social networking sites. In Proceedings of the 51st Hawaii International Conference on System Sciences (pp. 3977–3986). University of Hawaii at Manoa. https://doi.org/10.24251/HICSS.2018.501

Zhu, H., Han, C., & Li, Y. (2020). Combating fake news with deep contextualized representations using BERT. IEEE Transactions on Computational Social Systems, 7(3), 667–678. https://doi.org/10.1109/TCSS.2020.2986035

Downloads

Published

2026-07-31

How to Cite

[1]
Rani, U. and Mittal, K. 2026. Towards Transparent Fake News Detection Systems Using Large Language Models. International Journal of Convergent Research. 3, 1 (Jul. 2026), 1–11.