Hoax News Detection on Urban Issues Using Naive Bayes with Chi-Square Feature Selection
DOI:
https://doi.org/10.31571/ijcmasted.v1i1.1038Keywords:
Hoax News, Naive Bayes, Chi-Square, Text Classification, Feature SelectionAbstract
The rapid spread of hoax news in digital media has become a serious concern because it can mislead the public, influence opinions, and potentially trigger social unrest, particularly in urban-related issues. The increasing volume of online information also makes manual verification difficult, highlighting the need for automatic hoax detection systems using machine learning approaches. This study aims to develop a hoax news detection model using the Multinomial Naive Bayes algorithm combined with Chi-Square feature selection to improve classification efficiency. The dataset consists of 1,000 Indonesian news articles related to urban issues, consisting of hoax and factual news collected through web scraping from fact-checking websites and online news portals. The research process includes text preprocessing stages such as case folding, tokenizing, stopword removal, and stemming, followed by TF-IDF feature extraction. Chi-Square feature selection was applied to reduce feature dimensionality by selecting the most relevant terms. The dataset was divided into training and testing sets using an 80:20 ratio. Model performance was evaluated using a confusion matrix with accuracy, precision, recall, and F1-score metrics. The experimental results show that the Multinomial Naive Bayes model without feature selection achieved the highest accuracy of 98.5%. Meanwhile, the application of Chi-Square feature selection with 200 selected features achieved an accuracy of 90.5%, while using 100 features resulted in an accuracy of 86.5%. These results indicate that although feature selection slightly reduces classification performance, it significantly reduces feature dimensionality while maintaining acceptable performance. Overall, this study demonstrates that Multinomial Naive Bayes is effective for hoax news classification, while Chi-Square feature selection provides an efficient alternative when dimensionality reduction is required.
References
Agma, A. R. (2025). Hoaks dan Disinformasi di Media Sosial: Strategi Literasi Digital untuk Meningkatkan Kesadaran Publik. Jurnal Komunikasi Dan Media, 1(1), 30–36. https://jurnal.samudrailmu.com/index.php/jkomed/article/view/28
Rahutomo, F., Yanuar, I., Pratiwi, R., & Ramadhani, D. M. (2019). Naïve Bayes’s Experiment On Hoax News Detection In Indonesian Language. Jurnal Penelitian Komunikasi Dan Opini Publik, 23(1), 1–15. https://doi.org/10.33299/jpkop.23.1.1805
Ratiasasadara, P. W., Sudarno, & Tarno. (2023). Analisis sentimen penerapan ppkm pada twitter menggunakan naïve bayes classifier dengan seleksi fitur chi-square 1,2,3. 11, 580–590. https://doi.org/10.14710/j.gauss.11.4.580-590
Reddy, H., Raj, N., Gala, M., & Basava, A. (2020). Text-mining-based Fake News Detection. https://doi.org/10.1007/s11633-019-1216-5
Sami’un, D. C., Sugiharto, A., & Jie, F. (2024). Chi Square Feature Selection For Improving Sentiment Analysis Of News Data Privacy Treats. Journal of Theoretical and Applied Information Technology, 102(18), 6601–6610. http://jatit.org/volumes/Vol102No18/3Vol102No18.pdf
Saraswati, N. W. S., Yanti, C. P., Muku, I. D. M. K., & Dewi, D. A. P. R. (2025). Evaluation Analysis of the Necessity of Stemming and Lemmatization in Text Classification. Jurnal Manajemen, Teknik Informatika, Dan Rekayasa Komputer, 24(2), 321–332. https://doi.org/10.30812/matrik.v24i2.4833
Sheviraa, S., Suarjaya, I. M. A. D., & Buana, P. W. (2022). Pengaruh Kombinasi dan Urutan Pre-Processing pada Tweets Bahasa Indonesia. Jurnal Ilmiah Teknologi Dan Komputer, 3(2). https://doi.org/10.24843/jtrti.2022.v03.i02.p06
Siino, M., Tinnirello, I., & Cascia, M. La. (2024). Is text preprocessing still worth the time ? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers. Information Systems, 121, 102342. https://doi.org/10.1016/j.is.2023.102342
Snegha, M. (2024). News Article Classification Using Navie Bayes Algorthim. International Journal of Progressive Research in Engineering Managemengt and Science (IJPREMS), 4(3), 425–427.
Downloads
Published
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Bimli, Wilda Susanti (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All articles published in the Proceeding of IJC-MaSTEd (International Joint Conference on Mathematics, Science, Technology, and Education) are published as open access and are licensed under the Creative Commons Attribution–ShareAlike 4.0 International License (CC BY-SA 4.0).
This license permits anyone to read, download, copy, distribute, adapt, and reuse the published work in any medium or format, including for commercial purposes, provided that proper credit is given to the original author(s) and source, a link to the license is included, and any modifications are clearly indicated. Any derivative works must be distributed under the same license (CC BY-SA 4.0).
Author Rights
Authors retain full copyright of their work. By submitting and publishing in the Proceeding of IJC-MaSTEd, authors grant the publisher a non-exclusive right to publish, distribute, and archive the article as part of the conference proceedings.
Authors are allowed to share, deposit, and reuse their published articles for academic, educational, and research purposes, provided proper attribution is given.
User Rights
Users may freely access and reuse all published content without prior permission, as long as appropriate attribution is provided and the same license terms are maintained.
Third-Party Material
Any third-party material included in an article must be clearly identified and is not covered by the Creative Commons license unless explicitly stated.
By submitting a manuscript to the Proceeding of IJC-MaSTEd, authors confirm that they agree with the open access policy and the license terms stated above.

