Comparative Approachin Evaluatingthe Impact of TF-IDF on Machine Learning Performance for Phishing Email Detection

Authors

  • Muhammad Azhar Hong Kong Shue Yan University image/svg+xml
  • Vinaye Armoogum
  • Sheeba Armoogum
  • Hamudi Alamsyah

Abstract

Phishing emails remain a major cybersecurity threat, resulting in significant financial and data losses. Therefore, accurate and efficient detection techniques are needed to identify phishing messages effectively. This study compares four traditional machine learning algorithms—Support Vector Machine (SVM), Random Forest, Logistic Regression, and XGBoost—for phishing email detection, with particular emphasis on the effect of Term Frequency-Inverse Document Frequency (TF-IDF) as a text feature extraction technique. Experiments were performed using a labeled dataset consisting of phishing and legitimate emails under two experimental conditions: with and without TF-IDF. Model performance was evaluated using accuracy, precision, recall, F1-score, and Area Under the Curve (AUC). The results show that SVM combined with TF-IDF achieves the best classification performance, indicating its potential for applications that require high detection accuracy. Random Forest and XGBoost achieve slightly lower performance but provide a favorable balance between predictive performance and computational efficiency, making them suitable for resource-constrained environments. Overall, the findings demonstrate that TF-IDF can substantially improve classification performance and emphasize the importance of effective feature engineering in traditional machine learning-based phishing detection systems.

References

Downloads

Published

2026-01-31

How to Cite

Comparative Approachin Evaluatingthe Impact of TF-IDF on Machine Learning Performance for Phishing Email Detection. (2026). Indonesian Journal of Cyber-AI and Security Intelligence, 1(1), 49-59. https://journal.idnns.org/index.php/ijcasi/article/view/38