An Explainable Lightweight Phishing Website Detection Framework Using Consensus Feature Selection and Performance-Weighted Ensemble Learning
Singh A1*
DOI:10.31033/ABJAR/5.3.2026.120
1* Anjali Singh, Department of Computer Science and Engineering, Netaji Subhas University of Technology, Dwarka, New Delhi, India.
Phishing websites continue to pose a significant cybersecurity threat by deceiving users into disclosing sensitive information through fraudulent web pages that closely imitate legitimate services. Although machine learning has significantly improved phishing detection, many existing approaches rely on high-dimensional feature sets, exhibit limited interpretability, or incur substantial computational overhead. To address these challenges, this paper proposes an explainable lightweight phishing website detection framework that combines consensus feature selection with performance-weighted ensemble learning. The proposed framework first employs Recursive Feature Elimination (RFE), Mutual Information (MI), and SHAP (SHapley Additive exPlanations) to identify the most informative URL-based features through a consensus voting strategy. This process reduces the original 22 URL features to 10 highly discriminative features while preserving classification performance. Subsequently, multiple ensemble classifiers are trained using the selected features, and a Performance-Weighted Consensus Ensemble (PWCE) is constructed by combining classifier probabilities according to their validation performance. Experiments conducted on the publicly available PhiUSIIL Phishing URL dataset demonstrate that the proposed framework achieves an accuracy of 99.987%, an F1-score of 99.989%, and a ROC-AUC of 99.989% while substantially reducing feature dimensionality. Comprehensive evaluations, including five-fold cross-validation, SHAP-based explainability, runtime analysis, calibration assessment, and feature ablation studies, confirm the robustness, efficiency, and interpretability of the proposed framework. The experimental results indicate that accurate phishing detection can be achieved using a compact set of explainable URL features, making the proposed approach suitable for real-time cybersecurity applications.
Keywords: phishing detection, explainable artificial intelligence, SHAP, feature selection, ensemble learning, cybersecurity
| Corresponding Author | How to Cite this Article | To Browse |
|---|---|---|
| , Department of Computer Science and Engineering, Netaji Subhas University of Technology, Dwarka, New Delhi, India. Email: |
Singh A, An Explainable Lightweight Phishing Website Detection Framework Using Consensus Feature Selection and Performance-Weighted Ensemble Learning. Appl Sci Biotechnol J Adv Res. 2026;5(3):42-57. Available From https://abjar.vandanapublications.com/index.php/ojs/article/view/120 |


©