DATASET-DEPENDENT PERFORMANCE OF FILTER-BASED FEATURE SELECTION METHODS: A COMPARATIVE STUDY IN HIGH-DIMENSIONAL CLASSIFICATION
Main Article Content
Abstract
This study presents a comprehensive evaluation of filter-based feature selection methods in high dimensional classification mainly focusing on the effectiveness of these methods in multiple datasets and classifiers. In this research three filter based techniques including Chi-square, Mutual Information and ANOVA are evaluated across three benchmark datasets with different characteristics such as synthetic noisy data, real world data and structured multiclass data.The analysis looks into the interactions between feature selection methods, classifiers and feature subset sizes using multiple performance metrics such as accuracy, precision, recall and F1-score. The results reveal that the highest performing experimental configuration was achieved in ISOLET dataset using Mutual Information with SVM classifier at top 40% feature subset with accuracy of 95.4%, precision of 95.6%, F1Score and recall of 95.4%. For Bioresponse dataset best performance was achieved using mutual information with random forest achieving accuracy of 81.4%, Madelon datasr had Chi-square performed best with random forest and accuracy of 81.3%. Findings demonstrte that not one single method performed best for all datasets and depends on dataset characteristics and classifier selection.