Online feature subset selection for mining feature streams in big data via incremental learning and evolutionary computation

Publications

Online feature subset selection for mining feature streams in big data via incremental learning and evolutionary computation

Author : Dr Vivek Yelleti

Year : 2025

Publisher : Elsevier B.V.

Source Title : Swarm and Evolutionary Computation

Document Type :

Abstract

Online streaming feature subset selection (OSFSS) presents a noteworthy challenge when data samples arrive rapidly and in a time-dependent manner. The complexity of this problem is further exacerbated when features arrive as a stream. Despite several attempts to solve OSFSS over feature streams, existing methods lack scalability, cannot handle interaction effects among features, and fail to efficiently handle high-velocity feature streams. To address these challenges, we propose a novel wrapper-for OSFSS named OSFSS-W (wrapper-for OSFSS), specifically designed to mine feature streams within the Apache Spark environment. Our proposed method incorporates (i) two vigilance tests: for removing (a) irrelevant features and (b) redundant features (ii) incremental learning and (iii) a tolerance-based feedback mechanism that retains and utilizes previous knowledge while adhering to the predefined tolerance thresholds. Additionally, for the purpose of optimization, we introduce a Bare Bones Particle Swarm Optimization (BBPSO-L) algorithm driven by the logistic distribution. Further, the BBPSO-L is parallelized within Apache Spark, following an island-based approach. We evaluated the performance of the proposed algorithm on the datasets taken from the cybersecurity, bioinformatics, and finance domains. The results demonstrate that incorporating two vigilance tests coupled with a tolerance-based feedback mechanism significantly improved the median Area under the receiver operating characteristic curve (AUC) and median cardinality across all datasets.