Publicación

Selecting What Matters: A Regression-Based Embedding for Variable Influence Discovery

Rivadeneira, Franco

Resumen

Unsupervised feature selection remains a fundamental challenge in high-dimensional data analysis, particularly when labels are unavailable or unreliable. This paper introduces Influential Variable Embedding (IVE), a novel regressionbased algorithm that assigns directional influence scores to each variable by quantifying its ability to predict all other variables. Unlike projection-based methods such as PCA or t-SNE, IVE maintains full interpretability and yields a direct variable importance ranking, enabling exploratory data analysis and preprocessing for downstream tasks. Experiments on the TCGA Pan-Cancer gene expression dataset demonstrate IVE's ability to identify biologically relevant genes and capture intrinsic data structure, as reflected by its superior Silhouette Score. Despite lower class-separability in clustering evaluations, IVE excels in revealing globally influential features, making it a valuable tool for unsupervised analysis. Future work will explore scalability enhancements, nonlinear extensions, and integration with deep representation learning frameworks.

Autores y colaboradores

Authors

Rivadeneira, Franco