EcoSta 2026: Start Registration
View Submission - EcoSta2026
A1926
Title: Clustering high-dimensional continuous and categorical data via variable selection Authors:  Somi Cha - Chonnam National University (Korea, South) [presenting]
Kwangmin Lee - Chonnam National University (Korea, South)
Abstract: A test-based variable selection framework is proposed for clustering high-dimensional data containing both continuous and categorical variables. When only a small subset of variables carries the cluster structure, clustering based on all variables may fail to recover the underlying groups. The main idea is to use initial cluster assignments as pseudo-labels, turning variable selection for clustering into variable-wise testing with false discovery rate control. The procedure first obtains initial cluster assignments from a dummy-expanded representation of the data, identifies informative variables through multiple testing, and then re-clusters the data using only the selected variables. A misclustering bound is established for the initial clustering step under the corresponding dummy-expanded model, showing that the initial assignments become asymptotically close to the true cluster labels under a sufficient signal condition. Simulation experiments demonstrate that the proposed framework improves clustering accuracy compared with using all variables, and an application to the 2024 Seoul Survey on Foreign Residents demonstrates that the selected variables yield interpretable cluster profiles.