EcoSta 2026: Start Registration
View Submission - EcoSta2026
A1696
Title: Covariance-based clustering and biclustering via the heterogeneous block covariance model and variants Authors:  Yunpeng Zhao - Colorado State University (United States) [presenting]
Xiang Li - The George Washington University (United States)
Ning Hao - University of Arizona (United States)
Qing Pan - George Washington University (United States)
Mayson Zhang - Colorado State University (United States)
Abstract: Clustering methods are traditionally designed to group data points in a (possibly high-dimensional) Euclidean space. A different but equally important task involves clustering at the feature level, which arises naturally in gene expression analysis but has received relatively limited attention in the statistics literature. In bioinformatics, it is often desirable to identify sets of highly correlated genes whose expression levels are jointly regulated and act synergistically to carry out the same biological functions. The heterogeneous block covariance model (HBCM) addresses this problem by characterizing community structure directly within the covariance matrix while accounting for heterogeneity in how features connect within communities. A novel variational expectation-maximization algorithm enables efficient estimation of group memberships. Theoretical analysis establishes the consistency of membership recovery, and simulation studies demonstrate the superior performance of HBCM compared with existing methods. An extension of HBCM to a joint clustering and biclustering framework simultaneously partitions both features and samples. This generalization introduces new computational and theoretical challenges, for which principled solutions are proposed. Applications to real gene expression data highlight the model's ability to uncover biologically meaningful clusters and biclusters.