A1698
Title: A mixture of logistic-normal multinomial distribution and its extension to biclustering
Authors: Sanjeena Dang - Carleton University (Canada) [presenting]
Abstract: Compositional count data is becoming increasingly prevalent across a variety of fields, including bioinformatics driven by advances in next-generation sequencing technologies, and topic modelling driven by large-scale digitalization of document collections. Analyzing such compositional data presents many challenges because they are restricted to a simplex, which complicates standard statistical analysis. While a Dirichlet-multinomial (DM) model can be used for such tasks, it fails to capture complex correlation structures between the variables. In contrast, the logistic-normal multinomial (LNM) distribution provides greater flexibility in capturing correlations but is computationally intensive. A computationally efficient framework for parameter estimation in LNM-based models using variational Gaussian approximations has been developed to overcome these limitations. This approach significantly reduces the computational burden and enables application to large-scale datasets. Building on this framework, a recent development of the LMN mixture model-based biclustering approach that simultaneously clusters observations and features (i.e., variables) in a dataset is discussed. Clustering of features is simultaneously achieved by introducing a block-diagonal covariance structure. The proposed approach is demonstrated on both simulated and real datasets.