Publications scientifiques
Local Feature Relevance and Cluster Order Selection in Model-Based Clustering for Mixed Data via Variational Inference (ouvre dans un nouvel onglet)
Auteurs
Université de Yaoundé I
Autres auteurs
Wilson Toussile
Publications scientifiques
Wilson Toussile
Fotso Simeon
HAL (Le Centre pour la Communication Scientifique Directe)
<div> Clustering high-dimensional mixed-type data is challenging when the number of clusters is unknown and datasets are cluttered with irrelevant features that mask the true structure. We address both challenges with a unified Bayesian variational framework for finite mixture models. A sparse symmetric Dirichlet prior on overfitted mixtures automatically prunes redundant clusters, bypassing costly model selection. At its core is a Local Feature Relevance Model (LFRM), with latent binary activation variables letting each feature's discriminative power vary by cluster, capturing the realistic case where a feature separates some subgroups but is pure noise for others, with the common, population-wide profile recovered as a special case. These activation variables are embedded within exponential family distributions, so the framework handles continuous, count, and categorical data natively, via a modular Coordinate Ascent Variational Inference (CAVI) algorithm. Since mean-field approximations can be overconfident in high dimensions, a deterministic posthoc stage refines the in-model estimate by pruning spurious clusters and filtering features through a strict Expected Information Gain (EIG) criterion. Simulations confirm the design recovers the true cluster count, strictly controls the False Discovery Rate, and sharpens cluster-specific signal detection relative to a shared-relevance assumption. We demonstrate practical utility on a synthetic clinicogenomic cohort and the real-world UCI Heart Disease dataset, robustly isolating clinically relevant variables against a heterogeneous background. The framework is available as an open-source Julia package. </div>