Objective. Higher-order statistical dependencies can encode cooperative structure in complex systems, but are often missed by pairwise analyses and supervised pipelines optimized for discrimination. We aim to develop an unsupervised feature ranking framework that prioritizes genes by higher-order synergy within transcriptomic communities. Approach. We combine principal component analysis derived latent components with O-information to quantify gene level synergy within network communities. Using hepatocellular carcinoma (HCC) microarray data (GSE102079), we start from gene communities identified in our previous community detection workflow and analyze them without supervised pre filtering. For each community, we select latent components through permutation based criteria, compute O-information between each gene and the retained components, assess significance using surrogate testing, and rank genes by synergy. Top ranked genes are selected using knee based criteria on ranked synergy profiles. Main results. Synergy based ranking consistently highlights genes associated with functions related to HCC, such as metabolic remodeling, extracellular matrix dynamics, epithelial-mesenchymal transition, angiogenesis, immune modulation and oxidative stress responses. Several high-synergy genes supported by literature were previously discarded by Boruta, indicating that synergy prioritizes cooperative module roles rather than solely discriminative markers. Significance. The proposed method provides a reproducible, fully unsupervised tool for feature ranking based on higher-order interactions, complementing supervised feature selection approaches and offering a general domain agnostic tool applicable to any samples by features dataset.
Higher-order synergy-based ranking in transcriptomic communities via latent factors and O-information
Antonio, Lacalamita;Marlis, Ontivero-Ortega;Alessandro, Fania
;Nicola, Amoroso;Roberto, Bellotti;Sebastiano, Stramaglia;Alfonso, Monaco
2026-01-01
Abstract
Objective. Higher-order statistical dependencies can encode cooperative structure in complex systems, but are often missed by pairwise analyses and supervised pipelines optimized for discrimination. We aim to develop an unsupervised feature ranking framework that prioritizes genes by higher-order synergy within transcriptomic communities. Approach. We combine principal component analysis derived latent components with O-information to quantify gene level synergy within network communities. Using hepatocellular carcinoma (HCC) microarray data (GSE102079), we start from gene communities identified in our previous community detection workflow and analyze them without supervised pre filtering. For each community, we select latent components through permutation based criteria, compute O-information between each gene and the retained components, assess significance using surrogate testing, and rank genes by synergy. Top ranked genes are selected using knee based criteria on ranked synergy profiles. Main results. Synergy based ranking consistently highlights genes associated with functions related to HCC, such as metabolic remodeling, extracellular matrix dynamics, epithelial-mesenchymal transition, angiogenesis, immune modulation and oxidative stress responses. Several high-synergy genes supported by literature were previously discarded by Boruta, indicating that synergy prioritizes cooperative module roles rather than solely discriminative markers. Significance. The proposed method provides a reproducible, fully unsupervised tool for feature ranking based on higher-order interactions, complementing supervised feature selection approaches and offering a general domain agnostic tool applicable to any samples by features dataset.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


