Disentanglement learning to deconfound neuroimaging-environmental data: application to multi-site data harmonization in psychiatry
Ines W. Sampaio, Anna M. Bianchi, Stefan Borgwardt, …
DOI:
10.1109/ACCESS.2026.3712218
Abstract
Debiasing and deconfounding represent fundamental challenges in deep learning (DL) applications to neuroimaging data, where confounding effects can significantly compromise model reliability and generalizability. Traditional debiasing approaches often require preprocessing steps that may introduce data leakage or fail to integrate seamlessly with end-to-end DL pipelines. To address these limitations, we propose to leverage disentanglement representation learning for neuroimaging debiasing in DL analyses, employing a light-weight plug-in that can be flexibly embedded within neural network architectures. We demonstrate the application of this method in a multi-site neuroimaging harmonization study, with a site effect disentanglement strategy, focusing on psychiatric disorders. Our approach incorporates the deconfounding plug-in, based on a cross-covariance (xcov) loss penalty, in a brain-environment fusion DL model, for the extraction of relevant, confounder-free multi-source integrated latent features. Brain functional connectivity data, from functional magnetic resonance imaging, and environmental risk information were collected from the multi-site PRONIA cohort, comprising individuals with recent-onset depression and psychosis, at clinical high-risk, and controls.We compared our deconfounding strategy against the established ComBat harmonization approach, for site-effects removal, and for improving downstream classification tasks, testing psychiatric group and sex discrimination. Experimental results demonstrate that our deconfounder plug-in effectively mitigated site effects while improving identifiability of biological signals of interest, achieving comparable downstream group and sex classification performance compared to when ComBat was used. These findings highlight the broader potential of embedding deconfounding processes directly into DL pipelines, providing an end-to-end framework for addressing various confounding variables beyond site effects, advancing debiasing strategies and enabling fairer and more reliable DL models for neuroimaging analyses.