bioRxiv

Adversarial random forests for omics synthesis

Preprint: computational studyAI & data

Abstract

Data availability is critical for understanding complex disease pathways and developing robust predictive models. Although high-throughput omics technologies have improved insight into disease mechanisms, data acquisition from inaccessible tissues such as the central nervous system remains a major limitation, causing small sample sizes and complicating early prediction of neurodegenerative disorders such as Alzheimer's and Parkinson's diseases. Generative modeling has emerged as a powerful approach for synthesizing data to support downstream clustering and prediction with small sample size, but existing methods rarely handle high-dimensional tabular omics data effectively. Adversarial random forests (ARFs) provide a well-performing framework for tabular data generation but are not designed for high-dimensional settings. To address this limitation, we introduce high-dimensional ARF ( h -ARF), an extension of ARF optimized for integrated clinical and high-dimensional omics data. Using benchmarks across nine datasets and eight performance metrics, we show that h -ARF better preserves both feature distributions, and downstream clustering and prediction utilities compared with ARFs. The method is implemented in the opensource R package harf, available on CRAN.

The paper

Emory University; University of Bremen

bioRxiv, 15 Sep 2026, Preprint, not peer-reviewed

doi.org/10.64898/2026.09.09.750490PubMed 42780183