Integrating autoencoder deep neural networks with principal component-based methods for subpopulation identification
Loading...
Files
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Crop Breeding and Applied Biotechnology
Abstract
This study evaluated statistical dimensionality reduction techniques—Principal Component Analysis (PCA), Sparse PCA (SPCA), Independent PCA (IPCA), deep learning based on autoencoders (AE), and their hybrid combinations—for subpopulation identification in genomic data. High-dimensional single nucleotide polymorphism (SNP) datasets pose computational challenges and may hinder clustering performance. We compared methods based on clustering agreement with known subpopulation labels using Oryza sativa data (413 genotypes; 44,100 SNPs). PCA and SPCA exhibited strong performance (Adjusted Rand Index [ARI] = 0.715), with PCA achieving the lowest computational time (4.89 s). The standalone autoencoder performed poorly, likely due to the high?dimensional (p >> n) setting. Among the hybrid approaches, IPCA-AE achieved the highest agreement (ARI = 0.758), demonstrating that combining dimensionality reduction with nonlinear representation learning improves population structure identification.
Description
Citation
FERREIRA, Camila Azevedo. et al. Integrating autoencoder deep neural networks with principal component-based methods for subpopulation identification. Crop Breeding and Applied Biotechnology, Viçosa, v. 26, n. 3, p. 01-11, 2026.
