Integrating autoencoder deep neural networks with principal component-based methods for subpopulation identification

Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

Crop Breeding and Applied Biotechnology

Abstract

This study evaluated statistical dimensionality reduction techniques—Principal Component Analysis (PCA), Sparse PCA (SPCA), Independent PCA (IPCA), deep learning based on autoencoders (AE), and their hybrid combinations—for subpopulation identification in genomic data. High-dimensional single nucleotide polymorphism (SNP) datasets pose computational challenges and may hinder clustering performance. We compared methods based on clustering agreement with known subpopulation labels using Oryza sativa data (413 genotypes; 44,100 SNPs). PCA and SPCA exhibited strong performance (Adjusted Rand Index [ARI] = 0.715), with PCA achieving the lowest computational time (4.89 s). The standalone autoencoder performed poorly, likely due to the high?dimensional (p >> n) setting. Among the hybrid approaches, IPCA-AE achieved the highest agreement (ARI = 0.758), demonstrating that combining dimensionality reduction with nonlinear representation learning improves population structure identification.

Description

Citation

FERREIRA, Camila Azevedo. et al. Integrating autoencoder deep neural networks with principal component-based methods for subpopulation identification. Crop Breeding and Applied Biotechnology, Viçosa, v. 26, n. 3, p. 01-11, 2026.

Collections

Endorsement

Review

Supplemented By

Referenced By