Skip to main content
Sequencing & Pangenomes

Inferring Missing Genotypes from Partially Observed Single Nucleotide Polymorphism Data Using Machine Learning-Based Models

Authors
  • Weronika Zawadzka (Wroclaw University of Environmental and Life Sciences)
  • Joanna Szyda orcid logo (Wroclaw University of Environmental and Life Sciences)
  • Magdalena Frąszczak (Wroclaw University of Environmental and Life Sciences)

Abstract

The goal of genomic imputation is to predict missing genotypes from partially observed single-nucleotide polymorphisms (SNP), which often occurs when individuals genotyped by an oligonucleotide arrays are to be imputed to the level of whole genome sequence variation. Traditional tools, such as Beagle, predict missing genotypes using Hidden Markov Models while our study explores the use of deep learning for SNP imputation, investigating how different model architectures, training strategies, and masking approaches affect imputation accuracy. We evaluate four AutoEncoder (AE) based approaches: (i) a naïve AE implementing dense neural networks, (ii) an AE based on dense neural networks implementing input embedding, (iii) a convolutional AE, and (iv) a transformer-based AE utilising the BERT model. Imputation accuracy obtained by Beagle served as a baseline. The idea behind using DL based architectures was a better restoration of rare genotypes, thanks to a larger flexibility of exploring the boundaries of the parameter space potentially offered by DL approach. The material consists of a subset of bulls from the 1000 Bull Genomes Project, using a fragment of chromosome 28 comprising the first 500 SNPs, with 742 bulls for training and 186 bulls for testing. A major challenge in SNP data is the overrepresentation of the homozygous reference genotype, which can make evaluation metrics misleading - models may achieve high overall accuracy while performing poorly on rare genotypes involving rare SNPs with low minor allele frequency (MAF). To address this, architectures were optimized by exploring model depth, layer sizes, activation functions, convolution and window sizes, learning rates, and regularization. Masking strategies were designed to circumvent the influence of common genotypes on model optimisation. Loss functions such as focal loss or balanced cross-entropy, along with evaluation metrics like macro F1, were applied to better capture performance across all genotypes. Performance was also assessed for different MAF levels to evaluate how well models impute genotypes with varying allele frequency. Accurate imputation of rare variants is particularly important for downstream analyses, such as genome-wide association studies or genomic selection. Architecture, parameterization, and masking strongly affected imputation performance. BERT achieved the highest F1 (0.93) and F1-macro (0.91), outperforming Beagle (F1 = 0.88, near-real-time inference). The convolutional AE (F1 = 0.88, F1-macro = 0.84), less computationally heavy (training approx. 8 min on AMD Ryzen 7 4800HS vs. 3 h for BERT), performed comparably to Beagle. The other two models showed lower performance, highlighting the utility of BERT and the convolutional autoencoder for SNP imputation. In conclusion, this study presents the development of deep learning models for SNP imputation, addressing i) selection of input encoding, (ii) model training, (iii) handling of class imbalance inherent to SNP data, (iv) choice of appropriate loss functions, and (v) scaling and computational aspects.

Keywords: 2026

How to Cite:

Zawadzka, W., Szyda, J. & Frąszczak, M., (2026) “Inferring Missing Genotypes from Partially Observed Single Nucleotide Polymorphism Data Using Machine Learning-Based Models”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285635. doi: https://doi.org/10.31274/wcgalp.23763

Rights: 1

Downloads:
Download PDF
View PDF

58 Views

16 Downloads

Published on
2026-02-26

Peer Reviewed