Skip to main content
Software

Systematic High-Throughput Genomic Data Processing and Quality Control in Large-Scale Bovine Genetic Evaluation

Authors
  • Hailin Su orcid logo (STgenetics)
  • Haniel Cedraz (STgenetics)
  • Adam Utsunomiya (STgenetics)
  • Pablo Ross (STgenetics)
  • Nader Deeb (STgenetics)

Abstract

The successful implementation of genomic evaluations and genetic services in livestock requires robust and scalable pipelines for quality control (QC) and data curation of massive datasets. This study details our novel, automated pipelines designed to efficiently process and analyze over 2 million bovine genotype records, with continuous integration of tens of thousands new genotypes weekly, ensuring high-quality and accurate genetic services for our clients. 1). Identification and Consolidation of Homogeneous Genotypes: Data errors, such as mislabeled or duplicated samples, including those from clones, pose a significant challenge to genetic evaluation accuracy. We developed a two-step pipeline for identifying homogeneous genotypes. The initial step employs rapid matrix operations on a shared core set of approximately 5,000 Single Nucleotide Polymorphisms (SNPs) across various commercial panels to quickly generate a list of highly matching genotype candidates from the >2 million database. The second step performs a detailed concordance check across all shared SNPs (ranging from 6,000 to 60,000+ SNPs, depending on the panel overlaps) for each new genotype against its candidate list. This system ensures the correct identification and consolidation of identical or near-identical genotypes. 2). Genomic Parent Discovery at Big-Data Scale: Efficient identification of potential parents in a large-scale database is critical for pedigree reconstruction and validation. Our method leverages a multi-stage approach utilizing the Opposite Homozygosity (OH) test for efficient candidate screening. The first stage uses the common ~5,000 SNP panel to rapidly filter a list of potential parents. The second stage then performs a fine-grained OH test on all shared SNPs between the animal and each of its screened candidates. 3). Determining Sex from X Chromosome Genotypes: Accurate sex identification is vital for data quality checks. We investigated the use of X-chromosome SNP genotypes across different commercial panels to determine sex. The methodology partitions the X chromosome into several segments of similar base pair length. The heterozygosity rate is calculated within each segment, and these seven segment-specific heterozygosity rates are then fitted into a Support Vector Machine (SVM) model for robust sex prediction, accounting for variability across different SNP arrays. 4). Identifying 'Pure-Bred' Populations: Genomic breed composition prediction relies on the selection of representative 'pure-bred' base populations. We propose a distribution-based approach to objectively identify genetically "pure" individuals. Our method uses the distribution of the 'Deviation from the Population Mean Genotype' as the primary metric, which is computationally efficient for multi-million data points. This approach facilitates a rigorous definition of population purity, which enables our accurate, automated genomic breed percentage estimations. Finally, these automated pipelines have been integrated into our weekly data processing workflow, ensuring continuous and rigorous QC and genetic evaluation of newly submitted genotypes, maintaining the integrity and utility of our expansive bovine genomic database.

Keywords: 2026

How to Cite:

Su, H., Cedraz, H., Utsunomiya, A., Ross, P. & Deeb, N., (2026) “Systematic High-Throughput Genomic Data Processing and Quality Control in Large-Scale Bovine Genetic Evaluation”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285970. doi: https://doi.org/10.31274/wcgalp.23827

Rights: 1

Downloads:
Download PDF
View PDF

133 Views

22 Downloads

Published on
2026-02-25

Peer Reviewed