Novel strategies to improve the computational efficiency of large-scale genomic evaluations
- Renzo Bonifazi (Wageningen University & Research)
- Mario Calus (Wageningen University & Research)
- Sharif Islam (Wageningen University & Research)
- Natalia Leite (Topigs Norsvin)
- Matias Schrauf (Wageningen University & Research)
- Jan ten Napel
(Wageningen University & Research)
- Dianne van der Spek (Topigs Norsvin Research Center)
- Jeremie Vandenplas (Wageningen University & Research)
Abstract
Increasing volumes of phenotypic and genomic data are being collected for selection purposes in animal populations, enabled by reduced genotyping costs and new phenotyping technologies. However, such large-scale phenotypic and genomic datasets pose substantial computational challenges for routine genomic evaluations as genomic estimated breeding values (GEBVs) must be delivered within limited timeframes. This study aimed to develop and validate novel strategies that reduce data from older generations or summarise it into recent ones, and to assess the impact of those strategies on both accuracy and computational efficiency in large-scale single-step genomic evaluations. Initial strategies focused on reducing pedigree, genomic, or phenotypic data, or a combination of these. Additionally, two novel strategies were developed to summarise historical data into recent generations, either at the individual level using de-regressed proofs (DRP) and effective record contributions (dERC), or at the genomic level through estimated SNP effects and associated prediction error covariances (PEC). All strategies were implemented using data from a two-way crossbred pig population, comprising ~4.2M pedigree records, ~465K genotypes, and ~4.1M phenotypes for a piglet trait (heritability of 0.42). All strategies were validated against a single-step genomic evaluation using all available data (hereafter called full evaluation). Validation metrics included Pearson correlation (ρ), dispersion bias (regression slope, b1), and root mean square error (RMSE) expressed in genetic standard deviations, between GEBVs of the full evaluation and of each strategy. Validation animals were young selection candidates born in the last two years, with a phenotype, with or without genotype. Strategies that reduce data from older generations indicate that retaining three generations of pedigree from phenotyped and genotyped animals, the two most recent generations of phenotypic data, and genotypes of animals with own or progeny phenotypes reduces time up to 53% and memory use up to 38%, while maintaining good agreement with the full evaluation (ρ ≥ 0.99, 0.97 ≥ b1 ≥ 0.98, and RMSE ≤ 0.15). Novel strategies that summarise historical data into DRP and dERC for historical parents, or estimated SNP effects and PEC, further improved the agreement of the selection candidates' GEBVs: ρ ≥ 0.99, 0.99 ≥ b1 ≥ 1.00, and RMSE ≤ 0.11, and ρ ≥ 0.99, 0.98 ≥ b1 ≥ 1.00, and RMSE ≤ 0.04, respectively. The two strategies reduced computational time up to 76% and 34%, respectively. Finally, summarising historical data into DRP and dERC improved the consistency of parents' GEBVs (compared with truncating historical data, ρ improved from 0.82 to 0.98). Similar patterns were observed when selection candidates did not have phenotypes available at the moment of selection. Overall, our study shows that summarising historical data into more recent generations can maintain or improve the consistency of GEBVs while substantially reducing computational requirements in large-scale single-step genomic evaluations.
Keywords: 2026
How to Cite:
Bonifazi, R., Calus, M., Islam, S., Leite, N., Schrauf, M., ten Napel, J., van der Spek, D. & Vandenplas, J., (2026) “Novel strategies to improve the computational efficiency of large-scale genomic evaluations”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2248261. doi: https://doi.org/10.31274/wcgalp.23382
Rights: 1
Downloads:
Download PDF
View PDF
156 Views
32 Downloads