Pruning outdated phenotypic and genomic data to optimize Single-Step genomic evaluation
Abstract
Running single-step genomic evaluations for traits with extensive phenotypic and genomic datasets substantially increases computational demands. As national evaluation systems continue to expand, efficient data management is essential for maintaining sustainability and computational feasibility. This study investigates how outdated genomic and phenotypic records can be trimmed without compromising prediction accuracy.Calving difficulty in Irish dairy and beef herds is routinely scored by farmers on a four-point scale (1 = unassisted, 4 = veterinary assistance). A single-step genomic evaluation (ssGTBLUP) is applied to these traits dairy heifers (DH), dairy cows (DC), beef heifers (BH), and beef cows (BC) along with correlated traits birth weight (BWT) and birth size (BSize). The pedigree includes 27,651,328 animals, of which 2,258,091 are genotyped. For data pruning, de-regressed GEBVs (DRPG) were calculated to generate independent pseudo-observations. Effective record contributions (ERC) were derived from routine GEBV reliabilities to quantify each record's influence and used as model weights. Year of birth was used as a cut-off. In the first scenario, phenotypic records of animals born before 2010 were replaced with their sire's DRPG and ERC, reducing the dataset size (e.g., BC records decreased from 6,156,164 to 5,012,161, replaced by 94,570 sire DRPG). In the second scenario, these animals were replaced with their own DRPG maintaining the full dataset size but with a simplified model excluding contemporary group for older animals. For both scenarios, genotypes of pruned animals were excluded from the T-matrix construction. Genomic prediction accuracy was assessed at both animal and sire levels by correlating GEBVs from full and reduced datasets. In addition, correlations between GEBVs from the routine evaluation, the full dataset, and the pruned scenarios were examined to evaluate consistency across evaluations. The initial results from the data pruning analyses are promising. Pruning older phenotypic and genomic records substantially reduced computation time, with the ssGTBLUP model converging in 48 hours for the routine evaluation versus 31 hours for the pruned sire-based scenario using 10 CPUs, while maintaining high correlations (≈1.0) with the routine evaluation. Identical GEBVs mean and SD between the routine and pruned runs across all traits confirming strong consistency in predictions. Validation results also indicated improved prediction accuracy for all traits and both scenarios. For instance for DH trait accuracy increased from 0.67 (routine) to 0.77 (pruned sire-based). These findings indicate that selective pruning of outdated data can enhance computational efficiency without compromising, and in some cases improving, the accuracy of genomic predictions providing a practical framework for optimizing large-scale routine evaluations.
Keywords: 2026
How to Cite:
Evans, R., Naderi, S., Pabiou, T. & Olori, V., (2026) “Pruning outdated phenotypic and genomic data to optimize Single-Step genomic evaluation”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2286384. doi: https://doi.org/10.31274/wcgalp.23972
Rights: 1
Downloads:
Download PDF
View PDF
72 Views
19 Downloads