Skip to main content
Phenomics

On the promises and pitfalls of prediction: how vision-based phenotyping affects breeding value estimation

Authors
  • Yuuko Xue orcid logo (Aarhus University)
  • John Bastiaansen (Animal Breeding and Genomics, Wageningen University & Research)
  • Grum Gebreyesus (Aarhus University)

Abstract

Phenotyping is essential to genetic evaluation since the reliability of estimated breeding values (EBV) hinges on both the accuracy and the volume of trait records available. However, routine phenotyping is challenging, especially for traits that are expensive or invasive to measure, require destructive sampling, or only express late in life. Vision-based phenotyping offers a scalable alternative by extracting visually discernible features from images or video that correlate with the target traits, enabling high-throughput, non-invasive approximation to ground-truth phenotypes. These computer-vision models rely on machine learning and therefore require strict train-test separation, with test data used only once to evaluate whether models generalize well to unseen data. However, this introduces two major challenges. First, most machine learning models require large datasets for training, but for traits that are difficult to measure, available data is often limited. Data splitting further reduces such availability. As models are aimed to be applied to selection candidates, data collected from these individuals should be excluded from training to avoid data leakage. Second, even when test accuracy closely matches training accuracy, small disparities between predicted and true phenotypes may still influence EBV and alter rankings. Previous studies have examined how different data-splitting strategies in computer vision, such as k-fold cross-validation, affect phenotypic prediction accuracy, but the extent to which EBV is affected remains unclear. This uncertainty hinders the reliable replacement of manual measurements with vision-based predictions. Here, we investigated how discrepancies between true and predicted phenotypes generated by computer vision models propagate to EBV re-ranking. Traits with heritability of 0.1, 0.3, and 0.7 were simulated in populations of 2,000 individuals. Computer vision models were assigned training accuracies of 0.98 and 0.8, corresponding to test accuracy ranges of 0.7-0.95 and 0.6-0.75, and two data-splitting strategies (80:20 and 50:50, training: test) were evaluated. For each scenario, predictions were generated across 1,000 simulation rounds, and average relative deviations between estimated breeding values (EBVs) and predicted EBVs were calculated. Preliminary results indicate that relative differences between EBVs and predicted EBVs occurred even under near-perfect model performance and across all heritability. Increased residual variance did not bias EBVs at the population level but reduced EBV accuracy and increased shrinkage toward zero. Differences in residual variance between training and test sets had negligible effects on EBVs, suggesting that training records can be reused for model validation. Future analyses will investigate when re-ranking occur and validate these findings using real image data.

Keywords: 2026

How to Cite:

Xue, Y., Bastiaansen, J. & Gebreyesus, G., (2026) “On the promises and pitfalls of prediction: how vision-based phenotyping affects breeding value estimation”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2286398. doi: https://doi.org/10.31274/wcgalp.23983

Rights: 1

Downloads:
Download PDF
View PDF

100 Views

19 Downloads

Published on
2026-02-26

Peer Reviewed