Skip to main content
Estimation & Prediction

Integrating Multi-Omics Data in Deep Learning Prediction of Pig Production Traits

Authors
  • Levi Ayres (Animal Breeding and Genomics, Wageningen University & Research, 6700 AH, the Netherlands)
  • Mario Calus (Wageningen University & Research)
  • Peter Karlskov-Mortensen (University of Copenhagen)
  • Maria Luigi-Sierra (University of Copenhagen)
  • Miriam Piles (Institute of Agrifood Research and Tecnology - IRTA)
  • Yuliaxis Ramayo-Caldas (Institute of Agrifood Research and Tecnology - IRTA)
  • Ioanna-Theoni Vourlaki (Institute of Agrifood Research and Technology)

Abstract

Integrating multiple omics layers through deep learning provides a powerful approach to improve the prediction of complex production traits in livestock. Traditional genomic prediction models often rely solely on SNP data, limiting their ability to capture the full biological complexity underlying phenotypic variation. In this study, we explored deep learning frameworks for integrating genomic, transcriptomic, and epigenomic data to predict two economically important pig production traits: average daily gain (ADG) and backfat thickness (BFT).We analyzed data from 443 pigs from three commercial breeds-Yorkshire (YY), Duroc (DD), and Landrace (LL). For each animal, genome-wide SNP genotypes were obtained, alongside gene expression profiles and DNA methylation data (CpG sites) from longissimus dorsi muscle. We applied an intermediate omics integration strategy, processing each omic layer through separate subnetworks before merging them into a shared latent representation.To assess the impact of feature dimensionality and selection on predictive performance, we compared models using: (i) all available features, (ii) reduced representations obtained via dimensionality reduction methods such as principal component analysis (PCA), and (iii) selected subsets of biologically relevant or statistically important features identified using partial least squares (PLS) and multi-omics factor analysis (MOFA). For PLS and MOFA, different subset sizes of 10, 100, 200, 500, and 1,000 features were evaluated. Model performance was first evaluated to assess whether the models obtained using data from one breed could be transferred to others. Based on the results, we then implemented a stratified 5-fold cross-validation using the best-performing method from the previous step, ensuring that each fold contained animals from all breeds in approximately equal proportions. Predictive ability was quantified as the Pearson correlation (r) between observed and predicted phenotypes in the test sets.Our results showed that predictive correlations ranged from -0.23 to 0.49 for ADG and from -0.17 to 0.64 for BFT across all breeds and models. Among the feature-selection strategies, PLS provided the best performance, yielding the highest median correlations across subset sizes: 0.38 for DD, 0.40 for LL, and 0.42 for YY for ADG; 0.50 for DD, 0.40 for YY, and 0.49 for LL for BFT. Under the stratified 5-fold cross-validation framework, PLS produced correlations ranging from 0.17 to 0.59 for ADG and from 0.04 to 0.66 for BFT across folds. The best-performing PLS-selected subsets reached median predictive correlations of 0.51 for ADG and 0.46 for BFT. These findings highlight the importance of feature selection, showing that statistically selected omics features capture the most relevant information for trait prediction while reducing model complexity.

Keywords: 2026

How to Cite:

Ayres, L., Calus, M., Karlskov-Mortensen, P., Luigi-Sierra, M., Piles, M., Ramayo-Caldas, Y. & Vourlaki, I., (2026) “Integrating Multi-Omics Data in Deep Learning Prediction of Pig Production Traits”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285384. doi: https://doi.org/10.31274/wcgalp.23672

Rights: 1

Downloads:
Download PDF
View PDF

90 Views

31 Downloads

Published on
2026-02-26

Peer Reviewed