Leveraging Deep Learning Foundational Models for Genomic Prediction
Abstract
Traditional statistical approaches for genomic prediction (GP) of quantitative traits are largely restricted to linear additive effects of genetic variants. This framework limits models from capturing hierarchical relationships among the proximal and distal genomic contexts of variants, including regulatory domains, chromatin accessibility, and higher-order chromatin structure. Foundation models (FMs) are large neural networks trained to learn the DNA-sequence basis of diverse genomic and epigenomic tracks and offer a potential strategy for contextualizing variants. However, the generalizability of these models to livestock genomes remains largely unexplored. Here, we propose a strategy that leverages deep-learning FMs to derive per-variant feature embeddings that quantify long-range, sequence-mediated perturbations of regulatory potential for use as additive predictors in genome-wide regression. This implementation enables GP models to incorporate predicted multi-omic consequences of genetic variation that are inaccessible to conventional marker-based approaches. The proposed strategy integrates two modalities: (1) construction of variant-sequence embeddings that capture broader regulatory motif structure, and (2) combination of per-SNP embeddings with allelic dosages to generate individual-level predictors. The embedding framework is designed to encode DNA- and RNA-associated features, including sequence motifs, regulatory elements, and gene-expression dynamics. A publicly available pretrained DNA FM provides the initial framework for generating these embeddings, with future adaptation to porcine sequence through transfer learning. Multiplying the allelic-dosage matrix by the SNP-embedding matrix produces a per-individual representation of predicted regulatory potential that can be incorporated into genome-wide regression. Together, these modalities provide a biologically informed representation of how sequence-level variation may propagate to gene-regulatory and phenotypic effects. We evaluated this framework using 2,796 whole-genome-sequenced Duroc pigs and five agriculturally relevant traits: total teat number, back-fat thickness, loin-muscle depth, lean-meat percentage, and time spent eating per day. For approximately 800,000 biallelic SNPs, 196,608-bp reference and alternate sequence windows were processed with the deep learning sequence-to-function model, Enformer. Variant embeddings were derived from allele-specific differences in embeddings, and when combined with genotype dosages, to generated individual-level regulatory scores. BayesC models were fitted using genotype dosages alone, Enformer-derived scores alone, or both jointly, with predictive accuracy evaluated across five randomized 80/20 train-test splits. Enformer-derived scores captured non-zero phenotypic signal but underperformed genotype dosages alone, while the joint model produced no consistent improvement: accuracy decreased marginally for four traits and increased slightly for back-fat thickness. These results indicate that this implementations zero-shot Enformer embeddings are not sufficiently complementary to conventional dosages under an additive regression framework. Improvements will likely require livestock-specific fine-tuning, alternative embedding construction, or nonlinear integration of FM-derived features.
Keywords: 2026
How to Cite:
Weber, A. & Cheng, H., (2026) “Leveraging Deep Learning Foundational Models for Genomic Prediction”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2287220. doi: https://doi.org/10.31274/wcgalp.24237
Rights: 1
Downloads:
Download PDF
View PDF
51 Views
16 Downloads