Multi-task Genomic Representation Learning Based on DNA Large Language Models
Abstract
Pre-trained DNA large language models specifically designed for agricultural species remain limited, with most existing models trained primarily on the human genome or a small number of model organisms, restricting their ability to capture species-specific regulatory features in livestock genomes. Transformer-based DNA large language models (DNA LLMs) have emerged as powerful tools for learning sequence representations directly from raw genomic sequences; however, their application to large-scale and cross-species genomic analysis remains challenging. In particular, the quadratic computational complexity of standard Transformer architectures limits modeling of long genomic regions, and many existing DNA LLMs adopt context modeling strategies that do not fully integrate upstream and downstream regulatory information or explicitly account for the reverse-complement symmetry inherent to DNA sequences. In this study, we develop a cross-species DNA foundation model for mammalian genomes, with a focus on agricultural animals such as pig, cattle, and sheep. The model is trained using a self-supervised masked language modeling framework on reference genomes of multiple agricultural species, where random subsets of nucleotides are masked and predicted from their surrounding context to learn generalizable sequence representations. These pre-trained representations are subsequently transferred to downstream genomic prediction tasks through task-specific fine-tuning. Model evaluation is conducted using publicly available benchmark datasets curated from a recent large-scale functional annotation study of livestock genomes, together with additional animal epigenomic data collected from the NCBI Sequence Read Archive, spanning multiple species, tissues, and epigenomic assays. Across a comprehensive cross-species benchmark covering promoter and enhancer activity prediction, chromatin accessibility inference, and trait-associated sequence prediction, our model consistently outperforms conventional supervised baselines and six existing DNA foundation models, achieving an average improvement of approximately 5% in prediction accuracy, with particularly strong performance in cross-species transfer settings. These results demonstrate the effectiveness and robustness of reference-genome-based, masked self-supervised pre-training for DNA LLMs and highlight the importance of combining large-scale self-supervised learning with biologically informed fine-tuning for cross-species functional genomics. Overall, this work provides a unified and extensible framework for genomic modeling and epigenetic regulation analysis in agricultural animal species, offering practical value for functional genomics research and molecular breeding applications.
Keywords: 2026
How to Cite:
Wei, L., Zhai, C., Wang, Y. & Hu, X., (2026) “Multi-task Genomic Representation Learning Based on DNA Large Language Models”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285640. doi: https://doi.org/10.31274/wcgalp.23764
Rights: 1
Downloads:
Download PDF
View PDF
57 Views
15 Downloads