Skip to main content
Sequencing & Pangenomes

Multi-task Genomic Representation Learning Based on DNA Large Language Models

Authors
  • Lei Wei (China Agricultural University)
  • Chenghao Zhai (China Agricultural University)
  • Yuzhe Wang (China Agricultural University)
  • Xiaoxiang Hu (China Agricultural University)

Abstract

Pre-trained DNA large language models specifically designed for agricultural species remain limited, with most existing models trained primarily on the human genome or a small number of model organisms, restricting their ability to capture species-specific regulatory features in livestock genomes. Transformer-based DNA large language models (DNA LLMs) have emerged as powerful tools for learning sequence representations directly from raw genomic sequences; however, their application to large-scale and cross-species genomic analysis remains challenging. In particular, the quadratic computational complexity of standard Transformer architectures limits modeling of long genomic regions, and many existing DNA LLMs adopt context modeling strategies that do not fully integrate upstream and downstream regulatory information or explicitly account for the reverse-complement symmetry inherent to DNA sequences. In this study, we develop a cross-species DNA foundation model for mammalian genomes, with a focus on agricultural animals such as pig, cattle, and sheep. The model is trained using a self-supervised masked language modeling framework on reference genomes of multiple agricultural species, where random subsets of nucleotides are masked and predicted from their surrounding context to learn generalizable sequence representations. These pre-trained representations are subsequently transferred to downstream genomic prediction tasks through task-specific fine-tuning. Model evaluation is conducted using publicly available benchmark datasets curated from a recent large-scale functional annotation study of livestock genomes, together with additional animal epigenomic data collected from the NCBI Sequence Read Archive, spanning multiple species, tissues, and epigenomic assays. Across a comprehensive cross-species benchmark covering promoter and enhancer activity prediction, chromatin accessibility inference, and trait-associated sequence prediction, our model consistently outperforms conventional supervised baselines and six existing DNA foundation models, achieving an average improvement of approximately 5% in prediction accuracy, with particularly strong performance in cross-species transfer settings. These results demonstrate the effectiveness and robustness of reference-genome-based, masked self-supervised pre-training for DNA LLMs and highlight the importance of combining large-scale self-supervised learning with biologically informed fine-tuning for cross-species functional genomics. Overall, this work provides a unified and extensible framework for genomic modeling and epigenetic regulation analysis in agricultural animal species, offering practical value for functional genomics research and molecular breeding applications.

Keywords: 2026

How to Cite:

Wei, L., Zhai, C., Wang, Y. & Hu, X., (2026) “Multi-task Genomic Representation Learning Based on DNA Large Language Models”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285640. doi: https://doi.org/10.31274/wcgalp.23764

Rights: 1

Downloads:
Download PDF
View PDF

57 Views

15 Downloads

Published on
2026-02-25

Peer Reviewed