Development of pangenome for indigenous cattle using Next Generation Sequencing
Abstract
India is home to the world's largest cattle population, with over 50 distinct Bos indicus breeds that represent an invaluable genetic resource known for their exceptional adaptability and resilience. Despite their importance to the agricultural economy, the vast genomic diversity of these native dairy breeds is not fully captured by single, linear reference genomes. This limitation spurred the development of a comprehensive pangenome for Indian Desi cattle to better characterize their unique genetic architecture. To achieve this, 68 genomes from seven breeds were sequenced using Illumina short-read sequencing. A MaSuRCA-based pipeline was developed to assemble the reads and identify 13,065 non-reference novel sequences (NRNS) spanning approximately 41 Mbp. Utilizing the Animal QTL database for analysis and BWA for mapping, the study revealed that the majority of these NRNS were unique to Desi cattle. These sequences were significantly enriched in genic regions with functional roles associated with milk production quantitative trait loci (QTLs). This pangenome strategy markedly enhanced read mapping accuracy and uncovered previously hidden variants. The study was further expanded by sequencing five prominent Indian milch breeds using linked-read technology. Draft assemblies were generated via Supernova, followed by hybrid assemblies using RAGTAG, resulting in genomes ranging from 2.70 to 2.78 Gb with high BUSCO scores (94.1-95.5%). Comparison against the Brahman reference genome identified 6,844 non-reference unique insertions (NUIs), synonymous with the previously identified NRNS, totaling 7.57 Mb. These insertions were further genotyped across larger populations, revealing 2,312 Bos indicus common insertions (BICIs) that are ubiquitous across breeds. Notably, 926 BICIs were located within protein-coding genes tied to critical functions and QTLs for milk yield, reproduction, and health. SyRI-based synteny analysis demonstrated high conservation (87-95%) with the Brahman genome, yet identified substantial structural rearrangements including inversions, translocations, and duplications ranging from 19.84 to 153.16 Mb per genome. Synteny diversity analysis uncovered 10,643 perfectly collinear regions (87.3 Mb) and 6,622 hotspots of rearrangement (HOT regions; 55.18 Mb). These HOT regions were significantly enriched with immune-related genes, such as the MHC, NKC, and LRC clusters, providing a genetic basis for the superior disease resistance of Zebu cattle. To facilitate global research and ensure data accessibility, we developed GauPanDB. This user-friendly portal allows researchers to browse, visualize, compare, and download these critical genomic and pangenomic datasets. GauPanDB serves as a vital resource for genomics-assisted breeding and the long-term conservation of India's indigenous cattle.
Keywords: 2026
How to Cite:
Azam, S., Sahu, A., Kadivella, M., Gandham, R., Majumdar, S. & Rath, S., (2026) “Development of pangenome for indigenous cattle using Next Generation Sequencing”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2295601. doi: https://doi.org/10.31274/wcgalp.24348
Rights: 1
Downloads:
Download PDF
View PDF
88 Views
14 Downloads
