Skip to main content
Sequencing & Pangenomes

Improving representation of global cattle diversity with pangenomes built by the Bovine Pangenome Consortium

Authors
  • Alexander Leonard (ETH Zürich)
  • Tobias van Elst (ETH Zürich)
  • Lloyd Low (Adelaide University)
  • Ben Rosen (USDA-ARS)
  • Tim Smith (USDA-ARS)
  • Hubert Pausch (ETH Zürich)

Abstract

The Bovine Pangenome Consortium (BPC) was established to create a comprehensive genomic resource representing global bovine diversity. In its first phase, the BPC collected 209 genome assemblies (24 public and 185 directly contributed) from 16 lead contacts spanning all six major continents, with 194 genomes passing quality thresholds. This diverse collection represents 77 unique cattle breeds and eight non-cattle species. The average genome size was 3.03±0.22 Gb, with a mean contig NG50 (based on the ARS-UCD2.0 genome reference size of 2.77 Gb) of 62.5±29.8 Mb. The mean artiodactyla BUSCO completeness was 98.1%, with some variation due to different sex chromosome complements and assembly types. Cluster analyses with reference alignment-derived SNPs recovered the expected population structure, revealing distinct clusters for European taurine, indicine, and east Asian taurine cattle. We also identified SNPs unique to each of the 77 genomes. Some breeds like Singida White or Nelore, despite having many SNPs, contained disproportionately few unique SNPs, while other breeds like Sarda and Chianina contained disproportionately many unique SNPs. With limited sample size (often n=1), we cannot definitively separate assembly errors from true variation, although variant metrics like the transition-transversion ratio were within the expected range. Many genomes were substantially more complete than the current cattle reference genome based on their chromosomal genome size, which allowed us to quantify centromeric repeats like SAT1.715 and SAT1.709. Notably, genomes almost exclusively contained either a 1,400 or 1,401 bp repeat unit of SAT1.1715, with indicine cattle genomes predominantly carrying the 1,400 bp variant. In higher quality assemblies, we found around 300 Mb of centromeric sequence (~10% of the genome size), aligning with expectations for Artiodactyla and explaining the difference in expected genome size for the cattle reference genome (which largely lacks centromeres). We built multiple pangenomes with the 77 unique breed cattle genomes, including a gene-focused pangene graph as well as both pggb and minigraph-cactus base-level variation graphs. By deconstructing the graphs into reference-coordinate VCFs, we characterized structural variants (SVs) across subspecies and breed purposes. We also explored the feasibility of aligning short and long sequencing reads to the minigraph-cactus pangenome, both using a filtered graph approach as well as a personalised pangenome approach. While we observed an improvement in alignment for non-reference reads, the overall complexity and necessary computational resources remain significant challenges to overcome before adopting a pangenomic bovine reference for routine applications.

Keywords: 2026

How to Cite:

Leonard, A., van Elst, T., Low, L., Rosen, B., Smith, T. & Pausch, H., (2026) “Improving representation of global cattle diversity with pangenomes built by the Bovine Pangenome Consortium”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2251329. doi: https://doi.org/10.31274/wcgalp.23384

Rights: 1

Downloads:
Download PDF
View PDF

111 Views

28 Downloads

Published on
2026-02-25

Peer Reviewed