Skip to main content
Microbiome

Host genomic sequence recovery from metagenomic sequence data generated from rumen fluid collected by esophageal tubing

Authors
  • Reid Anema (South Dakota State University)
  • Samodha Fernando (University of Nebraska–Lincoln)
  • Michael Gonda (South Dakota State University)
  • Andrew Lakamp (University of Nebraska–Lincoln)
  • Warren Snelling (USDA-ARS)
  • Matthew Spangler (University of Nebraska–Lincoln)

Abstract

Investigations of host genetic control of rumen microbiome characteristics rely on host genomic and rumen metagenomic data being generated from two separate samples. This not only adds cost but also introduces the possibility of error in appropriately pairing genomic and metagenomic data. Esophageal tubing allows for the collection of both the microbiome and host cells in cattle, even if the latter was unintentional. Metagenomic sequencing often generates host DNA sequences which are discarded due to low resolution and because they are not the target of the research. To investigate the potential to recover host genomic data from rumen metagenomic studies, rumen samples from beef cattle (n=717) were collected through esophageal tubing and shotgun sequenced at an average depth of 20 million reads (20M). The reads were randomly subsampled to 0.5x (10M) and 0.1x (2M) coverages to simulate low-coverage sequencing of the rumen metagenome. Host reads that mapped to the ARS 2.0 Bos taurus genome were selected for further analysis. Average depths of host sequence data were 1.01x10-1, 5.30x10-2, and 1.10x10-2 for the 20M, 10M, and 2M groups, respectively. Average coverages of the host genome were low and showed a disproportionately greater decrease than would be expected given the downsampling proportions; likely a consequence of low host read abundance in the original 20M set. Only the full set had samples with > 1x coverage of the host genome (n=6). The number of samples with > 0.1x coverage of the host genome were 159, 98, and 8 for 20M, 10M, and 2M, respectively. The average number of variants were 3.39x105, 1.99x105, and 4.72x104 for 20M, 10M, and 2M, respectively. The number of samples with 0 indels increased as metagenomic reads decreased; 0 and 22 for 20M and 2M, respectively. When samples with no indels were removed to avoid inflation of values, SNP:Indel ratios displayed an inverse relationship to number of metagenomic reads. Quality of variant calls were assessed by the transition to transversion (Ti/Tv) ratio with ratios of ~0.5 indicative of false positives or artifacts and more acceptable SNP quality reflected by ratios approaching higher values, i.e. ~2.0. Ti/Tv ratios were similar among downsampled sets, with average values of 1.79, 1.78, and 1.76 for 20M, 10M, and 2M, respectively. The low average and large range in host genomic coverage recovered from the metagenomic sequence data likely reflects the low number of host cells captured via esophageal tubing. Although the coverage was extremely low for the majority of the samples, imputation of genotypes, and comparison of imputed genotypes to existing genotypes of the animals, would help determine the effectiveness of recovering host genomic data from microbiome samples taken via esophageal tubing. The USDA is an equal opportunity provider and employer.

Keywords: 2026

How to Cite:

Anema, R., Fernando, S., Gonda, M., Lakamp, A., Snelling, W. & Spangler, M., (2026) “Host genomic sequence recovery from metagenomic sequence data generated from rumen fluid collected by esophageal tubing”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2283202. doi: https://doi.org/10.31274/wcgalp.23492

Rights: 1

Downloads:
Download PDF
View PDF

51 Views

12 Downloads

Published on
2026-02-26

Peer Reviewed