Skip to main content
Multi-omics analysis

Impact of metagenomic sequencing depth on prediction accuracy

Authors
  • Andrew Lakamp (University of Nebraska–Lincoln)
  • Nirosh Aluthge (University of Nebraska–Lincoln)
  • Larry Kuehn (USDA-ARS)
  • Warren Snelling (USDA)
  • James Wells (USDA-ARS)
  • Kristin Hales (USDA-ARS)
  • Bryan Neville (USDA-ARS)
  • Samodha Fernando (University of Nebraska–Lincoln)
  • Matthew Spangler (University of Nebraska–Lincoln)

Abstract

Metagenomic information can aid in both genomic and phenotypic predictions of economically relevant traits. Financial restraints often result in a trade-off between the number of samples sequenced and sequencing depth. Therefore, it is critical to understand how changes in sequencing depth lead to changes in prediction accuracy to make optimal use of resources. This study utilized the host genomic and rumen metagenomic information of 717 beef cattle (417 steers, 300 heifers) to make phenotypic predictions for average daily dry matter intake (ADDMI) and average daily gain (ADG). Steers were fed one of two concentrate-based diets, while heifers were fed one of two forage-based diets. Metagenomic samples were sequenced at an average depth of 20 million reads (20M set) and downsampled to 50% (10M set), 25% (5M set), and 10% (2M set) of the original reads. Rumen microbial open reading frames (ORF) were predicted from each set of reads and used to define a metagenomic (co)variance matrix. Variance components were estimated for each model using all available data. Two cross-validation schemes were utilized to determine prediction accuracy: leave-one-diet-out (LODO) and random 4-fold validation (4F). Models which incorporated host genomic and metagenomic data explained more variation and generally had greater prediction accuracies than models with only a single random effect. For ADDMI, the microbiability estimates were slightly numerically different (lowest for 2M and increasing to 20M), but not statistically different when considering standard errors. Microbiability estimates of ADG were higher for the 10M and 20M set, compared to the 2M and 5M set, especially for models which contained a host genomic, a metagenomic, and an interaction term. Spearman correlations of metagenomic effect solutions, termed the estimated metagenomic value (EMV), between the four sets for all models was almost always >0.90. The 5M, 10M, and 20M sets were almost always more correlated with each other than with the 2M set. Prediction accuracy generally increased as sequencing depth increased for ADDMI and ADG. The magnitude of these differences was dependent on cross-validation scheme and complexity of the model. As model complexity increased, i.e., when host genomics and an interaction term were added, prediction accuracy generally increased. Thus, while metagenomic predictions depended on sequencing depth, trait, and model complexity, using data from an average sequencing depth of 5 or 10 million reads per sample yielded most of the benefits of sequencing at a depth of 20 million reads.The USDA is an equal opportunity provider and employer.

Keywords: 2026

How to Cite:

Lakamp, A., Aluthge, N., Kuehn, L., Snelling, W., Wells, J., Hales, K., Neville, B., Fernando, S. & Spangler, M., (2026) “Impact of metagenomic sequencing depth on prediction accuracy”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2283040. doi: https://doi.org/10.31274/wcgalp.23476

Rights: 1

Downloads:
Download PDF
View PDF

55 Views

14 Downloads

Published on
2026-02-26

Peer Reviewed