DNA Methylation Imputation in Whole-Genome Bisulfite Sequencing
Abstract
DNA methylation is one of the epigenetic mechanisms that regulates gene expression, development, and genome stability. Whole Genome Bisulfite Sequencing (WGBS) provides single-base resolution of methylation status, enabling comprehensive epigenomic analysis, but is challenged by data sparsity and missing values due to technical limitations in bisulfite conversion and sequencing. This study aimed to implement and benchmark methylation imputation strategies to improve data completeness and downstream analysis accuracy in WGBS datasets. Peripheral blood mononuclear cell (PBMC) samples from 28 Holstein cows were subjected to DNA extraction, bisulfite conversion, and sequencing at 30à— coverage using an Illumina NovaSeq platform. Sequencing reads were trimmed, aligned to the ARS-UCD2.0 reference genome, and CpG methylation levels called with Bismark, followed by stringent quality control and conversion to β-values. CpG sites and samples with excessive missing data were excluded. Methylation imputation was performed using a gradient boosting approach to predict methylation levels at missing or low-coverage CpG sites. We used a minimum coverage threshold of 5 reads per CpG (approximately 35% of the total CpGs), randomly sampling CpGs per sample to create consistent train/validation/test splits (60%/15%/25%) based on the smallest per-sample CpG count. During benchmarking, missing/low‑coverage sites were imputed, with additional masking of observed CpGs to estimate error on held‑out data. Across 28 PBMC samples, the algorithm achieved a mean RMSE of 0.13 (±0.004) and a mean accuracy of 0.973 (±0.002) on both validation and test sets, indicating high predictive accuracy and consistency. These results support an effective, scalable pipeline for methylation data imputation, enhancing the reliability and biological interpretability of WGBS-based epigenetic analyses in large-scale studies. Future work will extend benchmarking to additional imputation methods (e.g., smoothing and HMM-based approaches, random forests, deep learning) and diverse tissues/cell types to assess generalizability across genomic contexts and coverage regimes. Furthermore, ensemble weighting schemes that incorporate canonical methylation patterns and biological constraints may further improve robustness.
Keywords: 2026
How to Cite:
Oliveira, G., Simon, L., Freeborn, K., Vandenbogaert, M. & Emam, M., (2026) “DNA Methylation Imputation in Whole-Genome Bisulfite Sequencing”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285808. doi: https://doi.org/10.31274/wcgalp.23800
Rights: 1
Downloads:
Download PDF
View PDF
64 Views
16 Downloads