Evaluating Ensemble and Single-Model Methods for Reconstructing Incomplete Lactation Curves
Abstract
In dairy production systems, the accurate assessment of a cow's lactation curve is fundamental for evaluating performance, health, and management outcomes. However, daily milk yield (DMY) records often contain missing data points due to technical limitations, associated costs, recording errors, or interruptions in data collection, such as equipment malfunctions, health events, or changes in farm personnel. While many linear and nonlinear models exist, with varying parameterizations and capacity to capture peak, time-to-peak, and persistency, selecting a single model that generalizes across cows remains challenging. In this study, we evaluated an ensemble framework built from up to 47 published lactation‑curve models and directly compared its performance to each individual model, including MilkBot, a widely used benchmark for lactation modeling, to identify optimal imputation strategies for reconstructing lactation curves. Data from 2017 to 2021 was obtained from a U.S. commercial dairy farm using automated milking systems. After quality control (3.5 SD outlier removal) and retaining cows with ≥290 DMY observations per lactation, the final dataset comprised 850,940 records from 2,732 Holstein cows. To assess imputation accuracy, DMY values were systematically masked, retaining 4 to10 evenly spaced records per lactation to emulate sparse sampling. Performance was quantified by root mean squared error (RMSE, kg/day) between imputed and observed values. The performance of the ensemble models (EM) was compared to the predicted values obtained using the best individual model, selected based on the lowest AIC value for each cow, as well as to predictions from MilkBot function. As expected, model performance improved as the number of records per animal increased. The best single-model RMSE decreased from 8.68 (SD 3.05) with 4 records to 6.66 (SD 1.92) with 10 records, while the MilkBot model RMSE dropped from 9.48 (SD 3.63) to 7.55 (SD 3.47). Notably, the EM consistently achieved lower RMSE values than both benchmarks, declining from 7.76 (SD 2.63) with 4 records to 6.06 (SD 1.52) with 10 records, highlighting its superior performance in reconstructing lactation curves, especially under sparse data conditions. The number of converged models also increased from 25 at 4 records to 40-41 at 6-10 records, reflecting robust model fit across conditions. These results demonstrate that ensemble modeling offers a robust and accurate solution for lactation-curve imputation, particularly when data are sparse, and holds promise for improving genetic evaluations and management decisions in dairy herds. Continued refinement and implementation of ensemble approaches will further enhance the biological relevance and reliability of lactation curve analyses in commercial dairy production systems.
Keywords: 2026
How to Cite:
Oliveira, G., Vukasinovic, N., Liang, D. & Fonseca, P., (2026) “Evaluating Ensemble and Single-Model Methods for Reconstructing Incomplete Lactation Curves”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285667. doi: https://doi.org/10.31274/wcgalp.23772
Rights: 1
Downloads:
Download PDF
View PDF
57 Views
15 Downloads