Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Machine learning based identification of key production drivers of sheep population in Türkiye: a century-long analysis with multiple imputation techniques.

Created on 24 Jul 2026

Authors

Malik Ergin, Merve Mürüvvet Dağ, Bektaş Kadakoğlu

Published in

Tropical animal health and production. Volume 58. Issue 7. Jul 24, 2026. Epub Jul 24, 2026.

Abstract

For the first time, this study employs a century-long dataset (1925-2024) to reveal key factors that would be in relationship with sheep population (NSheep) in Türkiye using state-of-the-art machine learning algorithms. Due to the existence of missing values in the original dataset, missing observations were addressed through four imputation techniques-Next Observation Carried Backward (NOCB), Mean, MIDASpy, and Random Forest (RF)-generating four distinct datasets for comparative analysis. For revealing the key production factors related with NSheep, Extreme Gradient Boosting (XGB) and Multilayer Perceptron (MLP) algorithms were modeled via 5-fold cross-validation and multiple performance metrics (R², MSE, RMSE, MAE, and MdAPE). MLP produced lower prediction errors than XGB across all imputation techniques, though this difference was statistically confirmed only under NOCB and RF imputation (Diebold-Mariano test, P < 0.01 and P < 0.05, respectively); differences under MEAN and MIDASpy imputation were not significant. The highest overall accuracy was achieved by MLP with NOCB imputation (R² = 0.975), while XGB with RF imputation showed the weakest fit (R² = 0.917). Feature importance analyses consistently identified cattle population (NBovine) as the dominant variable associated with NSheep across all four imputation techniques and both algorithms, followed by meadow and pasture area (M&PH) for XGBoost and a more evenly distributed set of variables (M&PH, sheep meat production, goat population) for MLP. Given that NBovine and NSheep both increased steadily over the study period, this association is interpreted as reflecting shared structural growth among livestock subsectors rather than a causal effect of cattle population on sheep numbers. These dataset-specific, associative findings support the value of combining multiple imputation strategies with flexible machine learning algorithms to characterize structural interdependencies in long-term agricultural production data. Future studies incorporating chronologically ordered validation schemes and explicitly modeling structural breaks and policy shifts could further clarify the robustness of these associations, and extending this approach to other livestock species and regions would help establish their broader generalizability.

PMID:
42496899
Bibliographic data and abstract were imported from PubMed on 24 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 7
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement