Entessar Al-Jbawi*(1), Rasha Dannoura(²), and Abdullah Yacoub(³)
(1). Sugar Beet Research Department, Crops Research Administration, General Commission for Scientific Agricultural Research (GCSAR), Damascus, Syria.
(2). (GCSAR), Damascus, Syria.
(3). Faculty of Agriculture, Department of Rural Engineering, Damascus University, Damascus, Syria.
(*Corresponding author: dr.entessara@gmail.com or dr.entessara@gcsar.gov.sy).
Received: 01/05/2026 Accepted: 15/06/2026
Abstract
Machine learning (ML) techniques have become promising tools for predicting crop productivity due to their ability to analyze complex relationships between agronomic traits and yield performance and to improve prediction accuracy. However, studies comparing different machine learning models for predicting quinoa (Chenopodium quinoa Willd.) grain yield under deficit irrigation and organic fertilization conditions remain limited, particularly when dealing with small experimental datasets. This study aimed to compare the performance of four machine learning models, namely Linear Regression (LR), Random Forest (RF), Support Vector Regression (SVR), and Extreme Gradient Boosting (XGBoost), for predicting quinoa grain yield based on agronomic traits and yield components. The study was based on data obtained from a factorial field experiment consisting of 18 observations. Prior to model development, exploratory data analysis, correlation analysis, and feature selection using the Random Forest model were performed to identify the most influential variables affecting grain yield prediction. Model performance was evaluated using Leave-One-Out Cross-Validation (LOOCV), based on the coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE). In addition, the Wilcoxon Signed-Rank Test was applied to determine the statistical significance of differences among prediction errors of the evaluated models. Feature importance analysis revealed that grain weight per plant was the most influential factor for predicting grain yield, followed by organic fertilization, dry matter production, harvest index, and plant density. The results showed that the Linear Regression model achieved the best predictive performance among all evaluated models, with a coefficient of determination of: R2=0.9968, while the values of root mean square error and mean absolute error were: RMSE=0.0176, MAE=0.0142. The Linear Regression model clearly outperformed the XGBoost, Random Forest, and Support Vector Regression models. Furthermore, the Wilcoxon Signed-Rank Test confirmed that the differences in prediction errors between the Linear Regression model and the other models were statistically significant (P < 0.05). These findings indicate that simple statistical models may outperform more complex machine learning approaches when analyzing small experimental datasets characterized by strong linear relationships among variables. The study also provides an integrated analytical framework combining feature selection, cross-validation, and statistical comparison of machine learning models, contributing to improved crop yield prediction accuracy and offering potential applications in plant breeding programs and similar agricultural experiments.
Key words: Quinoa; Machine learning; Grain yield prediction; Linear Regression; Random Forest; Support Vector Regression; XGBoost; Feature selection; Leave-One-Out Cross-Validation; Deficit irrigation; Organic fertilization.
Download full PDF
How to cite this article – APA Style
Al-Jbawi, E., Dannoura, R., & Yacoub, A. (2026). Prediction of quinoa grain yield using machine learning models based on growth traits and yield components under deficit irrigation and organic fertilization. Research Journal of Science (RJS), 4(1), 1–32.
