A novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum
| dc.contributor.author | Iman Fatima | |
| dc.contributor.author | Maria Rasheed | |
| dc.contributor.author | Darain Fatima | |
| dc.contributor.author | Muhammad Maaz Hamid | |
| dc.contributor.author | Ines Hilali Jaghdam | |
| dc.contributor.author | Ammar T. Zahary | |
| dc.date.accessioned | 2026-07-03T21:02:52Z | |
| dc.date.issued | 2026-04-06 | |
| dc.description.abstract | Effort estimation remains a persistent challenge in agile software development, particularly at the user story level where iterative delivery is emphasized. Inaccurate estimations can result in misallocated resources, budget overruns, and project delays. This research introduces a novel stacking-based ensemble Machine Learning (ML) model designed to enhance the accuracy of user story effort estimation. To address the scarcity of granular data, a high-fidelity dataset comprising 160 user stories from 36 professional scrum-based projects was developed, incorporating 13 industry-validated effort drivers. Adopting the Design Science Research (DSR) methodology, a two-tier hierarchical ensemble model was implemented. In this architecture, Extra Trees, XGBoost, and Random Forest function as Tier-1 base learners, while Linear Regression is employed as the Tier-2 meta-learner to synthesize predictive outputs and optimize the final estimate. Furthermore, a functional web-based prototype interface was developed, operationalizing the model for real-time decision support. Model’s performance was assessed using standard regression metrics: Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE). The proposed stacking ensemble model outperformed all standalone individual models, achieving MAE = 0.51, MSE = 0.49, and RMSE = 0.70. These results demonstrate the model’s effectiveness in improving estimation precision, thereby facilitating superior sprint planning and resource optimization within Scrum teams. Future research will explore the integration of unstructured textual data through Natural Language Processing (NLP) techniques and the adoption of deep learning architectures to further expand the model’s predictive capabilities. | |
| dc.identifier.citation | Fatima, I., Rasheed, M., Fatima, D., Hamid, M., Jaghdam, I. H., & Zahary, A. T. (2026). A novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum. Journal of Big Data, 13, 77. https://doi.org/10.1186/s40537-026-01414-8 | |
| dc.identifier.doi | 10.1186/s40537-026-01414-8 | |
| dc.identifier.issn | 2196-1115 | |
| dc.identifier.openalex | W7150964467 | |
| dc.identifier.other | https://doi.org/10.1186/s40537-026-01414-8 | |
| dc.identifier.uri | https://link.springer.com/article/10.1186/s40537-026-01414-8 | |
| dc.identifier.uri | https://repository.ust.edu.ye/handle/123456789/2617 | |
| dc.language.iso | en | |
| dc.publisher | Springer Nature | |
| dc.subject | Computer science | |
| dc.subject | Machine learning | |
| dc.subject | Artificial intelligence | |
| dc.subject | Scrum | |
| dc.subject | Ensemble learning | |
| dc.subject | Ensemble forecasting | |
| dc.subject | Computational Science and Engineering | |
| dc.subject | Estimation | |
| dc.subject | Data mining | |
| dc.subject | Software | |
| dc.subject | Scrum | |
| dc.subject | Ensemble learning | |
| dc.subject | Ensemble forecasting | |
| dc.subject | Computational Science and Engineering | |
| dc.subject | Estimation | |
| dc.title | A novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum | |
| dc.type | Article | |
| oaire.citation.issue | 1 | |
| oaire.citation.title | Journal Of Big Data | |
| oaire.citation.volume | 13 | |
| oaire.version | http://purl.org/coar/version/c_970fb48d4fbd8a85 |
ملفات
حزمة الترخيص
1 - 1 من 1
جاري التحميل...
- الاسم:
- license.txt
- الحجم:
- 1.71 KB
- تنسيق:
- Item-specific license agreed to upon submission
- الوصف: