A novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum

dc.contributor.authorIman Fatima
dc.contributor.authorMaria Rasheed
dc.contributor.authorDarain Fatima
dc.contributor.authorMuhammad Maaz Hamid
dc.contributor.authorInes Hilali Jaghdam
dc.contributor.authorAmmar T. Zahary
dc.date.accessioned2026-07-03T21:02:52Z
dc.date.issued2026-04-06
dc.description.abstractEffort estimation remains a persistent challenge in agile software development, particularly at the user story level where iterative delivery is emphasized. Inaccurate estimations can result in misallocated resources, budget overruns, and project delays. This research introduces a novel stacking-based ensemble Machine Learning (ML) model designed to enhance the accuracy of user story effort estimation. To address the scarcity of granular data, a high-fidelity dataset comprising 160 user stories from 36 professional scrum-based projects was developed, incorporating 13 industry-validated effort drivers. Adopting the Design Science Research (DSR) methodology, a two-tier hierarchical ensemble model was implemented. In this architecture, Extra Trees, XGBoost, and Random Forest function as Tier-1 base learners, while Linear Regression is employed as the Tier-2 meta-learner to synthesize predictive outputs and optimize the final estimate. Furthermore, a functional web-based prototype interface was developed, operationalizing the model for real-time decision support. Model’s performance was assessed using standard regression metrics: Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE). The proposed stacking ensemble model outperformed all standalone individual models, achieving MAE = 0.51, MSE = 0.49, and RMSE = 0.70. These results demonstrate the model’s effectiveness in improving estimation precision, thereby facilitating superior sprint planning and resource optimization within Scrum teams. Future research will explore the integration of unstructured textual data through Natural Language Processing (NLP) techniques and the adoption of deep learning architectures to further expand the model’s predictive capabilities.
dc.identifier.citationFatima, I., Rasheed, M., Fatima, D., Hamid, M., Jaghdam, I. H., & Zahary, A. T. (2026). A novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum. Journal of Big Data, 13, 77. https://doi.org/10.1186/s40537-026-01414-8
dc.identifier.doi10.1186/s40537-026-01414-8
dc.identifier.issn2196-1115
dc.identifier.openalexW7150964467
dc.identifier.otherhttps://doi.org/10.1186/s40537-026-01414-8
dc.identifier.urihttps://link.springer.com/article/10.1186/s40537-026-01414-8
dc.identifier.urihttps://repository.ust.edu.ye/handle/123456789/2617
dc.language.isoen
dc.publisherSpringer Nature
dc.subjectComputer science
dc.subjectMachine learning
dc.subjectArtificial intelligence
dc.subjectScrum
dc.subjectEnsemble learning
dc.subjectEnsemble forecasting
dc.subjectComputational Science and Engineering
dc.subjectEstimation
dc.subjectData mining
dc.subjectSoftware
dc.subjectScrum
dc.subjectEnsemble learning
dc.subjectEnsemble forecasting
dc.subjectComputational Science and Engineering
dc.subjectEstimation
dc.titleA novel stacking-based ensemble machine learning model for accurate user story effort estimation in scrum
dc.typeArticle
oaire.citation.issue1
oaire.citation.titleJournal Of Big Data
oaire.citation.volume13
oaire.versionhttp://purl.org/coar/version/c_970fb48d4fbd8a85

ملفات

حزمة الترخيص

يظهر الآن 1 - 1 من 1
جاري التحميل...
صورة مصغرة
الاسم:
license.txt
الحجم:
1.71 KB
تنسيق:
Item-specific license agreed to upon submission
الوصف: