A Model for Measuring the Effect of Splitting Data Method on the Efficiency of Machine Learning Models: A Comparative Study

dc.contributor.authorShujaaddeen, Abeer Abdullah
dc.contributor.authorBa-Alwi, Fadl Mutaher
dc.contributor.authorZahary, Ammar T.
dc.contributor.authorAlhegami, Ahmed Sultan
dc.date.accessioned2026-06-14T22:56:37Z
dc.date.issued2024
dc.description.abstractThe performance of a classification model in machine learning is affected by many factors, such as the method of splitting the dataset and the type of machine learning technology used. Accuracy varies from method to method. This paper presents a comparison between the performance of the model in terms of the type of data set segmentation method used. And in terms of the machine learning technique used on the one hand. The machine learning techniques used are as follows: k-nearest Neighbors (KNN), Support Vector Machines (SVM), Decision Trees (DT), Random Forest (RF), and MLP. The two data set partition techniques used are k-fold data partitioning methods, the percentage split of the data set provided by the YTA that relates to the commercial and industrial profits tax. The paper shows that using percentage split, The Na�ve Bayes (NB) and Multi-Layer Perceptron (MLP) classifiers gave the same results when the researchers used the percentage split method, and both of them produced the highest result of 100%, but this result will lead to overfitting. Next, the k-Nearest Neighbors (KNN) classifier produced the best results, while Decision Trees (DT) and random forest (RF) produced the same results. SVM, however, produced the worst outcome. The NB gave the highest score when using the K-Fold validation method but, this resulted in overfitting. Next, the MLP classifier produced the best accuracy result, followed by KNN, the (DT) classifier, then the (RF) classifier and SVM, which produced the same result with confidence, and finally, the SVM classifier produced the worst classification. The researchers also sees that in certain ways, training the data with a percentage split yields superior outcomes. However, K-Fold validation produces superior outcomes when used with other methods.en_US
dc.identifier10.1109/eSmarTA62850.2024.10639022
dc.identifier.citationShujaaddeen, A. A., Ba-Alwi, F. M., Zahary, A. T., & Alhegami, A. S. (2024). A model for measuring the effect of splitting data method on the efficiency of machine learning models: A comparative study. In 2024 4th International Conference on Emerging Smart Technologies and Applications (eSmarTA) (pp. 1-13). IEEE. https://doi.org/10.1109/eSmarTA62850.2024.10639022en_US
dc.identifier.urihttps://repository.ust.edu.ye/handle/123456789/401
dc.identifier.urihttps://ieeexplore.ieee.org/document/10639022
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers (IEEE)en_US
dc.titleA Model for Measuring the Effect of Splitting Data Method on the Efficiency of Machine Learning Models: A Comparative Studyen_US
dc.typeConference Paperen_US

ملفات