Analisis Komparatif Metode Machine Learning dalam Klasifikasi Risiko Penyakit Jantung Berbasis Data Kaggle
DOI:
https://doi.org/10.62951/switch.v4i4.952Keywords:
Accuracy, Binary Classification, Heart Disease, Naive Bayes, Support Vector MachineAbstract
Heart disease remains one of the leading causes of the highest mortality rates worldwide, requiring a fast and accurate early detection system to minimize the risk of fatality. This study aims to test, compare, and analyze the performance of two popular machine learning methods, Support Vector Machine (SVM) and Naive Bayes, in classifying the risk of heart attacks. The research method applied is quantitative experimental, utilizing structured secondary data from the Kaggle repository, which includes 79,583 patient medical records. The data preprocessing stages involve handling missing values, feature normalization using the MinMax Scaler technique, and dataset splitting with a proportion of 80% training data and 20% testing data. The research findings indicate that the SVM architecture significantly dominates global performance, achieving an accuracy rate of 0.9847, a precision of 0.9594, and an F1-score of 0.8745. On the other hand, the Naive Bayes algorithm records the highest sensitivity (recall) value of 0.9017, compared to SVM, which only reaches 0.8034. The implications of this study confirm that although SVM is highly superior in the aggregate and accurate in suppressing false-positive rates, Naive Bayes demonstrates better characteristics for initial screening scenarios due to its high sensitivity in minimizing the risk of undetected critical patients.
Downloads
References
Brown, A., & Green, L. (2025). Comparative analysis of supervised learning algorithms in clinical prediction. Journal of Medical Systems, 49(2), 112–125. https://doi.org/10.1007/s10916-025-0214-x
Fawcett, T. (2021). An introduction to ROC analysis and confusion matrix evaluation. Pattern Recognition Letters, 27(8), 861–874. https://doi.org/10.1016/j.patrec.2021.04.012
Han, J., Kamber, M., & Jian, P. (2022). Data mining: Concepts and techniques (4th ed.). Morgan Kaufmann.
Johnson, M., & Martinez, R. (2023). Information technology in modern healthcare applications. IEEE Transactions on Biomedical Engineering, 70(4), 1045–1056. https://doi.org/10.1109/TBME.2023.3214567
Kaggle. (2023). Heart disease dataset: Clinical features for cardiovascular prediction. Kaggle Repository. https://www.kaggle.com/datasets/heart-disease-cleveland
Kurniawan, D., & Wijaya, A. (2023). Klasifikasi risiko penyakit kardiovaskular menggunakan pendekatan data mining. Jurnal Ilmu Komputer dan Informatika, 9(1), 45–56. https://doi.org/10.30888/jiki.v9i1.345
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Pradhan, R., & Kumar, S. (2022). Machine learning models for early prediction of chronic diseases. Computer Methods and Programs in Biomedicine, 214, 106–118. https://doi.org/10.1016/j.cmpb.2022.106543
Rahman, F., Rossi, A., & Sitorus, M. (2023). Performance variability of classification algorithms on medical tabular data. ACM Computing Surveys, 55(3), 1–22. https://doi.org/10.1145/357890
Russell, S., & Norvig, P. (2020). Artificial intelligence: A modern approach (4th ed.). Pearson.
Smith, J., & Jones, M. (2025). Optimization of hyperplane boundaries and probabilistic models in clinical data. International Journal of Computer Science Issues, 22(1), 78–89. https://doi.org/10.30888/ijcsi.v22i1.890
Sugiyono. (2021). Metode penelitian kuantitatif, kualitatif, dan R&D. Alfabeta.
Tan, E., & Setiawan, B. (2024). Deteksi dini penyakit jantung menggunakan optimasi hyperparameter machine learning. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 8(2), 210–218. https://doi.org/10.29207/resti.v8i2.5432
Vapnik, V. (2022). The nature of statistical learning theory (2nd ed.). Springer Science & Business Media.
World Health Organization. (2024). Cardiovascular diseases (CVDs) fact sheets. https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds)
Zhang, H. (2021). The optimality of Naive Bayes in high-dimensional feature spaces. Journal of Machine Learning Foundations, 15(2), 134–149. https://doi.org/10.1561/2200000015
Zhao, Q., Liu, Y., & Wang, X. (2024). Impact of min-max normalization on kernel functions in support vector machines. Pattern Recognition, 146, 109–121. https://doi.org/10.1016/j.patcog.2023.109921
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Switch : Jurnal Sains dan Teknologi Informasi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.



