Analisis Komparatif Metode Machine Learning dalam Klasifikasi Risiko Penyakit Jantung Berbasis Data Kaggle

Authors

  • Adisuputra Adisuputra Universitas Pertiba
  • Aditya Ahmad Fauzi Universitas Pertiba
  • Ditra Liandaputra Universitas Pertiba
  • Fitriyanti Fitriyanti Universitas Pertiba
  • Tri Dewi Yuni Utami Universitas Pertiba

DOI:

https://doi.org/10.62951/switch.v4i4.952

Keywords:

Accuracy, Binary Classification, Heart Disease, Naive Bayes, Support Vector Machine

Abstract

Heart disease remains one of the leading causes of the highest mortality rates worldwide, requiring a fast and accurate early detection system to minimize the risk of fatality. This study aims to test, compare, and analyze the performance of two popular machine learning methods, Support Vector Machine (SVM) and Naive Bayes, in classifying the risk of heart attacks. The research method applied is quantitative experimental, utilizing structured secondary data from the Kaggle repository, which includes 79,583 patient medical records. The data preprocessing stages involve handling missing values, feature normalization using the MinMax Scaler technique, and dataset splitting with a proportion of 80% training data and 20% testing data. The research findings indicate that the SVM architecture significantly dominates global performance, achieving an accuracy rate of 0.9847, a precision of 0.9594, and an F1-score of 0.8745. On the other hand, the Naive Bayes algorithm records the highest sensitivity (recall) value of 0.9017, compared to SVM, which only reaches 0.8034. The implications of this study confirm that although SVM is highly superior in the aggregate and accurate in suppressing false-positive rates, Naive Bayes demonstrates better characteristics for initial screening scenarios due to its high sensitivity in minimizing the risk of undetected critical patients.

Downloads

Download data is not yet available.

References

Brown, A., & Green, L. (2025). Comparative analysis of supervised learning algorithms in clinical prediction. Journal of Medical Systems, 49(2), 112–125. https://doi.org/10.1007/s10916-025-0214-x

Fawcett, T. (2021). An introduction to ROC analysis and confusion matrix evaluation. Pattern Recognition Letters, 27(8), 861–874. https://doi.org/10.1016/j.patrec.2021.04.012

Han, J., Kamber, M., & Jian, P. (2022). Data mining: Concepts and techniques (4th ed.). Morgan Kaufmann.

Johnson, M., & Martinez, R. (2023). Information technology in modern healthcare applications. IEEE Transactions on Biomedical Engineering, 70(4), 1045–1056. https://doi.org/10.1109/TBME.2023.3214567

Kaggle. (2023). Heart disease dataset: Clinical features for cardiovascular prediction. Kaggle Repository. https://www.kaggle.com/datasets/heart-disease-cleveland

Kurniawan, D., & Wijaya, A. (2023). Klasifikasi risiko penyakit kardiovaskular menggunakan pendekatan data mining. Jurnal Ilmu Komputer dan Informatika, 9(1), 45–56. https://doi.org/10.30888/jiki.v9i1.345

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Pradhan, R., & Kumar, S. (2022). Machine learning models for early prediction of chronic diseases. Computer Methods and Programs in Biomedicine, 214, 106–118. https://doi.org/10.1016/j.cmpb.2022.106543

Rahman, F., Rossi, A., & Sitorus, M. (2023). Performance variability of classification algorithms on medical tabular data. ACM Computing Surveys, 55(3), 1–22. https://doi.org/10.1145/357890

Russell, S., & Norvig, P. (2020). Artificial intelligence: A modern approach (4th ed.). Pearson.

Smith, J., & Jones, M. (2025). Optimization of hyperplane boundaries and probabilistic models in clinical data. International Journal of Computer Science Issues, 22(1), 78–89. https://doi.org/10.30888/ijcsi.v22i1.890

Sugiyono. (2021). Metode penelitian kuantitatif, kualitatif, dan R&D. Alfabeta.

Tan, E., & Setiawan, B. (2024). Deteksi dini penyakit jantung menggunakan optimasi hyperparameter machine learning. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 8(2), 210–218. https://doi.org/10.29207/resti.v8i2.5432

Vapnik, V. (2022). The nature of statistical learning theory (2nd ed.). Springer Science & Business Media.

World Health Organization. (2024). Cardiovascular diseases (CVDs) fact sheets. https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds)

Zhang, H. (2021). The optimality of Naive Bayes in high-dimensional feature spaces. Journal of Machine Learning Foundations, 15(2), 134–149. https://doi.org/10.1561/2200000015

Zhao, Q., Liu, Y., & Wang, X. (2024). Impact of min-max normalization on kernel functions in support vector machines. Pattern Recognition, 146, 109–121. https://doi.org/10.1016/j.patcog.2023.109921

Downloads

Published

2026-07-23

How to Cite

Adisuputra Adisuputra, Aditya Ahmad Fauzi, Ditra Liandaputra, Fitriyanti Fitriyanti, & Tri Dewi Yuni Utami. (2026). Analisis Komparatif Metode Machine Learning dalam Klasifikasi Risiko Penyakit Jantung Berbasis Data Kaggle. Switch : Jurnal Sains Dan Teknologi Informasi, 4(4), 200–206. https://doi.org/10.62951/switch.v4i4.952

Similar Articles

1 2 3 4 > >> 

You may also start an advanced similarity search for this article.