Klasifikasi Compliance dan Clustering Risk Profiling Supplier pada Data Procurement

Authors

  • Wahyu Rahman Hakim Universitas Muhammadiyah Bengkulu
  • Fitriah Fitriah Universitas Muhammadiyah Bengkulu

DOI:

https://doi.org/10.62951/bridge.v4i3.967

Keywords:

Classification, Clustering, K-Means, Missing Data, Random Forest

Abstract

The cost efficiency and supply chain quality of a company are strongly influenced by how well purchase orders comply with established policies and how consistently supplier performance is maintained. This study builds a classification model to predict the Compliance  status of a purchase order, while also grouping suppliers into risk profiles through Clustering. The data source used is the Procurement KPI Analysis Dataset from Kaggle, covering 777 rows of purchase order data from 5 suppliers. Prior to modelling, an investigation into missing value patterns pointed toward a Missing Not At Random (MNAR) mechanism on the Defective_Units attribute, so an imputation strategy accompanied by an indicator feature (missing flag) was applied instead of row deletion. Three classification algorithms were tested   Logistic Regression, Random Forest, and Gradient Boosting   while supplier risk grouping used the K-Means algorithm. Random Forest emerged as the best-performing model, recording an F1-Score of 0.899 and a Recall of 0.977, with Defect_Rate as the most decisive feature. On the Clustering side, two of the five suppliers fell into the Higher Risk category. These findings confirm that combining classification and Clustering, coupled with careful examination of Missing Data, produces a more complete understanding to support data-driven Procurement decision-making.

Downloads

Download data is not yet available.

References

Abdulla, A., Baryannis, G., & Badi, I. (2023). An integrated machine learning and MARCOS method for supplier evaluation and selection. Decision Analytics Journal, 9(October), 100342. https://doi.org/10.1016/j.dajour.2023.100342

Alam, S., Sohaib, M., Arora, S., & Asad, M. (2023). An investigation of the imputation techniques for missing values in ordinal data enhancing Clustering and classification analysis validity. Decision Analytics Journal, 9(October), 100341. https://doi.org/10.1016/j.dajour.2023.100341

Ali, R., Ashiquzzaman, S., & Ahmed, S. (2023). A decision support system for classifying supplier selection criteria using machine learning and random forest approach. Decision Analytics Journal, 7(March), 100238. https://doi.org/10.1016/j.dajour.2023.100238

Aljohani, A. (2023). Predictive Analytics and Machine Learning for Real-Time Supply Chain Risk Mitigation and Agility.

Barus, O. P., Nathasya, C., & Pangaribuan, J. J. (2023). Mathematical Modelling of Engineering Problems The Implementation of RFM Analysis to Customer Profiling Using K-Means Clustering. 10(1), 298–303.

Chen, J., & Wang, P. (2023). Deep Generative Imputation Model for Missing Not At Random Data. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23), October 21â•fi25, 2023, Birmingham, United Kingdom (Vol. 1, Issue 1). Association for Computing Machinery. https://doi.org/10.1145/3583780.3614835

Dube, L., & Verster, T. (2023). Enhancing classification performance in imbalanced datasets : A comparative analysis of machine learning models. 3(September), 354–379.

Hairani, H., Anggrawan, A., & Priyanto, D. (2023). Improvement Performance of the Random Forest Method on Unbalanced Diabetes Data Classification Using Smote-Tomek Link. 7(March), 258–264.

Hanisah, N., Malek, A., Fairos, W., Yaacob, W., Wah, Y. B., Nasir, S. A., Shaadan, N., & Indratno, S. W. (2023). Comparison of ensemble hybrid sampling with bagging and boosting machine learning approach for imbalanced data. 29(1), 598–608. https://doi.org/10.11591/ijeecs.v29.i1.pp598-608

Kaope, C., & Pristyanto, Y. (2023). The Effect of Class Imbalance Handling on Datasets Toward Classification Algorithm Performance. 22(2), 227–238. https://doi.org/10.30812/matrik.v22i2.2515.

Lee, K. J., Carlin, J. B., Simpson, J. A., Moreno-betancur, M., & Moreno-betancur, M. (2023). classification.

Ness, M. Van, Bosschieter, T. M., Halpin-gregorio, R., & Udell, M. (2023). The Missing Indicator Method : From Low to High Dimensions. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’23), August 6â•fi10, 2023, Long Beach, CA, USA (Vol. 1, Issue 1). Association for Computing Machinery. https://doi.org/10.1145/3580305.3599911.

Rahiminia, M., Razmi, J., Farahani, S. S., & Sabbaghnia, A. (2026). Cluster-based supplier segmentation : a sustainable data-driven approach. 5(3), 209–228. https://doi.org/10.1108/MSCRA-05-2023-0017.

Sisk, R., Sperrin, M., Peek, N., Smeden, M. Van, & Martin, G. P. (2023). Imputation and missing indicators for handling Missing Data in the development and deployment of clinical prediction models : A simulation study. 32(8), 1461–1477. https://doi.org/10.1177/09622802231165001

Wongvorachan, T., & He, S. (2023). A Comparison of Undersampling , Oversampling , and SMOTE Methods for Dealing with Imbalanced Classification in Educational Data Mining.

Zheng, G., Kong, L., & Brintrup, A. (2023). Federated machine learning for privacy preserving , collective supply chain risk prediction. 7543. https://doi.org/10.1080/00207543.2022.2164628

Downloads

Published

2026-08-05

How to Cite

Wahyu Rahman Hakim, & Fitriah Fitriah. (2026). Klasifikasi Compliance dan Clustering Risk Profiling Supplier pada Data Procurement. Bridge : Jurnal Publikasi Sistem Informasi Dan Telekomunikasi, 4(3), 25–34. https://doi.org/10.62951/bridge.v4i3.967

Similar Articles

<< < 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.