SCALABLE DETECTION OF COST INEFFICIENCY IN BPJS KESEHATAN CLAIMS USING XGBOOST WITH IMBALANCE LEARNING AND VAEX-BASED BIG DATA PROCESSING

  • Alpian Roymundus Siringo-ringo Departement of Informatics, Faculty of Electrical Engineering, Informatics, and Business, Institut Teknologi Kalimantan, Indonesia https://orcid.org/0009-0004-3061-6049
  • Ramadhan Paninggalih Departement of Informatics, Faculty of Electrical Engineering, Informatics, and Business, Institut Teknologi Kalimantan, Indonesia https://orcid.org/0009-0009-2552-7347
  • Rizal Kusuma Putra Departement of Informatics, Faculty of Electrical Engineering, Informatics, and Business, Institut Teknologi Kalimantan, Indonesia https://orcid.org/0009-0000-5467-4308
Keywords: BPJS kesehatan, Classification, XGBoost, Vaex

Abstract

BPJS Kesehatan Indonesia's national health insurance system, faces significant challenges in managing large-scale claim data, particularly in identifying cost inefficiencies such as abnormal claims, duplication, and misuse. Although machine learning methods have been widely applied in healthcare analytics, limited studies address both large-scale data processing and class imbalance issues in inefficiency detection. This study aims to develop a scalable classification model using XGBoost integrated with Vaex for efficient big data processing. The dataset was obtained from the JKN Healthkathon, consisting of millions of claim records with predefined inefficiency labels. Data preprocessing was performed using Vaex to enable out-of-core computation, followed by model development using XGBoost with various data splitting strategies and imbalance handling techniques, including random oversampling and undersampling. Model performance was evaluated using accuracy, precision, recall, and F1-score. The results indicate that the balanced split with random oversampling achieved the most stable performance, with precision, recall, and F1-score of 0.91. In contrast, stratified splitting without imbalance handling yielded higher accuracy but lower recall. This study contributes by proposing a scalable and efficient framework that integrates XGBoost with Vaex for large-scale healthcare claim analysis and systematically evaluates the impact of imbalance handling strategies on model performance. However, the use of predefined labels with undisclosed generation processes may introduce potential bias in the results.

Downloads

Download data is not yet available.

References

C. F. P. Zebua, D. Ardhila, Y. Yuriska, and F. P. Gurning, “ANALISIS PERAN PROGRAM JAMINAN KESEHATAN NASIONAL (JKN) DALAM MENGURANGI BEBAN FINANSIAL PASIEN: STUDI LITERATURE,” El-Mujtama J. Pengabdi. Masy., vol. 4, no. 2, pp. 802–811, August 2023, doi: https://doi.org/10.47467/elmujtama.v4i2.4407.

R. Annisa, S. Winda, E. Dwisaputro, and K. N. Isnaini, “MENGATASI DEFISIT DANA JAMINAN SOSIAL KESEHATAN MELALUI PERBAIKAN TATA KELOLA,” INTEGRITAS J. Antikorupsi, vol. 6, no. 2, pp. 209–224, December 2020.

BPJS Kesehatan, "PORTAL DATA JAMINAN KESEHATAN NASIONAL," Portal Data BPJS Kesehatan, 2024. [Online].

A. C. Nugraha and M. I. Irawan, “KOMPARASI DETEKSI KECURANGAN PADA DATA KLAIM ASURANSI PELAYANAN KESEHATAN MENGGUNAKAN METODE SUPPORT VECTOR MACHINE (SVM) DAN EXTREME GRADIENT BOOSTING (XGBOOST),” J. Sains dan Seni ITS, vol. 12, no. 1, May 2023, doi: https://doi.org/10.12962/j23373520.v12i1.107032.

G. G. Jerith et al., “EVALUATION OF CATBOOST FOR DIABETES PREVENTION IN COMPARISON TO XGBOOST: TO AVERT MORTALITY BY DEVELOPING AN AI MODEL CAPABLE OF PREDICTING THE ONSET OF DIABETES,” 2024 Int. Conf. Electron. Syst. Intell. Comput., pp. 319–324, Nov. 2024, doi: https://doi.org/10.1109/ICESIC61777.2024.10846686.

A. Mozzillo et al., “EVALUATION OF DATAFRAME LIBRARIES FOR DATA PREPARATION ON A SINGLE MACHINE”, [Online]. doi : https://doi.org/10.1002/9781394155408.ch2

G. A. Shafila, “IMPLEMENTASI METODE EXTREME GRADIENT BOOSTING (XGBOOST ) UNTUK KLASIFIKASI PADA DATA BIOINFORMATIKA (STUDI KASUS : PENYAKIT EBOLA , GSE 122692),” Bachelor [Thesis] Islamic University of Indonesia., 2020 [Online].

I. H. Sarker, “MACHINE LEARNING: ALGORITHMS, REAL-WORLD APPLICATIONS AND RESEARCH DIRECTIONS,” SN Comput. Sci., vol. 2, no. 3, pp. 1–21, March 2021, doi: https://doi.org/10.1007/s42979-021-00592-x.

Annisa, “STRUKTUR DATA TREE,” Fikti Umsu, pp. 1–10, 2023, [Online].

J. Han, M. Kamber, and J. Pei, DATA MINING: CONCEPTS AND TECHNIQUES. 2011. doi: https://doi.org/10.1016/C2009-0-61819-5

N. F. Bakri, ANALISIS KLASIFIKASI FINANCIAL DISTRESS DENGAN MEMBANDINGKAN METODE EXTREME GRADIENT BOOSTING DAN ARTIFICIAL NEURAL NETWORK MENGGUNAKAN RASIO KEUANGAN PADA PERUSAHAAN PERBANKAN TAHUN 2013–2022, Bachelor [Thesis] Hasanuddin University., 2024.

J. H. Friedman, “GREEDY FUNCTION APPROXIMATION: A GRADIENT BOOSTING MACHINE,” Ann. Stat., vol. 29, no. 5, pp. 1189–1232, November 2001, doi: https://doi.org/10.1214/aos/1013203451 .

A. N. Rachmi, “IMPLEMENTASI METODE RANDOM FOREST DAN XGBOOST PADA KLASIFIKASI CUSTOMER CHURN,” pp. 1–101,. Bachelor [Thesis] Islamic University of Indonesia, 2020.

T. Chen and C. Guestrin, “XGBOOST: A SCALABLE TREE BOOSTING SYSTEM,” Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., vol. 13-17-Augu, pp. 785–794, 2016, doi: https://doi.org/10.1145/2939672.2939785

S. E. Herni Yulianti, Oni Soesanto, and Yuana Sukmawaty, “PENERAPAN METODE EXTREME GRADIENT BOOSTING (XGBOOST) PADA KLASIFIKASI NASABAH KARTU KREDIT,” J. Math. Theory Appl., vol. 4, no. 1, pp. 21–26, 2022, doi: https://doi.org/10.31605/jomta.v4i1.1792.

L. Zhang and C. Zhan, MACHINE LEARNING IN ROCK FACIES CLASSIFICATION: AN APPLICATION OF XGBOOST. 2017. doi: https://doi.org/10.1190/IGC2017-351.

A. Mozzillo, “MAXIMIZING EFFICIENCY IN EXISTING DATA PREPARATION PIPELINES,” CEUR Workshop Proc., vol. 3478, pp. 720–726, 2023.

M. A. Breddels and J. Veljanoski, “VAEX: BIG DATA EXPLORATION IN THE ERA OF GAIA,” Astron. Astrophys., vol. 618, pp. 1–13, October 2018, doi: https://doi.org/10.1051/0004-6361/201732493.

R. Schaer, H. Müller, and A. Depeursinge, “OPTIMIZED DISTRIBUTED HYPERPARAMETER SEARCH AND SIMULATION FOR LUNG TEXTURE CLASSIFICATION IN CT USING HADOOP,” J. Imaging, vol. 2, no. 2, 2016, doi: https://doi.org/10.3390/jimaging2020019.

M. M. RAMADHAN, I. S. SITANGGANG, F. R. NASUTION, and A. GHIFARI, “PARAMETER TUNING IN RANDOM FOREST BASED ON GRID SEARCH METHOD FOR GENDER CLASSIFICATION BASED ON VOICE FREQUENCY,” DEStech Trans. Comput. Sci. Eng., no. cece, 2017, doi: https://doi.org/10.12783/dtcse/cece2017/14611.

A. A. Khan, “BALANCED SPLIT: A NEW TRAIN-TEST DATA SPLITTING STRATEGY FOR IMBALANCED DATASETS” 2022, [Online]. [Accessed: 2 March 2025]

Z. Karimi, Confusion Matrix. 2021, [Online].[Accessed: 29 January 2025]

Published
2026-08-24
How to Cite
[1]
A. R. Siringo-ringo, R. Paninggalih, and R. K. Putra, “SCALABLE DETECTION OF COST INEFFICIENCY IN BPJS KESEHATAN CLAIMS USING XGBOOST WITH IMBALANCE LEARNING AND VAEX-BASED BIG DATA PROCESSING”, BAREKENG: J. Math. & App., vol. 20, no. 4, pp. 3531-3544, Aug. 2026.