Perbandingan Klasifikasi Risiko Biaya Asuransi Kesehatan Menggunakan K-Nearest Neighbor dan Regresi Logistik Multinomial
Abstract
Penetapan premi asuransi kesehatan membutuhkan estimasi risiko biaya medis yang presisi. Penelitian ini bertujuan untuk membandingkan kinerja metode regresi logistik multinomial dan K-nearest neighbor (KNN) dalam mengklasifikasikan tingkat risiko biaya medis nasabah. Variabel prediktor yang dianalisis meliputi usia, indeks massa tubuh, status perokok, dan banyaknya anak. Variabel respons berupa biaya medis kontinu diubah menjadi tiga kategori risiko, yaitu rendah, sedang, dan tinggi, menggunakan persentil ke-33 dan ke-66. Evaluasi model dilakukan melalui teknik 5-fold cross validation. Hasil penelitian menunjukkan bahwa pendekatan non-parametrik KNN dengan parameter K=7 menghasilkan akurasi sebesar 89.13%, presisi 90.41%, recall 89.13%, dan f1-score 88.9%. Hasil-hasil ini lebih tinggi dibandingkan model regresi logistik multinomial yang hanya mencapai akurasi 81.27%. Dengan demikian, metode KNN relatif lebih handal dan stabil dalam mengklasifikasikan data risiko biaya medis yang memiliki variabilitas ekstrem. Accurate estimation of medical cost risk is essential in determining health insurance premium pricing. This study aims to compare the performance of the multinomial logistic regression and the K-nearest neighbor (KNN) methods in classifying policyholders’ medical cost risk levels. The predictor variables considered in this study include age, body mass index, smoking status, and number of children. The continuous medical cost variable was transformed into three risk categories, low, medium, and high, based on the 33rd and 66th percentiles. Model evaluation was conducted using a 5-fold cross-validation technique. The results demonstrate that the non-parametric KNN approach with K=7 achieved an accuracy of 89.13%, precision of 90.41%, recall of 89.13%, and f1-score of 88.9%. These results are outperforming the multinomial logistic regression model which only achieved an accuracy of 81.27%. These findings indicate that the KNN method is relatively more reliable in classifying medical cost risk data characterized by high variability and extreme observations.
Collections
- UF - Actuaria [123]

