Evaluasi Kinerja Model Geographical-XGBoost pada Data dengan Tingkat Heterogenitas Spasial dan Pola Linearitas Berbeda
Date
2026Author
Mutmainah, Zamrah
Susetyo, Budi
Rahardiantoro, Septian
Metadata
Show full item recordAbstract
Analisis data spasial sering menghadapi tantangan akibat adanya heterogenitas spasial dan hubungan nonlinear antarpeubah. Metode konvensional seperti Ordinary Least Squares (OLS) dan Geographically Weighted Regression (GWR) memiliki keterbatasan dalam menangkap kedua karakteristik tersebut secara bersamaan. OLS mengasumsikan hubungan yang bersifat global dan homogen, sedangkan GWR hanya mampu mengakomodasi heterogenitas spasial dengan asumsi hubungan linear. Di sisi lain, metode Extreme Gradient Boosting (XGBoost) memiliki kemampuan yang baik dalam memodelkan hubungan nonlinear, namun belum mempertimbangkan variasi spasial. Pengembangan metode Geographically Weighted-XGBoost (GW-XGBoost) telah menggabungkan aspek spasial dan nonlinear, tetapi masih menggunakan pendekatan hibrida yang berpotensi menyebabkan overfitting. Oleh karena itu, penelitian ini mengevaluasi kinerja metode Geographical-XGBoost (G-XGBoost) yang mengintegrasikan pembobotan geografis secara langsung ke dalam algoritma XGBoost.
Penelitian ini bertujuan mengevaluasi performa G-XGBoost dibandingkan dengan OLS, GWR, XGBoost, dan GW-XGBoost pada berbagai tingkat heterogenitas spasial dan pola hubungan linear maupun nonlinear. Evaluasi dilakukan melalui kajian simulasi menggunakan empat skenario yang mengombinasikan hubungan linear dan nonlinear dengan heterogenitas spasial rendah maupun tinggi. Selanjutnya, seluruh metode diterapkan pada data tingkat kemiskinan kabupaten/kota di Pulau Jawa tahun 2024 untuk menentukan model terbaik. Performa model dievaluasi menggunakan koefisien determinasi (R^2), Root Mean Squared Error (RMSE), dan Mean Absolute Error (MAE), sedangkan stabilitas model dinilai berdasarkan selisih performa antara data latih dan data uji.
Hasil kajian simulasi menunjukkan bahwa G-XGBoost memiliki performa yang paling stabil dan konsisten pada seluruh skenario yang diuji. Pada skenario linear, OLS dan GWR memberikan hasil yang kompetitif, namun performanya menurun pada data nonlinear. Sebaliknya, G-XGBoost mampu mempertahankan akurasi prediksi yang tinggi, baik pada hubungan linear maupun nonlinear, serta pada tingkat heterogenitas spasial yang berbeda. Selain itu, G-XGBoost menunjukkan gap antara data latih dan data uji yang paling kecil dibandingkan metode lain, sehingga memiliki kemampuan generalisasi yang lebih baik dan risiko overfitting yang lebih rendah dibandingkan GW-XGBoost.
Pada kajian empiris, G-XGBoost kembali menunjukkan performa terbaik dengan nilai R^2 sebesar 0,8115 dan RMSE sebesar 1,6241, yang merupakan hasil terbaik dibandingkan model pembanding. Analisis Local Variable Importance (LVI) menunjukkan bahwa tingkat kepentingan relatif peubah penjelas dalam proses prediksi tingkat kemiskinan berbeda antar kabupaten/kota di Pulau Jawa. Rata-rata lama sekolah merupakan salah satu peubah yang paling sering memiliki nilai kepentingan peubah (variable importance) tinggi, sedangkan persentase penduduk yang bekerja di sektor pertanian, persentase rumah tangga dengan alas lantai semipermanen, upah minimum, kepemilikan jaminan kesehatan, dan persentase penduduk berpendidikan minimal SMA teridentifikasi sebagai peubah dengan variable importance tinggi pada wilayah tertentu. Hasil penelitian ini menunjukkan bahwa G-XGBoost mampu mengakomodasi hubungan nonlinear dan heterogenitas spasial secara bersamaan, yang dapat menjadi referensi awal dalam mengidentifikasi peubah-peubah yang memiliki tingkat kepentingan relatif tinggi untuk dikaji lebih lanjut sesuai karakteristik masing-masing wilayah. Spatial data analysis often faces challenges arising from spatial heterogeneity and nonlinear relationships among variables. Conventional methods such as Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR) are limited in their ability to capture these two characteristics simultaneously. OLS assumes a global, homogeneous relationship across all observations, whereas GWR accommodates spatial heterogeneity while still relying on the assumption of linear relationships. On the other hand, Extreme Gradient Boosting (XGBoost) effectively models nonlinear relationships but does not explicitly account for spatial variation. The development of Geographically Weighted XGBoost (GW-XGBoost) combines spatial and nonlinear modeling capabilities; however, its sequential hybrid framework may increase the risk of overfitting. Therefore, this study evaluates the performance of Geographical-XGBoost (G-XGBoost), which directly integrates geographical weighting into the XGBoost algorithm.
This study aims to evaluate the performance of G-XGBoost compared with OLS, GWR, XGBoost, and GW-XGBoost across different levels of spatial heterogeneity and both linear and nonlinear relationship patterns. The evaluation was conducted through simulation studies involving four scenarios that combined low and high spatial heterogeneity with linear and nonlinear relationships. Subsequently, all models were applied to district- and city-level poverty data in Java Island for 2024 to determine the best-performing model. Model performance was assessed using the coefficient of determination (R^2) Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE), while model stability was evaluated based on the performance gap between the training and testing datasets.
The simulation results demonstrate that G-XGBoost achieved the most stable and consistent performance across all evaluated scenarios. In the linear scenarios, OLS and GWR produced competitive results; however, their performance deteriorated substantially under nonlinear conditions. In contrast, G-XGBoost consistently maintained high predictive accuracy across both linear and nonlinear relationships and under varying degrees of spatial heterogeneity. Furthermore, G-XGBoost exhibited the smallest gap between training and testing performance among all competing models, indicating superior generalization ability and a lower risk of overfitting than GW-XGBoost.
In the empirical study, G-XGBoost again achieved the best performance, with an R^2 value of 0.8115 and an RMSE of 1.6241, outperforming all competing models. The Local Variable Importance (LVI) analysis revealed that the relative importance of explanatory variables in predicting poverty varied across districts and cities in Java Island. Average years of schooling was among the variables that most frequently exhibited high variable importance values, while the percentage of the population employed in the agricultural sector, the percentage of households with semi-permanent flooring, minimum wage, health insurance coverage, and the percentage of the population with at least a senior high school education were also identified as variables with high variable importance in specific regions. These findings demonstrate that G-XGBoost is capable of simultaneously accommodating nonlinear relationships and spatial heterogeneity while providing an initial reference for identifying explanatory variables with relatively high variable importance for further investigation according to the characteristics of each region.

