Performance of The XGBoost Method in Measuring Uncertainty of Small Area Estimation
Date
2026Author
Amatullah, Fida Fariha
Fitrianto, Anwar
Notodiputro, Khairil Anwar
Metadata
Show full item recordAbstract
Sample surveys are one of the primary sources for producing socioeconomic indicators, including average per capita expenditure. However, at small-area levels such as sub-districts, limited sample sizes often lead to direct estimates with large variances and low precision. Small Area Estimation (SAE) provides an effective solution by incorporating auxiliary information to improve estimation accuracy. Among SAE approaches, the Fay–Herriot (FH) model is the most widely used; nevertheless, it relies on assumptions of linearity, normality, and independence of model components. In contrast, recent advances in machine learning have introduced methods such as Extreme Gradient Boosting (XGBoost), which can capture non-linear relationships and complex interactions without requiring strict parametric assumptions. This study aims to evaluate the performance of XGBoost in estimating average per capita expenditure at the sub-district level in Jambi Province and to assess the associated uncertainty.
The study used data from the 2021 National Socioeconomic Survey (SUSENAS) and the 2021 Village Potential Statistics (PODES) of Jambi Province. Prior to modeling, household expenditure data were cleaned by removing non-recurring expenditure components, followed by outlier treatment using Winsorization at the 5th and 95th percentiles. Exploratory analysis revealed a skewed distribution of direct estimates and substantial variation in sample sizes across sub-districts, supporting the application of SAE methods. For the EBLUP-FH model, a logarithmic transformation was applied to the response variable to improve its distributional characteristics and satisfy model assumptions. The results showed that, after transformation, both the area random effect and the sampling error became closer to normality, with p-values of 0.5263 and 0.05009, respectively. However, the independence assumption was not fully satisfied, as a relatively strong correlation of 0.8457 was observed between the two model components.
The evaluation results indicate that the EBLUP-FH model achieved an MSE of 1.81 × 10?, which increased to 2.14 × 10? after bootstrap resampling, while the corresponding MAPE values were 1.89% and 2.93%, respectively. These low MAPE values indicate a high degree of agreement between the EBLUP-FH estimates and the direct estimates. Furthermore, most shrinkage factor values ranged between 0.80 and 0.95, indicating that EBLUP-FH effectively employed the borrowing-strength mechanism to improve estimate stability in sub-districts with small sample sizes. Meanwhile, XGBoost demonstrated strong capability in capturing complex relationships among variables, although its bootstrap-based uncertainty measures were generally larger than those of EBLUP-FH.
A comparison of standard errors showed that EBLUP-FH generally produced more stable estimates than direct estimation. XGBoost provided competitive results and greater flexibility in handling complex data structures; however, its bootstrap results indicated relatively higher uncertainty. These findings suggest that, for estimating per capita expenditure at the sub-district level in Jambi Province, EBLUP-FH remains a highly effective approach, as it effectively utilizes auxiliary information through the borrowing-strength mechanism, resulting in precise and stable estimates. Nevertheless, the violation of the independence assumption indicates that machine learning approaches such as XGBoost remain promising alternatives or complementary methods for SAE applications involving complex socioeconomic data.
The findings of this study highlight the importance of SAE methods in improving the quality of per capita expenditure indicators for areas with limited sample sizes, thereby supporting more targeted, evidence-based policymaking. Moreover, the results demonstrate that integrating classical statistical methods with machine learning approaches offers significant potential for developing more flexible SAE frameworks capable of addressing the increasing complexity and heterogeneity of modern socioeconomic data.

