IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      Evaluasi Metode Seleksi Eigenvector pada Eigenvector Spatial Filtering dengan Model Boosting untuk Prediksi Konsentrasi PM2.5

      Thumbnail
      View/Open
      Cover (706.6Kb)
      Fulltext (2.802Mb)
      Lampiran (433.7Kb)
      Date
      2026
      Author
      Az-Zahra, Putri Nisrina
      Djuraidah, Anik
      Erfiani
      Metadata
      Show full item record
      Abstract
      Autokorelasi spasial merupakan fenomena ketika pola ketergantungan spasial menunjukkan hubungan yang sistematis antara nilai suatu peubah di suatu lokasi dengan nilai peubah yang sama di lokasi sekitarnya. Keberadaan karakteristik tersebut pada data spasial perlu diperhatikan karena dapat menyebabkan bias dalam model statistik apabila tidak ditangani dengan tepat. Salah satu metode yang efektif dalam menangani permasalahan tersebut adalah Eigenvector Spatial Filtering (ESF). ESF memanfaatkan vektor ciri dari matriks pembobot spasial untuk menangkap pola ketergantungan spasial dalam data dan mengikutsertakan pola tersebut ke dalam model sebagai peubah bebas tambahan sehingga dapat meningkatkan kualitas model prediksi. Pada penerapan ESF, pemilihan vektor ciri yang optimal menjadi salah satu aspek yang perlu diperhatikan. Tahapan ini penting dalam pemodelan untuk menghindari overfitting, meningkatkan interpretabilitas model, dan menurunkan beban komputasi. Meski efektif dalam seleksi parsial, berbagai metode pemilihan vektor ciri yang umum digunakan belum memberikan informasi eksplisit mengenai kontribusi individual masing-masing vektor ciri terhadap prediksi model. Sebagai alternatif, metode SHAP dapat diterapkan untuk melakukan seleksi fitur yang dalam hal ini yaitu vektor ciri, dengan basis yang kuat dalam teori permainan. Vektor ciri dengan nilai SHAP tertinggi dapat diprioritaskan untuk dipilih dalam pemodelan ESF, sementara vektor ciri dengan kontribusi rendah dapat dieliminasi. Metode ESF masih terbatas dalam menangkap hubungan nonlinier maupun interaksi yang kompleks antar peubah. Algoritma pembelajaran mesin telah menjadi pendekatan yang efektif karena kemampuannya dalam menangkap pola yang kompleks dan nonlinier. Oleh karena itu, integrasi ESF dengan algoritma pembelajaran mesin telah diterapkan seiring dengan berkembangnya data geospasial skala besar. Pada penelitian ini, akan dilakukan evaluasi metode pemilihan vektor ciri pada ESF dengan integrasi model boosting pada kasus kualitas udara dengan PM2.5 sebagai indikatornya. Particulate Matter (PM2.5) merupakan partikel halus berdiameter kurang dari 2,5 µm yang digunakan sebagai indikator utama kualitas udara. Ukurannya yang sangat kecil dapat menembus saluran pernafasan terdalam manusia dan menimbulkan gangguan pernapasan, penyakit kardiovaskular, serta peningkatan risiko kematian dini. Provinsi DKI Jakarta sebagai ibu kota Indonesia menghadapi permasalahan polusi udara yang semakin serius akibat tingginya kepadatan penduduk dan intensitas penggunaan transportasi yang tinggi sehingga konsentrasi PM2.5 kerap melampaui ambang batas aman. Penelitian ini mengkaji evaluasi metode seleksi eigenvector pada ESF yang dikombinasikan dengan model pembelajaran mesin berbasis Gradient Boosting Machine (GBM) untuk prediksi konsentrasi PM2.5 dengan tujuan mengatasi autokorelasi spasial dan meningkatkan akurasi model prediksi data spasial. Metode seleksi eigenvector yang dibandingkan meliputi kriteria eigenvalue positif, Indeks Moran, LASSO, dan SHAP (Shapley Additive exPlanations). Data simulasi dibangkitkan menggunakan tiga model dasar spasial (SEM, SAR, GSM) dengan berbagai tingkat autokorelasi spasial. Sedangkan data empiris meliputi konsentrasi PM2.5 diperoleh melalui stasiun pemantauan di Provinsi DKI Jakarta, peubah lingkungan dan meteorologi melalui Google Earth Engine, serta data lokasi jalan raya dan industri melalui OpenStreetMap. Hasil simulasi menunjukkan bahwa metode seleksi berbasis SHAP secara konsisten memberikan performa terbaik dengan nilai akurasi pemodelan tertinggi dan distribusi nilai yang paling sempit di seluruh skenario autokorelasi spasial dan model pembangkitan data. Metode ini mampu memilih jumlah eigenvector yang lebih sedikit dan lebih efektif sehingga menghasilkan model yang lebih stabil dan akurat dibandingkan metode lain. Pada data empiris, penambahan eigenvector melalui ESF meningkatkan kinerja model GBM dalam memprediksi konsentrasi PM2.5, dengan metode SHAP menunjukkan hasil terbaik pada skenario 10-fold cross-validation. Namun, pada pemisahan data berbasis spatial blocked cross-validation yang lebih ketat secara spasial, peningkatan kinerja model menjadi kurang signifikan. Penerapan evaluasi metode melalui data empiris pada kasus PM2.5 menunjukkan bahwa penambahan eigenvector secara signifikan meningkatkan kinerja model dibandingkan dengan model tanpa komponen spasial. Berdasarkan skenario 10-fold cross-validation, metode seleksi berbasis SHAP menghasilkan akurasi prediksi tertinggi dengan nilai, serta mampu menangkap ketergantungan spasial dan hubungan nonlinier secara efektif. Metode SHAP menunjukkan tingkat ketahanan yang baik melalui pemilihan komponen spasial yang stabil dan konsisten pada berbagai wilayah. Temuan ini menunjukkan keunggulan metodologis dari integrasi ESF dengan pembelajaran mesin dan seleksi fitur berbasis SHAP sehingga menghasilkan kerangka pemodelan spasial yang lebih interpretatif dan robust. Temuan baru dari penelitian ini adalah penerapan SHAP sebagai metode seleksi eigenvector yang tidak hanya meningkatkan akurasi prediksi tetapi juga memberikan interpretasi yang lebih jelas terhadap kontribusi pola spasial dalam model. Pendekatan ini mengatasi keterbatasan metode seleksi tradisional yang kurang eksplisit dalam menjelaskan peran masing-masing eigenvector. Implikasi penelitian ini mendukung pengembangan model prediksi kualitas udara yang lebih akurat dan interpretatif. Saran untuk penelitian selanjutnya adalah peningkatan kinerja model pada wilayah yang terpisah secara spasial dengan integrasi pemodelan spasiotemporal dan pengembangan matriks pembobot spasial yang lebih representatif. Temuan ini membuka peluang pengembangan lebih lanjut dengan menggabungkan pemodelan spasiotemporal dan penyempurnaan matriks pembobot spasial untuk meningkatkan representativitas dan kinerja model pada skala yang lebih luas. Implikasi praktis dari penelitian ini dapat digunakan pada pengembangan sistem peringatan dini kualitas udara di kawasan perkotaan yang memerlukan model prediksi yang akurat dan mudah diinterpretasikan untuk mendukung pengambilan keputusan pengendalian polusi udara.
       
      Spatial autocorrelation is a phenomenon in which spatial dependence patterns indicate a systematic relationship between the value of a variable at a given location and the value of the same variable at neighboring locations. The presence of this characteristic in spatial data must be considered because it can introduce bias in statistical models if not properly addressed. One effective method for handling this issue is Eigenvector Spatial Filtering (ESF). ESF utilizes the eigenvectors of a spatial weights matrix to capture spatial dependence patterns in the data and incorporates these patterns into the model as additional independent variables, thereby improving the quality of predictive models. In the implementation of ESF, selecting the optimal eigenvectors is an important aspect to be considered. This step is crucial in modeling to avoid overfitting, enhance model interpretability, and reduce computational burden. Although commonly used eigenvector selection methods are effective in partial selection, they do not yet provide explicit information regarding the individual contribution of each eigenvector to model prediction. Alternatively, the SHAP method may be applied to feature selection offering a robust basis in game theory. Eigenvectors with the highest SHAP values can be prioritized in ESF modeling, whereas those with low contributions can be eliminated. The ESF method remains limited in capturing nonlinear relationships and complex interactions between variables. Machine learning algorithms have become effective approaches because of their ability to model complex and nonlinear patterns. Thus, the integration of ESF with machine learning algorithms has been implemented alongside the growing availability of large-scale geospatial data. In this study, we evaluate eigenvector selection methods in ESF with the integration of boosting models in the case of air quality, using PM2.5 as the indicator. Particulate Matter (PM2.5) refers to fine particles with diameters less than 2.5 µm that are used as a primary indicator of air quality. Their extremely small size allows them to penetrate deep into the human respiratory tract and cause respiratory disorders, cardiovascular disease, and increased risk of premature death. The province of DKI Jakarta, as the capital of Indonesia, faces increasingly serious air pollution problems due to high population density and intense transportation usage, so that PM2.5 concentrations frequently exceed safe thresholds. This study examines the evaluation of eigenvector selection methods in ESF combined with a machine learning model based on Gradient Boosting Machine (GBM) for predicting PM2.5 concentrations, with the aim of addressing spatial autocorrelation and improving the accuracy of spatial data predictive models. The eigenvector selection methods compared include the positive eigenvalue criterion, Moran’s Index, LASSO, and SHAP (Shapley Additive exPlanations). Simulation data is generated using three basic spatial models (SEM, SAR, GSM) with various degrees of spatial autocorrelation, while empirical data covers PM2.5 concentrations obtained from monitoring stations in DKI Jakarta Province, environmental and meteorological variables via Google Earth Engine, and road network and industry location data via OpenStreetMap. Simulation results show that SHAP-based selection methods consistently deliver the best performance, with the highest modeling accuracy values and the narrowest value distributions across all spatial autocorrelation scenarios and data generation models. This method can select a smaller and more effective set of eigenvectors, resulting in more stable and accurate models compared to other methods. In empirical data, the addition of eigenvectors through ESF improves GBM model performance in predicting PM2.5 concentrations, with the SHAP method achieving the best results in the 10-fold cross-validation scenario. However, with spatially stricter data splitting using spatial blocked cross-validation, improvements in model performance have become less significant. Application of the method evaluation using empirical data in the PM2.5 case demonstrates that the addition of eigenvectors significantly enhances model performance compared to models without spatial components. Based on the 10-fold cross-validation scenario, SHAP-based selection methods yield the highest prediction accuracy, as well as effectively capturing spatial dependence and nonlinear relationships. The SHAP method displays strong robustness through stable and consistent selection of spatial components across regions. These findings demonstrate a methodological advantage from the integration of ESF with machine learning and SHAP-based feature selection, resulting in a more interpretable and robust spatial modeling framework. The novelty of this study is the application of SHAP as an eigenvector selection method, which not only improves prediction accuracy but also provides clearer interpretation of the contribution of spatial patterns in the model. This approach overcomes the limitations of traditional selection methods that are less explicit in explaining the role of each eigenvector. The implications of this research support the development of more accurate and interpretable air quality prediction models. Recommendations for future research include improving model performance in spatially separated regions by integrating spatiotemporal modeling and developing more representative spatial weights matrices. These findings open further opportunities for development by combining spatiotemporal modeling and refining the spatial weights matrix (SWM) to enhance representativeness and model performance on a larger scale. The practical implications of this research can be used in the development of early warning systems for urban air quality, which require accurate and easily interpreted predictive models to support decision-making in air pollution control.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/177734
      Collections
      • MF - School of Data Science, Mathematic and Informatics [138]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository