IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Undergraduate Final Assignments
      • UF - School of Data Science, Mathematic and Informatics
      • UF - Statistics and Data Sciences
      • View Item
      •   IPB Repository
      • Final Assignments
      • Undergraduate Final Assignments
      • UF - School of Data Science, Mathematic and Informatics
      • UF - Statistics and Data Sciences
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      Clustering Provinsi di Indonesia Berdasarkan Karakteristik Sosial ``Ekonomi Menggunakan K-Means dan DBSCAN

      Thumbnail
      View/Open
      Cover (464.3Kb)
      Fulltext (913.0Kb)
      Lampiran (836.9Kb)
      Date
      2026
      Author
      Salas, Muhammad Rizqa
      Rahardiantoro, Septian
      Anisa, Rahma
      Metadata
      Show full item record
      Abstract
      Indonesia memiliki heterogenitas sosial ekonomi dan kesenjangan pembangunan antarwilayah, sehingga diperlukan segmentasi wilayah sebagai dasar kebijakan yang tepat sasaran. Permasalahan ini didekati melalui algoritma K-Means dan Density-Based Spatial Clustering of Applications with Noise (DBSCAN) yang diterapkan pada 17 indikator sosial ekonomi 38 provinsi dari Badan Pusat Statistik mencakup sektor ekonomi, pendidikan, kesehatan, dan teknologi; karena jumlah peubah yang banyak, data terlebih dahulu direduksi dimensinya melalui Analisis Komponen Utama per sektor, Analisis Faktor, dan Uniform Manifold Approximation and Projection (UMAP), dengan seleksi VIF dan data tanpa penanganan sebagai pembanding. Penelitian ini bertujuan mengelompokkan provinsi berdasarkan indikator tersebut serta membandingkan kinerja kedua algoritma dalam menghasilkan gerombol yang optimal. Kesepuluh kombinasi metode dievaluasi menggunakan Silhouette Index, Davies-Bouldin Index (DBI), dan Calinski-Harabasz Index (CHI). K-Means dengan input UMAP terpilih sebagai metode terbaik (Silhouette Index = 0,4554; DBI = 0,7604; CHI = 46,03) dan menghasilkan tiga gerombol dengan karakteristik berbeda: satu gerombol bernilai rendah pada mayoritas peubah sektor pembangunan yang menghimpun Kawasan Timur Indonesia dan wilayah perbatasan, satu gerombol bernilai menengah, serta satu gerombol bernilai tinggi yang menghimpun sebagian besar provinsi di Sumatra, Jawa, Kalimantan, dan Sulawesi Utara. Secara rata-rata, K-Means (Silhouette Index = 0,3344) cenderung lebih baik dibanding DBSCAN (0,3237) karena sensitivitasnya terhadap hyperparameter e dan MinPts membuat beberapa kombinasi metode dengan DBSCAN pada kasus ini kurang mampu memisahkan karakteristik provinsi, sehingga hanya menghasilkan satu gerombol dan sisanya dianggap sebagai noise.
       
      Indonesia has socio-economic heterogeneity and development gaps between regions, so regional segmentation is needed as a basis for targeted policies. This problem is approached through the K-Means algorithm and the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) which is applied to 17 socio-economic indicators in 38 provinces from the Central Statistics Agency covering the economic, education, health, and technology sectors; Due to the large number of variables, the data is first dimensionally reduced through Key Component Analysis per sector, Factor Analysis, and Uniform Manifold Approximation and Projection (UMAP), with VIF selection and unhandled data as a comparison. This study aims to group provinces based on these indicators and compare the performance of the two algorithms in producing optimal clusters. The ten combinations of methods were evaluated using the Silhouette Index, the Davies-Bouldin Index (DBI), and the Calinski-Harabasz Index (CHI). K-Means with UMAP input selected as the best method (Silhouette Index = 0.4554; DBI = 0.7604; CHI = 46.03) and produced three gangs with different characteristics: one low- value gang in the majority of development sector variables that gathers the Eastern Region of Indonesia and border areas, one mid-value gang, and one high-value gang that gathers most of the provinces in Sumatra, Java, Kalimantan, and North Sulawesi. On average, K-Means (Silhouette Index = 0.3344) tends to be better than DBSCAN (0.3237) because of its sensitivity to e hyperparameters and MinPts makes some combination methods with DBSCAN in this case less able to separate the provincial characteristics, resulting in only one cluster and the rest being considered as noise.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/178210
      Collections
      • UF - Statistics and Data Sciences [166]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository