IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      KINERJA JARAK ROBUST MAHALANOBIS DENGAN MATRIX MINIMUM COVARIANCE DETERMINANT PADA ANALISIS GEROMBOL DATA MENGANDUNG PENCILAN

      Thumbnail
      View/Open
      Cover (643.9Kb)
      Fulltext (2.654Mb)
      Lampiran (5.308Mb)
      Date
      2026
      Author
      Fitrianti, Dwi
      Fitrianto, Anwar
      Kurnia, Anang
      Metadata
      Show full item record
      Abstract
      Analisis gerombol merupakan teknik pengelompokan amatan berdasarkan kemiripan karakteristiknya tanpa menggunakan informasi awal kelompok. Kinerja analisis gerombol dapat dipengaruhi oleh keberadaan pencilan, korelasi antarpeubah, dan bentuk gerombol. Jarak Euclidean tidak memperhitungkan korelasi antarpeubah, sedangkan jarak Mahalanobis klasik sensitif terhadap pencilan karena menggunakan penduga rataan dan matriks ragam-peragam klasik. Sebagai alternatif, digunakan jarak robust Mahalanobis dengan Minimum Covariance Determinant (RMD-MCD) dan Matrix Minimum Covariance Determinant (RMD-MMCD). Penelitian ini bertujuan mengkaji kinerja RMD MMCD pada metode K-means, K-medoids, DBSCAN, dan DB-Kmeans untuk data yang mengandung pencilan, serta menggerombolkan desa dan mengidentifikasi karakteristik wilayah di Provinsi Sumatera Selatan berdasarkan data Potensi Desa (PODES) tahun 2024. Penelitian dilakukan melalui kajian simulasi dan penerapan pada data PODES. Data simulasi terdiri atas 800 amatan dan enam peubah. Simulasi mencakup tiga pola gerombol, yaitu terpisah, beririsan, dan spiral. Korelasi antarpeubah ditetapkan sebesar 0,1, 0,5, dan 0,9. Persentase pencilan yang digunakan sebesar 0%, 1%, 5%, dan 10%. Tingkat pencilan ditetapkan sebesar 1s, 2s, dan 3s, dengan satu atau dua peubah yang terkontaminasi. Setiap skenario direplikasi sebanyak 50 kali. Metode K-means, K-medoids, DBSCAN, dan DB Kmeans diterapkan menggunakan jarak Euclidean, Mahalanobis klasik, RMD MCD, dan RMD-MMCD. Kinerja penggerombolan dievaluasi menggunakan Adjusted Rand Index (ARI), Rand Index (RI), Silhouette, Dunn Index, dan Davies Bouldin Index (DBI). Kajian simulasi menunjukkan bahwa pola gerombol merupakan faktor yang paling kuat memengaruhi kinerja penggerombolan, diikuti oleh persentase pencilan dan korelasi antarpeubah. Euclidean cenderung stabil pada pola terpisah dan metode berbasis pusat, sedangkan RMD-MCD dan RMD-MMCD memberikan hasil yang lebih baik pada beberapa kondisi berpencilan dan berpola kompleks. Mahalanobis klasik cenderung berkinerja lebih rendah karena sensitif terhadap pencilan. Pada DBSCAN, RMD-MMCD mampu mendeteksi 78,74% pencilan sebagai noise, dengan 10,56% amatan normal keliru teridentifikasi sebagai noise. RMD-MCD mendeteksi 78,51% pencilan, tetapi menghasilkan kesalahan identifikasi amatan normal yang lebih tinggi, yaitu 11,86%. Kajian empirik penggerombolan desa berdasarkan data Podes di Sumatera Selatan menghasilkan bahwa kombinasi DBSCAN dan RMD-MMCD memberikan hasil terbaik secara umum berdasarkan nilai Silhouette tertinggi sebesar 0,5290 dan DBI terendah sebesar 0,0973, dengan Dunn Index sebesar 0,1699. Kombinasi tersebut menghasilkan dua gerombol utama dan satu kelompok noise. Karakteristik gerombol utama sangat dipengaruhi oleh kemudahan akses, kelengkapan fasilitas pelayanan umum, dan keaktifan kegiatan pemerintahan desa, yang mencakup desa desa reguler dan pusat administrasi lokal. Sementara itu, kelompok noise berhasil mengidentifikasi desa-desa pinggiran kota dengan karakteristik khusus yang menyimpang dari pola mayoritas. Dengan demikian, Euclidean lebih sesuai untuk pola gerombol terpisah dan metode berbasis pusat. RMD-MMCD direkomendasikan terutama untuk data berstruktur matriks yang mengandung pencilan, memiliki korelasi tinggi, dan pola gerombol kompleks karena mampu mendeteksi pencilan dengan kesalahan identifikasi amatan normal yang lebih rendah dibandingkan RMD-MCD. Penerapan DBSCAN dengan RMD-MMCD pada data PODES juga mampu membentuk gerombol wilayah berdasarkan perbedaan akses, fasilitas pelayanan umum, dan kegiatan pemerintahan desa serta mengidentifikasi wilayah dengan karakteristik khusus sebagai noise.
       
      Cluster analysis is a technique for grouping observations based on the similarity of their characteristics without using prior information about group membership. The performance of cluster analysis may be affected by the presence of outliers, correlations among variables, and cluster shapes. Euclidean distance does not account for correlations among variables, whereas the classical Mahalanobis distance is sensitive to outliers because it relies on classical estimators of the mean and covariance matrix. As alternatives, robust Mahalanobis distance based on the Minimum Covariance Determinant (RMD-MCD) and the Matrix Minimum Covariance Determinant (RMD-MMCD) are considered. This study aims to evaluate the performance of RMD-MMCD in K-means, K-medoids, DBSCAN, and DB-Kmeans for data containing outliers, as well as to cluster villages and identify regional characteristics in South Sumatra Province using the 2024 Village Potential Statistics (PODES) data. The study was conducted through a simulation study and an application to the PODES data. The simulated data consisted of 800 observations and six variables. The simulation included three cluster patterns: well-separated, overlapping, and spiral. Correlations among variables were set at 0.1, 0.5, and 0.9. The proportions of outliers were 0%, 1%, 5%, and 10%. The outlier severity levels were set at 1s, 2s, and 3s, with either one or two contaminated variables. Each scenario was replicated 50 times. K-means, K-medoids, DBSCAN, and DB Kmeans were implemented using Euclidean distance, classical Mahalanobis distance, RMD-MCD, and RMD-MMCD. Clustering performance was evaluated using the Adjusted Rand Index (ARI), Rand Index (RI), Silhouette, Dunn Index, and Davies–Bouldin Index (DBI). The simulation study showed that cluster pattern was the most influential factor affecting clustering performance, followed by the proportion of outliers and the correlation among variables. Euclidean distance tended to perform consistently for well-separated cluster patterns and center-based clustering methods, whereas RMD-MCD and RMD-MMCD produced better results under several conditions involving outliers and complex cluster patterns. The classical Mahalanobis distance generally showed lower performance because of its sensitivity to outliers. In DBSCAN, RMD-MMCD detected 78.74% of the outliers as noise, while incorrectly identifying 10.56% of normal observations as noise. RMD-MCD detected 78.51% of the outliers but produced a higher misclassification rate for normal observations, at 11.86%. The empirical clustering analysis of villages in South Sumatra using the PODES data showed that the combination of DBSCAN and RMD-MMCD provided the best overall performance, as indicated by the highest Silhouette value of 0.5290 and the lowest DBI value of 0.0973, with a Dunn Index of 0.1699. This combination produced two main clusters and a set of noise observations. The characteristics of the main clusters were strongly influenced by accessibility, the availability of public service facilities, and the level of village government activity. These clusters consisted of both regular villages and local administrative centers. Meanwhile, the noise group successfully identified peri-urban villages with distinctive characteristics that differed from the general pattern of most villages. Therefore, Euclidean distance is more suitable for well-separated cluster patterns and center-based clustering methods. RMD-MMCD is particularly recommended for matrix-variate data containing outliers, strong correlations, and complex cluster patterns because it can detect outliers with a lower rate of incorrectly identifying normal observations than RMD-MCD. The application of DBSCAN with RMD-MMCD to the PODES data was also able to form regional clusters based on differences in accessibility, public service facilities, and village government activities, while identifying areas with distinctive characteristics as noise.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/177725
      Collections
      • MF - School of Data Science, Mathematic and Informatics [138]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository