IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Undergraduate Final Assignments
      • UF - School of Data Science, Mathematic and Informatics
      • UF - Statistics and Data Sciences
      • View Item
      •   IPB Repository
      • Final Assignments
      • Undergraduate Final Assignments
      • UF - School of Data Science, Mathematic and Informatics
      • UF - Statistics and Data Sciences
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      Perbandingan Kinerja Partition-Based dan Density-Based Clustering pada Data Mengandung Outlier

      Thumbnail
      View/Open
      Cover (444.9Kb)
      Fulltext (971.0Kb)
      Lampiran (197.6Kb)
      Date
      2026
      Jenis/Type
      Skripsi
      Subtype
      Undergraduate Theses
      Author
      HASANAH, DELITA NUR
      Wijayanto, Hari
      Rahman, La Ode Abdul
      Metadata
      Show full item record
      Abstract
      Keberadaan outlier merupakan salah satu tantangan dalam analisis clustering karena dapat memengaruhi struktur data dan menurunkan ketepatan hasil pengelompokan. Penelitian ini bertujuan membandingkan kinerja metode partitionbased clustering (K-Means dan K-Medoids) dengan density-based clustering (DBSCAN) pada data yang mengandung outlier, serta menerapkannya pada data Indeks Kualitas Lingkungan Hidup (IKLH) kabupaten/kota di Pulau Jawa. Data simulasi dibangkitkan menggunakan distribusi normal multivariat melalui 12 skenario yang mengombinasikan kondisi overlap dan non-overlap, jumlah observasi sebanyak 100 dan 500, serta proporsi outlier sebesar 0%, 5%, dan 10%, dengan 100 ulangan pada setiap skenario. Kinerja metode dievaluasi menggunakan euclidean distance antara pusat cluster hasil clustering dan pusat clustersebenarnya. Pada data empiris, evaluasi dilakukan menggunakan Silhouette Index, DaviesBouldin Index, dan Calinski-Harabasz Index. Hasil penelitian menunjukkan bahwa peningkatan proporsi outlier meningkatkan nilai euclidean distance pada seluruh metode, dengan K-Means sebagai metode yang paling sensitif terhadap outlier. Sebaliknya, DBSCAN secara konsisten menghasilkan euclidean distance terendah pada data simulasi serta memperoleh nilai Silhouette Index tertinggi dan DaviesBouldin Index terendah pada data empiris. Hasil tersebut menunjukkan bahwa DBSCAN merupakan metode yang lebih robust terhadap keberadaan outlier dan mampu menghasilkan kualitas clustering yang lebih baik dibandingkan K-Means dan K-Medoids.
       
      Outliers are one of the major challenges in cluster analysis because they can distort the data structure and reduce clustering accuracy. This study aims to compare the performance of partition-based clustering methods (K-Means and K-Medoids) and the density-based clustering method (DBSCAN) on data containing outliers and to apply these methods to the Environmental Quality Index (EQI) data of regencies and municipalities in Java Island. Simulation data were generated from a multivariate normal distribution under 12 scenarios combining overlap and nonoverlap cluster conditions, sample sizes of 100 and 500 observations, and outlier proportions of 0%, 5%, and 10%, with 100 replications for each scenario. The clustering performance was evaluated using the Euclidean distance between the estimated cluster centers and the true cluster centers. For the empirical data, clustering quality was assessed using the Silhouette Index, Davies–Bouldin Index, and Calinski–Harabasz Index. The results indicate that increasing the proportion of outliers generally increases the Euclidean distance for all clustering methods, with K-Means being the most sensitive to outliers. In contrast, DBSCAN consistently produced the lowest Euclidean distance in the simulation study and achieved the highest Silhouette Index and the lowest Davies–Bouldin Index in the empirical analysis. These findings demonstrate that DBSCAN is more robust to the presence of outliers and provides better clustering performance than K-Means and KMedoids.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/177259
      Collections
      • UF - Statistics and Data Sciences [166]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository