IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      Evaluasi Kinerja Biclustering Algoritma Cheng and Church dan QUBIC pada Data Produksi Perikanan Budidaya di Indonesia

      Thumbnail
      View/Open
      Cover (644.0Kb)
      Fulltext (1.893Mb)
      Lampiran (652.2Kb)
      Date
      2026
      Author
      Pomalingo, Dewi Zulyani
      Aidi, Muhammad Nur
      Sadik, Kusman
      Metadata
      Show full item record
      Abstract
      Biclustering merupakan sebuah metode penggerombolan objek dan peubah secara simultan sehingga mampu menemukan pola lokal yang tidak selalu terlihat melalui metode clustering klasik. Beragam algoritma biclustering memiliki cara kerja dan tujuan optimasi yang berbeda, sementara belum terdapat pedoman khusus yang dapat dijadikan acuan untuk memilih algoritma paling sesuai. Penelitian ini bertujuan untuk mengevaluasi kinerja dua algoritma biclustering, yaitu CC (Cheng and Church) dan QUBIC (qualitative biclustering) pada data simulasi dan data produksi perikanan budidaya di Indonesia tahun 2023. Algoritma CC mengidentifikasi bicluster berdasarkan homogenitas numerik melalui minimisasi Mean Squared Residue (MSR), sedangkan QUBIC mengidentifikasi kesamaan pola kualitatif melalui proses diskretisasi. Data simulasi dibangkitkan dalam matriks berukuran 34 × 14 dengan background berdistribusi normal baku. Dua bicluster dengan model konstan memiliki rataan 5 dan 10 serta ragam 0,25 disisipkan pada matriks tersebut. Skenario simulasi terdiri atas ukuran bicluster kecil 10 × 4, sedang 12 × 5, dan besar 15 × 6, serta tingkat tumpang tindih 0%, 10%, dan 20%. Nilai pada daerah tumpang tindih ditentukan menggunakan rata-rata terbobot. Setiap kombinasi skenario diulang sebanyak 100 kali, kemudian baris dan kolom matriks diacak sebelum algoritma CC dan QUBIC diterapkan. Kemampuan algoritma dalam mengidentifikasi bicluster aktual dievaluasi menggunakan indeks Liu dan Wang. Hasil simulasi menunjukkan bahwa QUBIC menghasilkan indeks Liu dan Wang yang lebih tinggi daripada CC pada seluruh skenario, dengan rataan keseluruhan masing-masing sebesar 0,870 dan 0,382. Uji Wilcoxon berpasangan menunjukkan perbedaan kinerja yang signifikan, sehingga QUBIC secara umum lebih baik dalam mengidentifikasi bicluster aktual. CC memberikan kinerja terbaik pada bicluster berukuran besar, sedangkan QUBIC memberikan kinerja terbaik pada ukuran kecil dan sedang. Data empiris berupa volume produksi 14 jenis perikanan budidaya pada 34 provinsi telah distandardisasi. Pada algoritma CC, threshold optimal ditentukan melalui pengujian beberapa nilai ambang batas d yang menghasilkan ASR rendah dengan kestabilan struktur bicluster. CC menghasilkan 7 bicluster pada threshold optimal d = 0,009, mencakup 33 provinsi dengan nilai Average Squared Residue (ASR) sebesar 0,0052. Pada algoritma QUBIC, pemilihan kombinasi parameter dilakukan dengan mempertimbangkan jumlah bicluster nontrivial, rata-rata volume bicluster, dan nilai ASR. QUBIC menghasilkan 8 bicluster pada kombinasi parameter r = 1, q = 0,06, dan c = 0,75 yang mencakup 31 provinsi dengan ASR sebesar 0,4360. Nilai ASR CC yang lebih rendah dibanding QUBIC menunjukkan bahwa CC lebih unggul dalam membentuk bicluster yang homogen pada data produksi perikanan budidaya. Hal ini juga terlihat pada plot profil bicluster CC yang cenderung lebih berimpit dan sejajar, sedangkan profil bicluster QUBIC lebih menyebar, dan umumnya menunjukkan pola naik turun yang searah. Jumlah keanggotaan provinsi yang tercakup menunjukkan bahwa kedua algoritma mampu menangkap pola pada sebagian besar provinsi di Indonesia, meskipun CC memiliki cakupan yang sedikit lebih luas. Nilai indeks Liu dan Wang sebesar 0,1951 menunjukkan bahwa struktur keanggotaan hasil bicluster CC dan QUBIC memiliki tingkat kemiripan yang relatif rendah. Dengan demikian, meskipun kedua algoritma diterapkan pada data yang sama, yakni produksi perikanan budidaya, CC dan QUBIC cenderung menghasilkan penggerombolan yang berbeda.
       
      Biclustering is a method that simultaneously clusters objects and variables, thereby enabling the identification of local patterns that may not be detected by conventional clustering methods. Various biclustering algorithms have different mechanisms and optimization objectives, while no specific guideline is currently available for selecting the most appropriate algorithm. This study aims to evaluate the performance of two biclustering algorithms, namely CC (Cheng and Church) and QUBIC (Qualitative Biclustering), using simulated data and Indonesian aquaculture production data from 2023. The CC algorithm identifies biclusters based on numerical homogeneity through the minimization of Mean Squared Residue (MSR), whereas QUBIC identifies similarities in qualitative patterns through a discretization process. The simulated data were generated as a 34 × 14 matrix with a standard normally distributed background. Two constant-model biclusters with means of 5 and 10 and a variance of 0,25 were embedded in the matrix. The simulation scenarios consisted of small (10 × 4), medium (12 × 5), and large (15 × 6) biclusters, with overlap levels of 0%, 10%, and 20%. Values in the overlapping regions were determined using a weighted average. Each combination of scenarios was replicated 100 times, after which the rows and columns of the matrix were randomly permuted before the CC and QUBIC algorithms were applied. The ability of the algorithms to identify the true biclusters was evaluated using the Liu and Wang index. The simulation results showed that QUBIC produced higher Liu and Wang index values than CC across all scenarios, with overall means of 0,870 and 0,382, respectively. The paired Wilcoxon test indicated a significant difference in performance, demonstrating that QUBIC was generally better at identifying the true biclusters. CC performed best for large biclusters, whereas QUBIC performed best for small and medium-sized biclusters. The empirical data consisted of the production volumes of 14 aquaculture commodities across 34 provinces and were standardized prior to analysis. For the CC algorithm, the optimal threshold was determined by evaluating several values of the threshold parameter d, with selection based on a low ASR value and the stability of the resulting bicluster structure. CC produced 7 biclusters at the optimal threshold of d=0,009, covering 33 provinces, with an Average Squared Residue (ASR) value of 0,0052. For QUBIC, the parameter combination was selected by considering the number of nontrivial biclusters, the average bicluster volume, and the ASR value. QUBIC produced 8 biclusters using the parameter combination r=1, q=0,06, and c=0,75, covering 31 provinces, with an ASR value of 0,4360. The lower ASR value obtained by CC indicates that CC was superior to QUBIC in forming numerically homogeneous biclusters in the aquaculture production data. This is also reflected in the bicluster profile plots, where CC shows more aligned profiles, while QUBIC shows more dispersed but similar trends. The number of provinces included in the biclusters indicates that both algorithms were able to capture patterns across most provinces in Indonesia, although CC provided slightly broader coverage. The Liu and Wang index value of 0,1951 indicates that the membership structures of the biclusters produced by CC and QUBIC had a relatively low level of similarity. Therefore, although both algorithms were applied to the same aquaculture production data, CC and QUBIC tended to produce different clustering structures.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/179451
      Collections
      • MF - School of Data Science, Mathematic and Informatics [179]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository