IPB University Logo

SCIENTIFIC REPOSITORY

IPB University Scientific Repository collects, disseminates, and provides persistent and reliable access to the research and scholarship of faculty, staff, and students at IPB University

AI Repository
 
Building and Categories


      View Item 
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      •   IPB Repository
      • Final Assignments
      • Master Final Assignments
      • MF - School of Data Science, Mathematic and Informatics
      • View Item
      JavaScript is disabled for your browser. Some features of this site may not work without it.

      Evaluasi Kinerja BERTopic dalam Pemodelan Topik Multidomain pada Teks Berita Berbahasa Indonesia

      Thumbnail
      View/Open
      Cover (587.5Kb)
      Fulltext (1.264Mb)
      Lampiran (748.9Kb)
      Date
      2026
      Jenis/Type
      Tesis
      Subtype
      Theses
      Author
      Saputra, Wawan
      Soleh, Agus Mohamad
      Saefuddin, Asep
      Metadata
      Show full item record
      Abstract
      Pemodelan topik berbasis Transformer semakin banyak digunakan untuk mengidentifikasi struktur tematik pada data teks berskala besar. Salah satu metode yang berkembang adalah BERTopic, yang mengintegrasikan representasi dokumen berbasis embedding, reduksi dimensi, clustering, dan representasi topik. Kinerja BERTopic dapat bervariasi berdasarkan kombinasi komponen dan parameter yang digunakan serta karakteristik korpus yang dianalisis. Kajian yang mengevaluasi BERTopic secara sistematis pada teks berita berbahasa Indonesia dari beberapa portal juga masih terbatas. Penelitian ini bertujuan mengevaluasi secara deskriptif kinerja BERTopic berdasarkan variasi model embedding, algoritma clustering, dan metode representasi topik, mengidentifikasi konfigurasi dengan TQ maksimum pada setiap domain, serta menginterpretasikan topik dan komposisi sumber berita pada pemodelan Multidomain. Data penelitian terdiri atas 1.846 artikel berita mengenai Program Makan Bergizi Gratis (MBG) yang berasal dari CNN Indonesia sebanyak 396 artikel, Kompas sebanyak 617 artikel, dan Detik sebanyak 833 artikel. Data dianalisis dalam empat domain, yaitu CNN Indonesia, Kompas, Detik, dan Multidomain yang merupakan gabungan ketiga portal. Eksperimen membandingkan empat model embedding, yaitu IndoBERT, IndoSBERT, multilingual BERT, dan multilingual MPNet. Embedding direduksi menggunakan UMAP dengan n_components 5, 10, dan 15, kemudian dikelompokkan menggunakan K-Means, BIRCH, dan HDBSCAN dengan variasi parameter 5, 10, 15, 20, dan 30. Representasi topik dibandingkan menggunakan c-TF-IDF dan KeyBERTInspired dengan rentang unigram (1,1) dan unigram-bigram (1,2). Sebanyak 720 konfigurasi dievaluasi pada setiap domain sehingga diperoleh 2.880 konfigurasi. Evaluasi dilakukan menggunakan Topic Coherence (TC), Topic Diversity (TD), dan Topic Quality (TQ), sedangkan proporsi outlier digunakan sebagai informasi tambahan untuk menilai cakupan dokumen pada konfigurasi HDBSCAN. Hasil menunjukkan bahwa kinerja komponen BERTopic bervariasi antardomain. Berdasarkan rata-rata TQ, IndoBERT menghasilkan nilai tertinggi pada CNN Indonesia, sedangkan IndoSBERT menghasilkan rata-rata tertinggi pada Kompas, Detik, dan Multidomain. HDBSCAN menghasilkan rata-rata dan nilai maksimum TQ tertinggi dibandingkan K-Means dan BIRCH pada seluruh domain. Demikian pula, c-TF-IDF dengan rentang unigram-bigram (1,2) menghasilkan rata-rata dan nilai maksimum TQ tertinggi pada seluruh domain dibandingkan metode representasi lainnya. Hasil tersebut menunjukkan bahwa model embedding bersifat lebih spesifik terhadap domain, sedangkan HDBSCAN dan c-TF-IDF unigram-bigram menunjukkan pola kinerja yang lebih konsisten dalam ruang konfigurasi yang dievaluasi. Konfigurasi terbaik CNN Indonesia menggunakan IndoBERT, UMAP dengan n_components 15, HDBSCAN dengan min_cluster_size 30, dan c-TF-IDF unigram-bigram, dengan TQ sebesar 0,7489 dan proporsi outlier 58,59%. Konfigurasi terbaik Kompas menggunakan IndoBERT, UMAP-5, HDBSCAN dengan min_cluster_size 15, dan c-TF-IDF unigram-bigram, dengan TQ tertinggi sebesar 0,7831 dan proporsi outlier 20,42%. Pada Detik, konfigurasi terbaik menggunakan IndoSBERT, UMAP-15, HDBSCAN dengan min_cluster_size 15, dan c-TF-IDF unigram-bigram, dengan TQ sebesar 0,7355 dan proporsi outlier 0,96%. Pada multidomain, beberapa konfigurasi multilingual MPNet mencapai TQ maksimum yang sama sebesar 0,7499. Salah satu konfigurasi tersebut menggunakan UMAP dengan n_components = 5, HDBSCAN dengan min_cluster_size = 15, dan c-TF-IDF unigram-bigram, serta menghasilkan TC sebesar 0,7499, TD sebesar 1,0000, dan tidak menghasilkan dokumen outlier. Perbedaan proporsi outlier tersebut menunjukkan bahwa TQ yang tinggi tidak selalu disertai cakupan dokumen yang tinggi sehingga kedua informasi perlu dipertimbangkan secara bersama dalam menafsirkan kinerja BERTopic. Interpretasi terhadap konfigurasi yang dipilih menghasilkan tiga topik pada CNN Indonesia, yaitu keracunan makanan siswa, anggaran dan tata kelola MBG, serta wacana pendanaan zakat. Kompas menghasilkan dua topik, yaitu pelaksanaan MBG di sekolah dan wacana zakat untuk pendanaan MBG, sedangkan Detik menghasilkan dua topik berupa pelaksanaan dan isu keamanan pangan MBG serta mitra dapur dan pembayaran MBG. Pada Multidomain terbentuk dua topik, yaitu pelaksanaan umum program MBG dan susu, sapi perah, serta rantai pasok. Topik pelaksanaan umum mencakup 1.799 dari 1.846 dokumen dan memiliki komposisi portal yang relatif mengikuti komposisi awal korpus, sedangkan topik susu, sapi perah, dan rantai pasok terdiri atas 47 dokumen dengan proporsi dokumen Kompas yang relatif lebih besar. Hasil tersebut menunjukkan bahwa pemodelan Multidomain menghasilkan struktur topik yang lebih ringkas dan dapat menggambarkan tema bersama lintas portal, sedangkan pemodelan domain tunggal memberikan rincian tema yang tidak selalu terbentuk sebagai topik tersendiri setelah korpus digabungkan. Secara keseluruhan, hasil penelitian menunjukkan bahwa evaluasi BERTopic perlu mempertimbangkan kombinasi komponen pemodelan, kualitas topik, cakupan dokumen, dan karakteristik domain.
       
      Transformer-based topic modeling has increasingly been used to identify thematic structures in large-scale text corpora. One increasingly adopted approach is BERTopic, which integrates embedding-based document representation, dimensionality reduction, clustering, and topic representation. However, BERTopic performance may vary across different modeling components, parameter settings, and corpus characteristics. Systematic evaluations of BERTopic on Indonesian-language news texts collected from multiple news portals remain limited. This study aimed to descriptively evaluate BERTopic performance across different embedding models, clustering algorithms, and topic representation methods; identify the configurations achieving the maximum TQ in each domain; and interpret the resulting topics and news-source composition in Multidomain modeling. The dataset consisted of 1,846 news articles concerning Indonesia's Free Nutritious Meal Program (Makan Bergizi Gratis, MBG), comprising 396 articles from CNN Indonesia, 617 from Kompas, and 833 from Detik. The data were analyzed in four domains: CNN Indonesia, Kompas, Detik, and a Multidomain corpus combining the three portals. The experiment compared four embedding models: IndoBERT, IndoSBERT, multilingual BERT, and multilingual MPNet. The embeddings were reduced using UMAP with n_components of 5, 10, and 15 and subsequently clustered using K-Means, BIRCH, and HDBSCAN with parameter values of 5, 10, 15, 20, and 30. Topic representations were compared using c-TF-IDF and KeyBERTInspired with unigram (1,1) and unigram-bigram (1,2) ranges. A total of 720 configurations were evaluated for each domain, resulting in 2,880 configurations. Performance was assessed using Topic Coherence (TC), Topic Diversity (TD), and Topic Quality (TQ), while the outlier proportion was used as additional information to assess document coverage in HDBSCAN configurations. The results showed that the performance of BERTopic components varied across domains. Based on mean TQ, IndoBERT achieved the highest value for CNN Indonesia, whereas IndoSBERT achieved the highest mean TQ for Kompas, Detik, and the Multidomain corpus. HDBSCAN produced the highest mean and maximum TQ values across all domains compared with K-Means and BIRCH. Similarly, c-TF-IDF with the unigram-bigram (1,2) range consistently achieved the highest mean and maximum TQ values across all domains compared with the other topic representation methods. These results indicate that embedding performance was more domain-specific, whereas HDBSCAN and c-TF-IDF with unigram-bigram representation exhibited more consistent performance patterns within the evaluated configuration space. The best CNN Indonesia configuration used IndoBERT, UMAP with n_components = 15, HDBSCAN with min_cluster_size = 30, and c-TF-IDF unigram-bigram, achieving a TQ of 0.7489 with an outlier proportion of 58.59%. The best Kompas configuration used IndoBERT, UMAP-5, HDBSCAN with min_cluster_size = 15, and c-TF-IDF unigram-bigram, achieving the highest overall TQ of 0.7831 with an outlier proportion of 20.42%. For Detik, the best configuration used IndoSBERT, UMAP-15, HDBSCAN with min_cluster_size = 15, and c-TF-IDF unigram-bigram, producing a TQ of 0.7355 and an outlier proportion of 0.96%. In the multidomain corpus, multiple multilingual MPNet configurations achieved the same maximum TQ of 0.7499. One of these configurations used UMAP with n_components = 5, HDBSCAN with min_cluster_size = 15, and c-TF-IDF unigram-bigram, producing a TC of 0.7499, a TD of 1.0000, and no outlier documents. The differences in outlier proportions demonstrate that a high TQ does not necessarily correspond to high document coverage. Therefore, both aspects should be considered when interpreting BERTopic performance. Interpretation of the selected configurations identified three topics in CNN Indonesia: student food poisoning, MBG budgeting and governance, and the discourse on zakat-based funding. Kompas produced two topics: MBG implementation in schools and the discourse on using zakat to fund MBG. Detik also produced two topics: MBG implementation and food-safety issues, and kitchen partners and MBG payments. The Multidomain corpus produced two topics: general implementation of the MBG program and milk, dairy cattle, and the supply chain. The general implementation topic contained 1,799 of the 1,846 documents and had a source composition relatively similar to that of the overall corpus, whereas the milk, dairy cattle, and supply-chain topic contained 47 documents and had a relatively higher proportion of documents from Kompas. These findings indicate that Multidomain modeling produces a more compact topic structure, whereas single-domain modeling provides more detailed themes that do not necessarily remain distinct after the corpora are combined. Overall, the findings indicate that BERTopic evaluation should consider combinations of modeling components, topic quality, document coverage, and domain characteristics.
       
      URI
      http://repository.ipb.ac.id/handle/123456789/179954
      Collections
      • MF - School of Data Science, Mathematic and Informatics [180]

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository
        

       

      Browse

      All of IPB RepositoryCollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

      My Account

      Login

      Application

      google store

      Copyright © 2020 Library of IPB University
      All rights reserved
      Contact Us | Send Feedback
      Indonesia DSpace Group 
      IPB University Scientific Repository
      UIN Syarif Hidayatullah Institutional Repository
      Universitas Jember Digital Repository