Show simple item record

dc.contributor.advisorAfendi, Farit Mochamad
dc.contributor.advisorMasjkur, Mohammad
dc.contributor.authorIZZATI, NAIYA DZIL
dc.date.accessioned2026-08-12T01:19:15Z
dc.date.available2026-08-12T01:19:15Z
dc.date.issued2026
dc.identifier.urihttp://repository.ipb.ac.id/handle/123456789/178331
dc.description.abstractBanjir merupakan salah satu bencana hidrometeorologi yang paling sering terjadi di Indonesia dan sering menjadi topik pembahasan masyarakat di media sosial. Penelitian ini bertujuan menganalisis pengaruh tahapan preprocessing terhadap performa model IndoBERT dan IndoBERTweet pada klasifikasi sentimen terkait banjir Aceh. Data penelitian terdiri atas 5403 cuitan yang diberi label secara manual ke dalam tiga kelas sentimen, yaitu negatif, netral, dan positif. Penelitian menerapkan tiga skenario preprocessing dengan tingkat kompleksitas yang berbeda, kemudian performa model dievaluasi menggunakan balanced accuracy dan F1-score macro. Hasil penelitian menunjukkan bahwa IndoBERT mencapai performa terbaik pada skenario preprocessing minimal dengan nilai balanced accuracy sebesar 0,8575 dan F1-score macro sebesar 0,8428. Sedangkan, IndoBERTweet memperoleh performa terbaik dengan nilai balanced accuracy sebesar 0,8677 F1-score macro sebesar 0,8435, sekaligus menjadi model dengan performa tertinggi pada penelitian ini. IndoBERTweet menunjukkan performa yang lebih konsisten dibandingkan IndoBERT pada seluruh skenario preprocessing. Sebaliknya, performa IndoBERT cenderung menurun seiring bertambahnya tahapan preprocessing. Temuan ini menunjukkan bahwa pengaruh preprocessing berbeda pada setiap model sehingga pemilihan tahapan preprocessing perlu disesuaikan dengan karakteristik model dan data yang digunakan.
dc.description.abstractFloods are one of the most frequent hydrometeorological disasters in Indonesia and are often discussed on social media. This study aimed to analyze the effect of preprocessing on the performance of IndoBERT and IndoBERTweet models for sentiment classification related to the 2025 Aceh flood. The dataset consisted of 5,403 tweets that were manually labeled into three sentiment classes: negative, neutral, and positive. Three preprocessing scenarios with different levels of complexity were applied, and model performance was evaluated using balanced accuracy and macro F1-score. The results showed that IndoBERT achieved its best performance under the minimal preprocessing scenario, with a balanced accuracy of 0.8575 and a macro F1-score of 0.8428. Meanwhile, IndoBERTweet achieved its best performance under a more complex preprocessing scenario, with a balanced accuracy of 0.8677 and a macro F1-score of 0.8435, making it the best-performing model in this study. In general, IndoBERTweet showed more consistent performance than IndoBERT across all preprocessing scenarios. In contrast, the performance of IndoBERT decreased as the preprocessing stages became more complex. These findings indicate that the effect of preprocessing differs across models, and therefore preprocessing strategies should be adjusted to the characteristics of the model and the data used.
dc.description.sponsorship
dc.language.isoid
dc.publisherIPB Universityid
dc.titlePengaruh Tahapan Preprocessing terhadap Performa Model IndoBERT dan IndoBERTweet pada Analisis Sentimen Banjir Aceh 2025id
dc.title.alternativeA Comparative Study of IndoBERT and IndoBERTweet Model under Different Preprocessing Scenarios for Sentiment Analysis of the 2025 Aceh Flood
dc.typeSkripsi
dc.subject.keywordanalisis sentimenid
dc.subject.keywordbanjirid
dc.subject.keywordIndoBERTid
dc.subject.keywordIndoBERTweetid
dc.subject.keywordpreprocessingid
dc.subtypeUndergraduate Theses


Files in this item

Thumbnail
Thumbnail
Thumbnail

This item appears in the following Collection(s)

Show simple item record