Evaluasi Ketahanan IndoBERT dalam Analisis Sentimen Berita Ekonomi dan Implikasinya terhadap Prediksi Harga Saham
Date
2026Jenis/Type
TesisSubtype
ThesesAuthor
RESILOY, UNIQUE DESYRRE A.
Sartono, Bagus
Notodiputro, Khairil Anwar
Wigena, Aji Hamim
Metadata
Show full item recordAbstract
Pemanfaatan analisis sentimen berita ekonomi dalam prediksi harga saham
memerlukan model yang tidak hanya memiliki kinerja klasifikasi yang baik, tetapi
juga mampu mempertahankan kinerjanya ketika karakteristik bahasa dan kualitas
label berubah. Selain itu, informasi sentimen yang dihasilkan model belum tentu
selalu memberikan tambahan informasi yang bermanfaat bagi prediksi harga
saham. Penelitian ini bertujuan mengevaluasi ketahanan IndoBERT-Finansial
terhadap perubahan karakteristik linguistik dan kualitas label, mengevaluasi kinerja
model terpilih pada berita ekonomi empiris berbahasa Indonesia, serta menilai
kontribusi indeks sentimen terhadap prediksi harga penutupan IHSG, BBRI, dan
TLKM menggunakan Long Short-Term Memory (LSTM) dan Gated Recurrent
Unit (GRU).
Kajian simulasi menggunakan 40.000 teks berita ekonomi sintetis yang
dibentuk dalam empat skenario linguistik dengan lima replikasi. Evaluasi dilakukan
terhadap arsitektur Base dan Large pada kondisi label Original dan Relabeled serta
dua skema pelatihan, yaitu fine-tuning langsung (FT-L) dan fine-tuning bertahap
(FT-B). Ketahanan dievaluasi berdasarkan perubahan kinerja klasifikasi dan
konsistensi hasil antarkondisi dengan Macro F1 sebagai metrik utama. Pada kajian
empiris, model terpilih diterapkan pada 32.491 artikel berita ekonomi periode
2021–2025. Probabilitas kelas hasil klasifikasi digunakan untuk membentuk indeks
sentimen harian yang kemudian ditambahkan sebagai peubah eksogen dalam model
LSTM dan GRU untuk prediksi lima horizon, yaitu H+1 hingga H+5. Kontribusi
sentimen dinilai dengan membandingkan galat model tanpa sentimen dan dengan
sentimen.
Hasil simulasi menunjukkan bahwa ketahanan IndoBERT-Finansial tidak
hanya berkaitan dengan ukuran arsitektur, tetapi juga dengan kondisi label, skema
pelatihan, dan interaksi antarfaktor. Kondisi label dan skema fine-tuning
memberikan pengaruh yang signifikan terhadap Macro F1, sementara tidak
ditemukan satu arsitektur maupun skema fine-tuning yang secara konsisten unggul
pada seluruh kondisi pengujian. Berdasarkan keseluruhan evaluasi, IndoBERTFinansial Base dengan FT-L dipilih untuk kajian empiris. Model tersebut
menghasilkan Balanced Accuracy sebesar 0,8510, Macro F1 sebesar 0,8439, dan
Weighted F1 sebesar 0,8512 pada data uji.
Penambahan indeks sentimen tidak memberikan pengaruh yang seragam
terhadap prediksi harga saham. Pada IHSG dan BBRI, perbedaan galat antara model
tanpa sentimen dan dengan sentimen tidak signifikan. Pada TLKM, penambahan
sentimen menurunkan MAPE LSTM dari 3,0328% menjadi 2,8680% dan
menghasilkan perbaikan yang signifikan, tetapi pada GRU justru meningkatkan
MAPE dari 2,7984% menjadi 2,8661% dengan perbedaan yang juga signifikan.
Temuan ini menunjukkan bahwa kualitas klasifikasi sentimen yang baik tidak
secara otomatis menghasilkan peningkatan prediksi harga saham. Manfaat sentimen
berita bergantung pada aset dan arsitektur model yang menggunakannya, sehingga
informasi sentimen lebih tepat diperlakukan sebagai informasi tambahan yang perlu
dievaluasi sesuai konteks pemodelannya. The use of economic news sentiment analysis for stock price prediction
requires a model that not only achieves good classification performance but is also
able to maintain its performance when linguistic characteristics and label quality
change. In addition, sentiment information produced by the model does not
necessarily provide useful additional information for stock price prediction. This
study aims to evaluate the robustness of IndoBERT-Financial under changes in
linguistic characteristics and label quality, assess the performance of the selected
model on empirical Indonesian economic news, and examine the contribution of
sentiment indices to the prediction of the closing prices of IHSG, BBRI, and TLKM
using Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) models.
The simulation study used 40,000 synthetic economic news texts generated
under four linguistic scenarios with five replications. The evaluation compared
Base and Large architectures under Original and Relabeled label conditions and
two training schemes, namely direct fine-tuning (FT-L) and staged fine-tuning (FTB). Robustness was evaluated based on changes in classification performance and
the consistency of results across conditions, with Macro F1 as the primary metric.
In the empirical study, the selected model was applied to 32,491 economic news
articles from 2021–2025. Class probabilities from the sentiment classification
results were used to construct daily sentiment indices, which were then included as
exogenous variables in LSTM and GRU models for five-step-ahead forecasting
from H+1 to H+5. The contribution of sentiment was assessed by comparing the
errors of models without sentiment and models with sentiment.
The simulation results show that the robustness of IndoBERT-Financial is not
determined solely by model size, but is also associated with label condition, training
scheme, and interactions among factors. Label condition and fine-tuning scheme
had significant effects on Macro F1, while no single architecture or fine-tuning
scheme consistently outperformed the others across all testing conditions. Based
on the overall evaluation, IndoBERT-Financial Base with FT-L was selected for
the empirical study. The selected model achieved a Balanced Accuracy of 0.8510,
a Macro F1 of 0.8439, and a Weighted F1 of 0.8512 on the test data.
Adding the sentiment index did not produce a uniform effect on stock price
prediction. For IHSG and BBRI, the differences in prediction error between models
without sentiment and models with sentiment were not statistically significant. For
TLKM, adding sentiment reduced the LSTM MAPE from 3.0328% to 2.8680% and
produced a significant improvement, whereas in GRU it increased the MAPE from
2.7984% to 2.8661%, with the difference also being significant. These findings
indicate that good sentiment classification performance does not automatically lead
to improved stock price prediction. The usefulness of news sentiment depends on
the aset and the forecasting architecture in which it is incorporated, so sentiment
information is more appropriately treated as supplementary information whose
contribution should be evaluated within the specific modeling context.

