Show simple item record

dc.contributor.advisorHerdiyeni, Yeni
dc.contributor.advisorHardhienata, Medria Kusuma Dewi
dc.contributor.authorDjunaid, Muhammad Quwwamul Haq
dc.date.accessioned2026-08-14T06:20:55Z
dc.date.available2026-08-14T06:20:55Z
dc.date.issued2026
dc.identifier.urihttp://repository.ipb.ac.id/handle/123456789/178723
dc.description.abstractPerkembangan Large Language Model (LLM) mendorong pemanfaatannya dalam Personalized Learning System untuk mendukung pembelajaran yang lebih adaptif. Namun, kesesuaian respons LLM terhadap karakteristik kognitif pengguna masih belum banyak dievaluasi. Penelitian ini bertujuan mengevaluasi kinerja LLM dalam menghasilkan respons yang sesuai dengan karakteristik kognitif pengguna pada sistem LogiCT berbasis Computational Thinking. Penelitian menggunakan pendekatan evaluatif kuantitatif dengan metode LLM-as-a-Judge dan dua Human Expert Judges. Evaluasi dilakukan terhadap 48 respons yang mewakili delapan tipe kognitif dan enam level pedagogis. Respons dihasilkan menggunakan Gemma 3 27B dan dievaluasi oleh GPT-5.2 melalui OpenRouter API menggunakan skala 0–2. Hasil evaluasi menunjukkan bahwa kesesuaian LLM-as-a-Judge terhadap Human Expert Judge 1 dan Human Expert Judge 2 menghasilkan accuracy masing-masing sebesar 85,42% dan 68,75%, Weighted F1-Score sebesar 0,82 dan 0,73, Weighted Kappa sebesar 0,17 dan 0,14, serta Gwet's AC2 sebesar 0,96 dan 0,85. Hasil tersebut menunjukkan bahwa LogiCT mampu menghasilkan respons yang sesuai dengan karakteristik kognitif pengguna.
dc.description.abstractThe rapid advancement of Large Language Models (LLMs) has driven their adoption in Personalized Learning Systems to support more adaptive learning. However, the alignment of LLM-generated responses with users' cognitive characteristics has received limited attention. This study evaluates the performance of LLM-generated responses in the LogiCT system for Computational Thinking using a quantitative evaluative approach with an LLM-as-a-Judge and two Human Expert Judges. A total of 48 responses representing eight cognitive types across six pedagogical levels were generated using Gemma 3 27B and evaluated by GPT-5.2 via the OpenRouter API on a 0–2 scale. The agreement between the LLM-as-a-Judge and Human Expert Judges achieved accuracy values of 85.42% and 68.75%, Weighted F1-Scores of 0.82 and 0.73, Weighted Kappa values of 0.17 and 0.14, and Gwet's AC2 values of 0.96 and 0.85, respectively. These findings indicate that LogiCT generally produces responses aligned with users' cognitive characteristics.
dc.description.sponsorship
dc.language.isoid
dc.publisherIPB Universityid
dc.titleEvaluasi Kinerja Large Language Model pada Personalized Learning System Berdasarkan Kerangka Kognitifid
dc.title.alternativePerformance Evaluation of Large Language Models in Personalized Learning Systems Based on a Cognitive Framework
dc.typeSkripsi
dc.subject.keywordCognitive frameworkid
dc.subject.keywordComputational Thinkingid
dc.subject.keywordLarge Language Modelid
dc.subject.keywordLLM-as-a-Judgeid
dc.subject.keywordPersonalized Learning Systemid
dc.subtypeUndergraduate Theses


Files in this item

Thumbnail
Thumbnail
Thumbnail

This item appears in the following Collection(s)

Show simple item record