Show simple item record

dc.contributor.advisorHerdiyeni, Yeni
dc.contributor.advisorHardhienata, Medria Kusuma Dewi
dc.contributor.authorZulkifli, Shyfa Kanaya
dc.date.accessioned2026-08-06T07:38:44Z
dc.date.available2026-08-06T07:38:44Z
dc.date.issued2026
dc.identifier.urihttp://repository.ipb.ac.id/handle/123456789/177435
dc.description.abstractPerkembangan Large Language Model (LLM) membuka peluang baru dalam pengembangan Personalized Learning Systems (PLS), namun efektivitas pedagogisnya dalam pembelajaran Computational Thinking (CT) masih belum banyak dievaluasi secara sistematis. Penelitian ini bertujuan menganalisis dan mengevaluasi kerangka pedagogis pemanfaatan LLM untuk mendukung PLS pada pembelajaran CT. Penelitian ini menggunakan pendekatan evaluatif kuantitatif dengan dataset sintetis yang terdiri atas 144 respons, merupakan hasil kombinasi berbagai level pedagogis dan profil kognitif pengguna terhadap soal CT jenjang SD, SMP, dan SMA. Agen LLM menghasilkan 144 respons berbasis model Gemma 3 27B. Seluruh respons dievaluasi menggunakan metode LLM-as-a-Judge (GPT-5.2), lalu divalidasi oleh pakar manusia (human expert judges) terhadap 48 sampel terpilih menggunakan purposive sampling. Hasil penelitian menunjukkan bahwa seluruh respons yang dihasilkan sesuai dengan profil pengguna tanpa ditemukan respons kosong maupun tidak relevan. Evaluasi LLM-as-a-Judge menghasilkan 90,97% respons dengan skor pedagogis tertinggi pada keseluruhan dataset. Analisis kesesuaian antara LLM-as-a-Judge dan human expert judges pada 48 sampel menunjukkan tingkat kesepakatan yang tinggi, khususnya pada level pedagogis dasar hingga menengah. Hasil evaluasi mencatatkan tingkat akurasi berkisar antara 70,80%-95,80% dan koefisien Gwet's AC2 antara 0,65-0,95, dengan variasi yang dipengaruhi oleh acuan evaluator manusia. Temuan ini menunjukkan bahwa integrasi profil kognitif dan level pedagogis pada LLM mampu menghasilkan pembelajaran yang adaptif. Selain itu, penelitian ini juga menunjukkan bahwa pendekatan LLM-as-a-Judge berpotensi menjadi mekanisme evaluasi pedagogis yang reliabel dan memiliki skalabilitas tinggi, meskipun tetap memerlukan keterlibatan evaluator manusia untuk konteks penilaian dengan kompleksitas dan abstraksi tinggi.
dc.description.abstractThe advancement of Large Language Models (LLMs) has opened new opportunities for the development of Personalized Learning Systems (PLS). However, their pedagogical effectiveness in Computational Thinking (CT) education has not yet been systematically evaluated. This study aims to analyze and evaluate a pedagogical framework for leveraging LLMs to support PLS in CT learning. A quantitative evaluative approach was employed, utilizing a synthetic dataset of 144 responses generated from combinations of various pedagogical levels and cognitive profiles applied to CT problems spanning elementary, junior high, and senior high school levels. Responses were generated by an LLM agent using the Gemma 3 27B model, subsequently evaluated through an LLM-as-a-Judge method employing GPT-5.2 across the entire 144 responses, and further validated by human expert judges on a purposively selected subset of 48 samples. The results indicate that all generated responses were consistent with their corresponding user profiles, with no empty or irrelevant responses observed. The LLM-as-a-Judge evaluation yielded the highest pedagogical scores for 90.97% of the responses across the full dataset. The agreement analysis between LLM-as-a-Judge and human expert judges on the 48-sample subset revealed varying levels of agreement depending on which human evaluator served as the reference standard, with accuracy ranging from 70.8% to 95.8% and Gwet's AC2 coefficients ranging from 0.655 to 0.953, indicating an overall high level of alignment between LLM-as-a-Judge and human expert judgments, particularly at lower to intermediate pedagogical levels. These findings suggest that the integration of cognitive profiles and pedagogical levels within LLMs can produce adaptive learning experiences. Furthermore, this study demonstrates that the LLM-as-a-Judge approach holds strong potential as a reliable and scalable mechanism for pedagogical evaluation, although human evaluator involvement remains necessary for assessment contexts involving high complexity and abstraction.
dc.description.sponsorship
dc.language.isoid
dc.publisherIPB Universityid
dc.titleEvaluasi Penggunaan Large Language Model (LLM) dalam Pengembangan Personalized Learning System (PLS) dengan Kerangka Pedagogisid
dc.title.alternativeEvaluation of Large Language Model (LLM) Utilization in the Development of a Personalized Learning System (PLS) Using a Pedagogical Framework
dc.typeSkripsi
dc.subject.keywordComputational Thinkingid
dc.subject.keywordEvaluasi Pedagogisid
dc.subject.keywordLarge Language Modelid
dc.subject.keywordPembelajaran Adaptifid
dc.subtypeUndergraduate Theses


Files in this item

Thumbnail
Thumbnail
Thumbnail

This item appears in the following Collection(s)

Show simple item record