Optimasi Arsitektur ANN menggunakan Algoritma Genetika untuk Prediksi Produksi Pertanian Berbasis Data Sintetis CTGAN
Abstract
Keterbatasan amatan pada pemodelan prediksi produksi pertanian rentan memicu overfitting dan menyulitkan pembelajaran pola interaksi kompleks. Untuk mengatasinya, Conditional Tabular Generative Adversarial Network (CTGAN) diterapkan sebagai metode sintesis data, dipadukan dengan Artificial Neural Network (ANN) yang pengoptimalan hyperparameter-nya dilakukan oleh Genetic Algorithm (GA) guna memodelkan hubungan nonlinier. Penelitian ini bertujuan membentuk sintetis menggunakan CTGAN, membangun model prediksi produksi padi menggunakan ANN-GA berbasis data sintetis dan membandingkannya dengan regresi linier berganda, serta mengidentifikasi pengaruh peubah penjelas terhadap produksi padi menggunakan kerangka Leave-One-Out Cross-Validation (LOOCV). Dari tujuh skenario arsitektur CTGAN yang diuji, skenario dengan hyperparameter seluruhnya bernilai besar paling sering menghasilkan data sintetis terbaik berdasarkan nilai JS Distance dan selisih korelasi Spearman. Model regresi linier berganda menghasilkan MAPE LOOCV sebesar 13,37%, lebih rendah dibandingkan ANN-GA sebesar 19,95%. Namun model regresi hanya dapat mempertahankan dua peubah penjelas dan tidak selalu dapat memenuhi asumsi parametrik, sehingga inferensi pengaruh peubah penjelas pada kasus ini cenderung terbatas. Di sisi lain, model ANN-GA hanya menghasilkan satu titik pencilan pada sebaran APE dan mengakomodasi seluruh peubah penjelas yang digunakan. Analisis Permutation Feature Importance (PFI) pada ANN-GA mengidentifikasi dampak perubahan iklim, pengendalian organisme pengganggu tanaman, dan Normalized Difference Vegetation Index (NDVI) sebagai peubah paling berpengaruh terhadap produksi padi. Limited observations in agricultural production predictive modeling are highly susceptible to triggering overfitting and complicate the learning of complex interaction patterns. To address this, a Conditional Tabular Generative Adversarial Network (CTGAN) was applied for data synthesis, combined with an Artificial Neural Network (ANN) optimized by a Genetic Algorithm (GA) to model nonlinear relationships. This study aims to generate synthetic data using CTGAN, develop a synthetic data-based ANN-GA rice production prediction model to compare against multiple linear regression, and identify explanatory variables' effects using a Leave-One-Out Cross-Validation (LOOCV) framework. Among seven tested CTGAN architectures, the scenario with exclusively large hyperparameters most frequently produced the best synthetic data based on JS Distance and Spearman correlation differences. The multiple linear regression yielded a LOOCV MAPE of 13.37%, lower than the ANN-GA's 19.95%. However, the regression model only retained two explanatory variables and could not consistently fulfill parametric assumptions, thereby limiting inference. Conversely, the ANN-GA model produced only one outlier in the APE distribution and accommodated all explanatory variables. Permutation Feature Importance (PFI) analysis on the ANN-GA identified climate change impacts, plant pest control, and the Normalized Difference Vegetation Index (NDVI) as the most influential variables on rice production.

