Analisis Pengaruh Pengurangan Fitur Berkorelasi Tinggi serta RFE terhadap Klasifikasi Kanker Payudara dengan SVM dan Random Forest
Date
2026Jenis/Type
SkripsiSubtype
Undergraduate ThesesAuthor
Rachmalia, Adinda
Julianto, Mochamad Tito
Khatizah, Elis
Metadata
Show full item recordAbstract
Kanker payudara merupakan salah satu jenis kanker dengan tingkat kejadian yang tinggi pada perempuan sehingga diperlukan metode klasifikasi yang akurat untuk mendukung diagnosis dini. Penelitian ini bertujuan menganalisis pengaruh penghapusan fitur berkorelasi tinggi berdasarkan nilai Variance Inflation Factor (VIF) serta penggunaan Recursive Feature Elimination (RFE) terhadap klasifikasi kanker payudara menggunakan algoritma Support Vector Machine (SVM) dan random forest. Dataset yang digunakan adalah Wisconsin Breast Cancer Diagnostic dari UCI Machine Learning Repository. Penelitian dilakukan menggunakan empat skenario, yaitu baseline, penghapusan fitur berkorelasi tinggi berdasarkan nilai VIF, RFE dengan 15 fitur terpilih, serta kombinasi VIF dan RFE. Hasil penelitian menunjukkan bahwa kombinasi VIF dan RFE menghasilkan performa terbaik pada model SVM menggunakan 15 fitur terpilih, sedangkan model random forest menghasilkan performa terbaik pada skenario RFE tanpa VIF. Secara umum, penghapusan dan seleksi fitur mampu mengurangi jumlah fitur dari 30 menjadi 15 tanpa menurunkan performa klasifikasi secara signifikan. Breast cancer is one of the most common cancers among women, making accurate classification methods essential for supporting early diagnosis. This study aims to analyze the effect of removal of highly correlated features based on Variance Inflation Factor (VIF) values and the application of Recursive Feature Elimination (RFE) on breast cancer classification using Support Vector Machine (SVM) and random forest algorithms. The dataset used in this study was the Wisconsin Breast Cancer Diagnostic obtained from the UCI Machine Learning Repository. Four experimental scenarios were evaluated, namely the baseline scenario, removal of highly correlated features based on VIF values, RFE with 15 selected features, and a combination of VIF and RFE. The results showed that the combination of VIF and RFE produced the best performance for the SVM model using 15 selected features, whereas the random forest model achieved its best performance under the RFE-only scenario. Overall, feature elimination and feature selection reduced the number of features from 30 to 15 without significantly decreasing classification performance.
Collections
- UF - Mathematics [164]

