Abstract:Abstract: Gross Regional Domestic Product (GRDP) is one of the most important socio-economic indicators. In order to gain a more comprehensive understanding of the current economic situation and regional differences, estimating…
imating GRDP using integration of satellite imagery and official statistics data can provide valuable information. This research estimates the GRDP value in 2022 by using data in 2019 to 2021 related to two aspects, agriculture and non-agriculture. Soil adjusted vegetation index (SAVI), enhanced vegetation index (EVI), and land cover (LC) used as agriculture aspect, while nighttime light (NTL), human settlement index (HSI), land area, and population per regency/city used as non-agriculture aspect. GRDP estimation are produced with machine learning approach using support vector machine (SVM) and random forest (RF) method. Correlation test on each variable shows only land area that does not have a significant correlation with GRDP. RF model then chosen as the best model with RMSE, MSE, MAE, and R2 value of 0.2549; 0.5049; 0.7727; and 0.2543, respectively. The estimated values acquired in several regencies/cities have rather near, some even very close to the official statistics values.
Keywords: GRDP; satellite imagery; machine learning; random forest; support vector machine
Abstrak: Produk Domestik Regional Bruto (PDRB) merupakan salah satu indikator sosio-ekonomi yang penting. Penghitungan nilai PDRB dengan pendekatan yang melibatkan kombinasi data citra satelit dan statistik resmi dapat memberikan informasi serta pemahaman yang lebih komprehensif. Penelitian ini melakukan estimasi nilai PDRB pada tahun 2022 menggunakan data tahun 2019 hingga 2021 dengan melibatkan dua aspek, agrikultur dan non-agrikultur. Data soil adjusted vegetation index (SAVI), enhanced vegetation index (EVI), dan tutupan lahan (land cover/LC) digunakan sebagai aspek agrikultur, sementara data citra cahaya malam (NTL), human settlement indeks (HSI), luas wilayah kabupaten/kota, dan jumlah populasi per kabupaten/kota digunakan sebagai aspek non-agrikultur. Estimasi PDRB dihasilkan dengan menggunakan pendekatan machine learning berupa support vector machine (SVM) dan random forest (RF). Pengecekan korelasi antarvariabel menunjukkan bahwa hanya variabel luas wilayah tidak berpengaruh signifikan terhadap nilai PDRB. Model random forest kemudian dipilih sebagai model terbaik dengan nilai evaluasi RMSE, MSE, MAE, dan berturut-turut sebesar 0.2549, 0.5049, 0.7727, dan 0.2543. Nilai estimasi yang diperoleh di beberapa kabupaten/kota cukup mendekati, bahkan ada yang sangat dekat dengan nilai statistik resmi.
Kata kunci: PDRB; citra satelit; machine learning; random forest; support vector machine
Abstract:Abstract: There are a great number of academics that are now conducting research on sentiment analysis by employing supervised and machine learning techniques. The research can be carried out with the assistance of a variety…
iety of sources, including reviews of movies, reviews of Twitter, reviews of online products, blogs, discussion forums, and other social networks. With the progress of technology, individuals may now effortlessly utilize social media platforms to access and share information, as well as express their viewpoints to the general public, without any constraints of distance or time. Twitter is a social media network that serves as a repository for opinions. Diverse techniques are employed to provide optimal and realistically precise pressure detection. The analysis and discussion affirm that the Support Vector Machine (SVM) was effectively employed in this study, utilizing public opinion data on television program reviews in Indonesia. An SVM classifier is employed to examine the Twitter data set by utilizing various parameters. The study successfully completed the preprocessing process by collecting a total of 400 data points, consisting of 320 reviews from 4 television shows for training data and 80 reviews for testing. The data was filtered and classified using SVM, with 200 positive and 200 negative data points for comparison. The experiment utilized the SVM method using TF-IDF to achieve the most accurate test results. The test accuracy was 80%, while the training data accuracy reached 100%.
Keywords: Sentiment Analysis; Support Vector Machine; Television Shows Review, TF-IDF,
Abstrak: Saat ini, banyak akademisi sedang menyelidiki analisis sentimen melalui pemanfaatan teknik yang diawasi dan pembelajaran mesin. Kajian dapat dilakukan dengan menggunakan beberapa sumber seperti review film, review Twitter, review produk online, blog, forum diskusi, atau jejaring sosial lainnya. Dengan kemajuan teknologi, masyarakat kini dapat dengan mudah memanfaatkan platform media sosial untuk mengakses dan berbagi informasi, serta menyampaikan pandangan mereka kepada masyarakat umum, tanpa batasan jarak dan waktu. Twitter adalah jaringan media sosial yang berfungsi sebagai gudang opini. Beragam teknik digunakan untuk menghasilkan deteksi tekanan yang optimal dan presisi secara realistis. Analisis dan pembahasan menegaskan bahwa Support Vector Machine (SVM) efektif digunakan dalam penelitian ini, memanfaatkan data opini publik tentang review program televisi di Indonesia. Pengklasifikasi SVM digunakan untuk memeriksa kumpulan data Twitter dengan memanfaatkan berbagai parameter. Penelitian berhasil menyelesaikan proses preprocessing dengan mengumpulkan total 400 titik data yang terdiri dari 320 review dari 4 acara televisi untuk data pelatihan dan 80 review untuk pengujian. Data disaring dan diklasifikasikan menggunakan SVM, dengan 200 titik data positif dan 200 titik data negatif sebagai perbandingan. Percobaan ini menggunakan metode SVM dengan menggunakan TF-IDF untuk mencapai hasil pengujian yang paling akurat. Akurasi pengujiannya mencapai 80%, sedangkan akurasi data pelatihan mencapai 100%.
Kata kunci: Analisis Sentimen, Review Tayangan Televisi, TF-IDF, Support Vector Machine
Abstract:Abstract: Since information and communication technology has become ingrained in our daily lives, it has become easier to access information. However, there are some concerns. One of them is about fake news. The aim of this…
his study is to develop an Indonesian system for detecting false news by utilizing news headlines. The methods used are linear kernel support vector ma- chine and n-gram. According to the findings of the performance test that was carried out, the linear kernel support vector machine model employing the term frequency inverse document frequency unigram feature performs better than utilizing bigram. The precision value generated from the model performance test is 1.00. This means that the degree of accuracy in matching the requested information regarding fake news detection with the answers provided by the system is very good. Then the recall value generated is 0.99. This means the linear kernel support vector machine model using unigram news features is very effective for detecting fake news according to the text classification approach.
Keywords: classification; fake news; n-gram; support vector machine
Abstrak: Dengan adanya integrasi teknologi informasi dan komunikasi dalam kehidupan mem- buat kemudahan dalam mengakses informasi. Walaupun demikian, terdapat kekhawatiran akan beberapa hal. Salah satu di antaranya adalah berita palsu. Tujuan penelitian ini adalah merancang sistem deteksi berita palsu berbahasa Indonesia berdasarkan judul berita. Metode yang digunakan adalah Support Vector Machine kernel linier dan n-gram. Berdasarkan hasil uji performa, model Support Vector Machine kernel linier yang menggunakan fitur term frequency inverse document frequency unigram menunjukkan kinerja yang lebih baik dibandingkan bi- gram. Nilai precision yang dihasilkan dari uji performa model sebesar 1,00. Ini berarti derajat akurasi dalam mencocokkan informasi yang diminta mengenai deteksi berita palsu dengan ja- waban yang diberikan oleh sistem sangat baik. Kemudian nilai recall yang dihasilkan sebesar 0,99. Ini berarti model Support Vector Machine kernel linier dengan menggunakan fitur berita unigram sangat efektif untuk mendeteksi berita palsu menurut pendekatan teks klasifikasi.
Kata kunci: klasifikasi; berita palsu; n-gram; support vector machine
Abstract:Abstract: The JKN Mobile application is a mobile application created to facilitate healthcare administration in Indonesia since 2017. The application has been downloaded by over 10 million users and has received 484,000…
diverse reviews, including positive, negative, and neutral feedback. The average rating given by users is 4.5 out of 5 stars. This research aims to perform sentiment analysis on user reviews found in the Google Play Store review column. The methods used for sentiment analysis are Naive Bayes, K-Nearest Neighbor (K-NN), and Support Vector Machine (SVM). The test results show that with a 10% test data and 90% training data proportion, the SVM method achieves the highest accuracy of 95%. Naive Bayes follows with an accuracy of 87%, and K-NN with an accuracy of 75%.
Keywords: JKN mobile application, sentiment analysis, naive bayes, k-nearest neighbor (K-NN), support vector machine (SVM).
Abstrak: Aplikasi Mobile JKN adalah sebuah aplikasi yang dibuat untuk mempermudah administrasi kesehatan di Indonesia sejak tahun 2017. Aplikasi ini telah diunduh lebih dari 10 juta pengguna dengan 484 ribu ulasan beragam positif, negatif, dan netral. Rata-rata rating yang diberikan pengguna adalah 4,5 bintang dari 5 bintang. Penelitian ini bertujuan untuk melakukan analisis sentimen terhadap ulasan pengguna yang terdapat di kolom review Google Play Store. Metode yang digunakan untuk analisis sentimen adalah Naive Bayes, K-Nearest Neighbor (K-NN), dan Support Vector Machine (SVM). Hasil pengujian menunjukkan bahwa dengan menggunakan proporsi data uji sebesar 10% dan data training sebesar 90%, metode SVM mencapai akurasi tertinggi sebesar 95%. Diikuti oleh Naive Bayes dengan akurasi 87%, dan K-NN dengan akurasi 75%.
Kata kunci: JKN mobile, analisis sentimen, naïve bayes, k-nearest neighbor (K-NN), support vector machine (SVM).
Abstract:Abstract: Glaucoma is the second leading eye disease of blindness after cataracts. An ophthalmologist does a glaucoma examination with an eye screening that will produce a retinal image. The diagnosis's result of the retinal…
inal image is subjective because each doctor has dissent and a condition experienced. This research builds a system to identify retinal images in the category of glaucoma or normal patients. The purpose of this system as a tool is to help ophthalmologists diagnose glaucoma. This process begins by changing the colour of an image to validity. The image is extracted using the Gray Level Co-occurrence Matrix (GLCM) and produces five features than the result of the five features used as input to neural network Learning Vector Quantization (LVQ). The amount of retinal image data used 60 data for learning and 20 for testing. And the number of neurons used is 12, and the epoch of as many as 1000 was obtained based on the results comparison of variations against 8, 10, 12, 18, and 20 neurons with 500, 900, 1000, and 1100 epochs. The learning process results from the value of weights that will be used in the testing process. The results of this study obtained an accuracy rate of 85%, a precision of 89%, and a recall of 80%.
Keywords: glaucoma; LVQ; GLCM
Abstrak: Glaukoma merupakan penyakit mata penyebab kebutaan nomor dua setelah katarak. Pemeriksaan penyakit glaukoma dilakukan oleh dokter spesialis mata dengan cara skrining mata yang akan menghasilkan citra retina. Hasil diagnosa citra retina oleh dokter bersifat subjektif karena setiap dokter memiliki pendapat yang berbeda serta kondisi yang dialami. Penelitian ini membangun sistem untuk mengidentifikasi citra retina mata kedalam kategori penderita glaukoma atau normal. Tujuan pembuatan sistem ini sebagai alat bantu bagi dokter mata dalam mendiagnosa glukoma. Proses ini diawali dengan mengubah warna citra menjadi keabuan, kemudian citra tersebut diekstraksi menggunakan Gray Level Co-occurrence Matrix (GLCM) dan menghasilkan lima nilai fitur yang kemudian hasil dari kelima fitur tersebut digunakan sebagai masukan jaringan syaraf tiruan Learning Vector Quantization (LVQ). Jumlah data citra retina yang digunakan sebanyak 60 data untuk learning dan 20 data untuk testing. Dan untuk jumlah neuron yang digunakan yaitu 12 dan epoch sebanyak 1000 yang didapat berdasarkan hasil perbandingan variasi terhadap neuron sejumlah 8, 10, 12, 18, 20 dengan jumlah epoch sejumlah 500, 900, 1000 dan 1100. Hasil dari proses learning yaitu nilai bobot yang akan digunakan sebagai bobot di proses testing. Hasil penelitian ini diperoleh tingkat akurasi sebesar 85%, presisi sebesar 89% dan recall sebesar 80%.
Kata kunci: glaukoma; LVQ; GLCM
Abstract:Abstract: Non-performing loan (NPL) is a risk that credit unions must face and to avoid that, prospective debtors need to be surveyed. With previous loan data, support vector machine and naïve bayes can be used as classification…
ssification methods to give a decision about NPL. We use a data set with 61 data and process the data with orange 3.30 application to see the difference between SVM using linear (SVM-L), polynomial (SVM-P), RBF (SVM-R) and sigmoid (SVM-S) kernel with naïve bayes. We use a cross validation technique with various folds to measure the classification results and a convusion matrix to measure the data training classification results. Naïve bayes scores the highest in terms of accuracy and SVM-R scores the highest in terms of F1, precision and recall. SVM-P scores the lowest in terms of accuracy, F1, precision and recall. Naïve bayes scores the highest in terms of proportion of predicted for true negative class and proportion of actual for true positive class. SVM-S scores the highest in terms of proportion of predicted for true positive class and proportion of actual for true negative class. SVM-P scores the lowest in both proportion of predicted and proportion of actual.
Keywords: classification; naïve bayes; non-performing loan; support vector machine
Abstrak: Kredit macet merupakan resiko yang sering dialami koperasi simpan pinjam, sehingga perlu dilakukan survei terhadap calon debitur agar kredit menjadi sehat. Dengan menggunakan data pemberian kredit sebelumnya, support vector machine dan naïve bayes digunakan sebagai metode klasifikasi untuk memberikan keputusan macet atau tidaknya kredit anggota koperasi Mutiara Sejahtera. Data set yang berjumlah 61 data diolah menggunakan aplikasi Orange 3.30 dan dilihat perbandingan antara metode SVM dengan kernel linear, polynomial, RBF dan sigomoid dengan metode naïve bayes. Cross validation dengan jumlah fold bervariasi digunakan sebagai nilai ukur klasifikasi dan convusion matrix digunakan sebagai nilai ukur klasifikasi data training. Hasil yang diperoleh adalah naïve bayes memiliki nilai accuracy tertinggi dan SVM kernel RBF memiliki nilai F1, precision dan recall tertinggi. SVM kernel polynomial memiliki nilai terendah untuk accuracy, F1, precision dan recall. Naïve bayes memiliki nilai tertinggi untuk proportion of predicted (PoP) kelas true negative dan proportion of actual (PoA) kelas true positive. SVM kernel sigmoid memiliki nilai tertinggi untuk PoP kelas true positive dan PoA kelas true negative. SVM kernel polynomial memiliki nilai terendah baik untuk PoP maupun PoA true negative dan kelas true positive.
Kata kunci: klasifikasi; kredit macet; naive bayes; SVM
Abstract:Abstract: The technique of multiclass classification based on SVMs has been widely used. SVM optimization will be accomplished by examining the extraction features of Principal Component Analysis (PCA), Box-Cox Transformation,…
ation, and Recursive Feature Elimination (RFE). The dataset contains 13,611 rows and 17 variables, generated from the UCI repository's multiclass dry bean data. Barbunya, Bombay, Cal, Dermas, Horoz, Seker, and Sira are just a few of the dry bean kinds available. The dataset was tested using SVM Linear kernel and SVM Radial Basis.According to the results, the combination of scale-center-BoxCox-SVM Radial extraction achieves the maximum accuracy of 93.16 percent and the shortest processing time of 6.10 minutes. 96.00 percent, 100 percent, 96.71 percent, 95.16 percent, 97.60 percent, 97.74 percent, and 91.95 percent, according to bean class.RFE-SVM Radial has a 91.18 percent accuracy and a processing time of 6.55 minutes. BoxCox outperforms conventional techniques in terms of prediction accuracy while requiring less training time.
Keywords: Bean, PCA, BoxCox, SVM, RFE
Abstrak: Klasifikasi Multikelas menggunakan SVM telah banyak digunakan. Pada penelitian ini akan diuji fitur ekstraksi Principal Component Analysis, Box Cox Transformation dan fitur eliminisi Recursive Feature Elimination untuk mendapatkan optimasi SVM. Dataset berasal dari data multikelas kacang kering UCI repository dengan jumlah 13.611 baris dan 17 variabel. Kelas kacang kering yakni : Barbunya, Bombay, Cal, Dermas, Horoz, Seker dan Sira. Dataset diuji menggunakan kernel SVM Linier dan SVM Radial Basis. Didapatkan hasil, bahwa kombinasi fitur ekstraksi : scale-center-BoxCox-SVM Radial memiliki akurasi terbaik yakni 93,16% dan waktu proses 6,10 menit. Klasifikasi berdasarkan kelas kacang berturut-turut 96,00%,100%, 96,71%, 95,16%, 97,60%, 97,74% dan 91,95%. RFE- SVM Radial hanya memberikan akurasi sebesar 91,18 % dengan waktu proses sebesar 6.55 menit. Penggunaan BoxCox dibandingkan dengan lainnya, memberikan hasil prediksi lebih baik dan namun tidak mempercepat waktu pelatihan.
Kata kunci: BoxCox; Kacang; PCA; RFE; SVM
Abstract:Abstract: Question classification is a computer science system, which aims to analyze questions and can label each question based on existing categories. Questions can be collected from several materials or topics that are…
re many and different. Therefore, the researcher intends to create a classification system for quiz questions Data Warehouse and Business Intelligence which can be grouped into topics Data Warehouse, Business Intelligence, Data Analytics, and Performance Measurement. One way to solve this problem is by approach machine learning. In this study, researchers used a comparison of machine learning algorithms, namely the algorithm NaïveBayes and SupportVectorMachine using SMOTE and methods Cross-Validation The results of this study show the best accuracy results and are very helpful. The results obtained in the method cross-validation before SMOTE resulted in an accuracy rate of 82.02% for the results after going through the SMOTE stage of 94.79% on the algorithm Naïve Bayes, while the algorithm SupportVectorMachine get accuracy of 81.39% in the process before SMOTE for the results after going through SMOTE of 96.52%.
Keywords: Cross-Validation; Machine Learning; Naive Bayes; Support Vector Machine; Question Classification
Abstrak: Klasifikasi pertanyaan merupakan sebuah sistem ilmu komputer, yang bertujuan untuk menganalisis pertanyaan serta dapat memberi label pada setiap pertanyaan berdasarkan kategori yang ada. Pertanyaan soal dapat dikumpulkan dari beberapa materi atau topik yang banyak dan berbeda. Oleh karena itu, bermaksud untuk membuat sistem klasifikasi pertanyaan soal kuis Data Warehouse dan Business Intelligence yang dapat dikelompokkan menjadi topik Data Warehouse, Business Intelligence, Data Analitik, dan Pengukuran Kinerja. Cara yang dapat dilakukan untuk permasalahan ini dengan menggunakan pendekatan MachineLearning. Pada penelitian kali ini menggunakan perbandingan algoritma MachineLearning yaitu algoritma NaïveBayes dan SupportVectorMachine menggunakan metode SMOTE dan Cross-Validation. Hasil penelitian ini menunjukkan hasil akurasi yang terbaik dan sangat membantu. Hasil yang diperoleh pada metode cross-validation sebelum SMOTE menghasilkan tingkat akurasi sebesar 82.02% untuk hasil sesudah melalui tahap SMOTE sebesar 94.79 % pada algoritma Naïve Bayes, sedangkan pada algoritma Support Vector Machine menghasilkan akurasi sebesar pada proses sebelum SMOTE 81.39% untuk hasil sesudah melalui SMOTE sebesar 96.52%.
Kata kunci: Klasifikasi Pertanyaan; Pembelajaran Mesin; Naive Bayes; Support Vector Machine; Cross-Validation
Abstract:Abstract: The human eye can distinguish objects from digital images, however, computers do not have the ability as human eyes that can directly distinguish objects from digital images. Therefore the bag of visual words method…
ethod was created. Bag of visual words is a method for presenting digital images based on local features. Bag of visual words illustrates how an image can be taken its characteristics, so that computers can distinguish objects on digital images. The test results show that the bag of visual words are still not maximal in classifying digital image categories, especially the chair category, which is only able to produce the most accurate accuracy of 75%. To improve the performance quality of bag of visual words in classifying digital image categories, especially the chair category, you can add an approach to determine the good number of K in clustering the visual words pattern.
Keywords: Bag Of Visual Words, Classification, Digital Image, Speed-Up Robust Feature, Support Vector Machine
Abstrak: Secara kasat mata manusia bisa membedakan objek pada citra digital, namun, komputer tidak memiliki kemampuan sebagai mata manusia yang dapat secara langsung membedakan objek pada citra digital. Maka dari itu diciptakanlah metode bag of visual words. Bag of visual words adalah metode untuk menyajikan citra digital berdasarkan fitur lokal. Bag of visual words menggambarkan bagaimana suatu gambar dapat diambil karakteristiknya, sehingga komputer dapat membedakan objek pada citra digital. Hasil pengujian menunjukkan bag of visual words masih belum maksimal dalam mengklasifikasi kategori citra digital khususnya kategori chair, yang hanya mampu menghasilkan akurasi paling akurat sebesar 75 %. Untuk meningkatkan kualitas kinerja bag of visual words dalam mengklasifikasi kategori citra digital khususnya kategori chair, dapat menambahkan pendekatan untuk menentukan jumlah K yang baik dalam mengkluster pola visual words.
Kata kunci: Bag Of Visual Words, Klasifikasi, Citra Digital, Speed-Up Robust Feature, Support Vector Machine
Abstract:Meat is one of the essential food ingredients in meeting the nutritional needs. The current problem lies in the consumers' lack of knowledge on how to differentiate between pork, beef, goat, and lamb meat. This is because…
e when the meat is already cut, their appearances may seem similar at first glance. Many consumers are unaware of the practice of mixing different types of meat for consumption. One way to classify animal meat is by using image processing. In this research, an image processing system is created to classify meat, specifically pork, beef, goat, and lamb. Support Vector Machine (SVM) is a development of Machine Learning that can be used in classifying images into specific classes. SVM method as a classifier is performed using a confusion matrix. The test results show the highest accuracy value obtained in the class of Goat Meat 91.4%, the highest precision in the class of goat meat 80%, the highest recall in the class of beef 81.3%, and the highest F1-score in the class of beef 0.76.