Search Articles & Publications

Showing 73 articles found for "Classification"

OPTIMIZATION OF CART ALGORITHM BASED ON ANT BE COLONY FEATURE SELECTION FOR STUNTING DIAGNOSIS

Subarkah, Pungkas, Ikhsan, Ali Nur, Wahyudi, Rizki, Rofiqoh, Dayana
Abstract: Abstract: One of the main health problems in children is stunting which is one of the concerns in the Sustainable Development Goals (SDGs). Specifically in Indonesia, the prevalence of stunting in 2024 is 21.6%. This figure… ure is still relatively high, because the target prevalence of stunting is 14%. This study aims to implement machine learning knowledge through the Classification And Regression Trees (CART) algorithm based on Ant Be Colony (ABC) feature selection which aims to determine the increase in accuracy in analyzing stunting datasets. The data used comes from Kaggle which consists of 16500 datasets. The dataset consists of gender, age, birth length, birth weight, body length, body weight, breastfeeding and stunting status. The research methods used are data collection, data preprocessing, classification, and evaluation using K-fold cross validation. The results obtained in this research are the implementation of the CART algorithm obtained a value of 89.86% and the results of CART with Ant Be Colony (ABC) feature selection, which obtained an accuracy value of 93.65%. This shows that there is an increase in the accuracy value in the use of CART algorithm optimization and Ant Be Colony (ABC) feature selection by 3.76%. With the research results that have been obtained, it can be categorized as excellent accuracy value excellent. It is hoped that further research can be carried out by adding other classification algorithms or adding feature selection.             Keywords: classification; feature selection; optimazation; stunting   Abstrak: Salah satu masalah kesehatan utama pada anak adalah stunting yang menjadi salah satu perhatian dalam Sustainable Development Goals (SDGs). Khusus di Indonesia angka Pravelensi stunting pada tahun 2024 di angka 21.6%. Angka ini masih tergolong tinggi, karena target angka pravelensi stunting ialah 14%. Penelitian ini bertujuan untuk mengimplementasikan pengetahuan machine learning melalui algoritma Classification And Regression Trees (CART) berbasis seleksi fitur Ant Be Colony (ABC) yang bertujuan untuk mengetahui peningkatan akurasi dalam menganalisis dataset stunting. Data yang digunakan bersumber dari Kaggle yang terdiri dari 16500 dataset. Dataset terdiri dari jenis kelamin, usia, panjang lahir, berat lahir, panjangg badan, berat badan, menyusui dan status stunting.  Metode penelitian yang digunakan adalah pengumpulan data, preprocessing data, klasifikasi, dan evaluasi menggunakan K-fold cross validation. Hasil yang diperoleh pada penelitian ini adalah Implementasi algoritma CART memperoleh nilai sebesar 89,86% dan hasil seleksi fitur CART dengan Ant Be Colony (ABC) memperoleh nilai akurasi sebesar 93,65%. Hal ini menunjukkan adanya peningkatan nilai akurasi pada penggunaan optimasi algoritma CART dan pemilihan fitur Ant Be Colony (ABC) sebesar 3,76%. Dengan hasil penelitian yang telah diperoleh dapat dikategorikan nilai akurasi yang diperoleh sangat baik. Diharapkan dapat dilakukan penelitian selanjutnya dengan menambahkan algoritma klasifikasi lain atau menambahkan seleksi fitur.   Kata kunci: klasifikasi; optimalisasi; seleksi fitur; stunting

APPLICATION OF PARTICLE SWARM OPTIMIZATION SUPPORT VECTOR MACHINE FOR ELECTRICAL INSTALLATION CERTIFICATION PREDICTION

Priyono, Priyono, Panca Saputra, Elin, Suswandi, Suswandi, Rahman, Taufik
Abstract: Abstract: Feature selection is a crucial process that is very important to improve the performance of machine learning models, in accordance with data preprocessing. The feature selection process can be considered as a global… lobal combinatorial optimization problem in machine learning, which reduces the number of features, eliminates irrelevant data, and produces acceptable classification accuracy. The purpose of this study is to predict or determine the results of the electrical installation operation feasibility test based on data and obtain attribute selection features, and obtain accuracy level results. The Particle Swarm Optimization (PSO) approach is used to select the right characteristics to determine the results of the electrical installation operation feasibility test because attribute selection is needed in data analysis, because the PSO method will increase accuracy than just SVM in determining attribute selection. If SVM is used with PSO, the accuracy value is 96% and AUC is 0.994%, while the SVM method produces an accuracy level of 94.89% and AUC of 0.994%. With this finding, the accuracy value increases by 2%, making it a very good categorization category. It has been proven that the use of Particle Swarm Optimization (PSO) based algorithms can improve and improve results.             Keywords: PSO; SVM; Certification   Abstrak: Pemilihan fitur merupakan proses krusial yang sangat penting untuk meningkatkan kinerja model machine learning, sesuai dengan praproses data. Proses pemilihan fitur dapat dianggap sebagai masalah optimasi kombinatorial global dalam pembelajaran mesin, yang mengurangi jumlah fitur, menghilangkan data yang tidak relevan, dan menghasilkan akurasi klasifikasi yang dapat diterima. Tujuan dari penelitian ini adalah untuk memprediksi atau menentukan hasil uji kelayakan operasi instalasi listrik berdasarkan data dan memperoleh fitur pemilihan atribut, serta memperoleh hasil tingkat akurasi. Pendekatan Particle Swarm Optimization (PSO) digunakan untuk memilih karakteristik yang tepat untuk menentukan hasil uji kelayakan operasi instalasi listrik karena pemilihan atribut diperlukan dalam analisis data, karena metode PSO akan meningkatkan akurasi dari pada hanya SVM dalam menentukan pemilihan atribut. Jika SVM digunakan dengan PSO, nilai akurasinya adalah 96% dan AUC sebesar 0,994%, sedangkan metode SVM menghasilkan tingkat akurasi sebesar 94,89% dan AUC sebesar 0,994%. Dengan temuan ini, nilai akurasi meningkat sebesar 2%, menjadikannya kategori kategorisasi yang sangat baik. Telah terbukti bahwa penggunaan algoritma berbasis Particle Swarm Optimization (PSO) dapat meningkatkan dan memperbaiki hasil.   Kata kunci: PSO; SVM; Sertifikasi  

IMPLEMENTATION OF K-NEAREST NEIGHBOR ALGORITHM FOR CLASSIFICATION OF LUNG CANCER CAUSES

Almeyda, Hanindiya Putri, Khoiri, Zidan Fathannul, Haris, M Sabirin, Alkaff, Nabilah Husen, Sukmadiningtyas, Sukmadiningtyas
Abstract: Abstract: Lung cancer is most deadly cancers in the world. Identification and classification of the causes of understanding lung cancer is essential for developing more effective prevention and treatment strategies. The… issue is that a lot of individuals are unaware about the characteristics and causes of lung cancer. The purpose of this study is to apply the K-Nearest Neighbor (K-NN) algorithm in the classification of the causes of lung cancer and provide education to the public must be aware of the traits of lung cancer patients and, to stay away from the causes of lung cancer. The dataset used consists of 309 samples with 16 relevant attributes. The K-NN algorithm was trained and tested to assess its ability to classify the factors that cause lung cancer. The results showed an accuracy of 90.32%, with a precision for the "YES" class of 96% and the "NO" class of 67%. The recall value for the "YES" class was 92% and for the "NO" class was 80%. The implementation of this algorithm gives good results in classification and can help in early detection and prevention of lung cancer which can be used in the development of more effective prevention and early diagnosis strategies. Keywords: lung cancer; k-nearest neighbor; classification; machine learning     Abstrak: Kanker paru-paru tergolong jenis penyakit kanker yang memperoleh angka kematian paling tinggi di dunia. Identifikasi dan klasifikasi penyebab kanker paru-paru sangat penting untuk pengembangan strategi pencegahan dan pengobatan yang lebih efektif. Masalah yang terjadi adalah banyak orang yang belum mengetahui tentang ciri-ciri dan penyebab-penyebab dari kangker paru tersebut. Tujuan penelitian ini adalah mengimplementasikan algoritma K-Nearest Neighbor (K-NN) dalam klasifikasi penyebab kanker paru-paru serta memberikan edukasi kepada masyarakat banyak agar mengetahui ciri-ciri orang yang mengidap kangker paru-paru dan tentunya untuk menghindari penyebab-penyebab dari kangker paru-paru tersebut. Dataset yang digunakan terdiri dari 309 sampel dengan 16 atribut yang relevan. Algoritma K-NN kemudian dilatih dan diuji untuk menilai kemampuannya dalam mengklasifikasikan faktor-faktor penyebab kanker paru-paru. Hasil penelitian menunjukkan akurasi sebesar 90.32%, dengan skor precision untuk kelas "YES" sebesar 96% dan kelas "NO" sebesar 67%. Nilai recall untuk kelas "YES" adalah 92% dan untuk kelas "NO" sebesar 80%. Implementasi algoritma ini memberikan hasil yang baik dalam klasifikasi dan dapat membantu dalam deteksi dini serta pencegahan kanker paru-paru yang dapat digunakan dalam pengembangan strategi pencegahan dan diagnosis dini yang lebih efektif.   Kata kunci: kanker paru-paru; k-nearest neighbor; klasifikasi; machine learning

ANALYSIS OF PUBLIC OPINION SENTIMENT REGARDING POLICE INSTITUTIONS BASED ON TWITTER USING THE SUPPORT VECTOR MACHINE (SVM) METHOD

Sirojudin, Said Ahmad, Susanti, Try, Aribangsa, Mhd Theo
Abstract: Abstract: Twitter occupies the top position of the most popular social media platform in Indonesia. Police and other related issues were the subject of much discussion. The aim of this research is to analyze public sentiment… ment towards the National Police Agency using Twitter with the support vector machine method. The research started by crawling Twitter data. The data contains a total of 6,925 entries for three keywords. Next, we move on to the preprocessing stage consisting of (cleaning, case folding, tokenization, and filtering). Next is the tf-idf feature extraction stage, finally the classification and evaluation stage. The results of manual data inspection (73:27) showed accuracy of 70.66%, precision of 70.68%, and recall of 99.76%. Testing the second data (82:18), found accuracy 86%, precision 86.21%, recall 99.71%. The results of manual data checking (82:18) showed accuracy of 70.66%, precision of 70.68%, recall of 99.76%. Testing the second data (82:18), found accuracy 86%, precision 86.21%, recall 99.71%. From the data system testing results (80:20), accuracy was 87.55%, positive precision 87.53%, negative precision 88.24%, positive recall 99.48%, and negative recall. the rate is 99.48.% – The result is 21.43%. Data testing results (60:40) showed accuracy of 86.89%, positive precision of 86.84%, negative precision of 88.46%, positive recall of 99.61%, and negative recall of 16.43%. Single test data validation system (80:20), accuracy 87.55, overall test cross validation system (k fold 5 accuracy) 86.673%. Keywords: data mining;police agencies;support vector machines   Abstrak: Twitter menduduki posisi teratas platform media sosial terpopuler di Indonesia. Polisi dan masalah terkait lainnya menjadi pokok bahasan banyak pembicaraan. Tujuan penelitian ini untuk menganalisis sentimen masyarakat terhadap Badan Kepolisian Nasional menggunakan Twitter dengan  metode support vector machine. Penelitian dimulai dengan  crawling  data Twitter. Data memuat total 6.925 entri dari tiga kata kunci. Selanjutnya beralih ke tahap preprocessing terdiri dari (pembersihan, pelipatan kasus, tokenisasi, dan pemfilteran). Selanjutnya tahap ekstraksi fitur tf-idf, terakhir tahap klasifikasi dan evaluasi. Hasil pemeriksaan data manual (73:27) menunjukkan akurasi 70,66%, presisi 70,68%, dan recall 99,76%. Menguji data kedua (82:18), menemukan akurasi 86%, presisi 86,21%, recall 99,71%. Hasil pemeriksaan data secara manual (82:18) menunjukkan akurasi 70,66%, presisi 70,68%, recall 99,76%. Menguji data kedua (82:18), menemukan akurasi 86%, presisi 86,21%, recall 99,71%. Dari hasil pengujian sistem data (80:20), akurasi 87,55%, presisi positif 87,53%, presisi negatif 88,24%, recall positif 99,48%, dan recall negatif. tarifnya adalah 99,48.% – Hasilnya 21,43%. Hasil pengujian data (60:40) menunjukkan akurasi 86,89%, presisi positif 86,84%, presisi negatif 88,46%, recall positif 99,61%, dan recall negatif 16,43%. Uji tunggal sistem validasi data (80:20), akurasi 87,55, uji keseluruhan sistem validasi silang  (akurasi k fold 5) 86,673%.   Kata Kunci: data mining;instansi kepolisian;mesin vektor pendukung

SENTIMENT ANALYSIS OF PEGIPEGI.COM ON GOOGLE PLAYSTORE WITH NAÏVE BAYES ALGORITHM

Hardian, Riski, Oktaviana, Luzi Dwi, Hamdi, Aulia
Abstract: Abstract: Today, many users use online platforms rather than offline platforms for ticket bookings, involving a wide range of services such as flights, hotels, trains, buses, and entertainment. PegiPegi.com, as one of the… e fastest growing online travel agencies in Indonesia, demonstrates success by understanding the value of technology and maintaining strong partnerships. Users of this platform often provide reviews, viewing user reviews can be done manually but this will have a less effective impact, so it needs to be done automatically with sentiment analysis. This research the Naïve Bayes method in sentiment analysis of PegiPegi.com reviews, with a focus on understanding customer satisfaction and service improvement. By combining these approaches, this research contributes to a deeper understanding of user responses to OTA services and presents the evaluation results of the Multinomial Naive Bayes classification model with an accuracy rate of 89.5%. The high precision in the Negative class demonstrates the model's ability to identify negative reviews. However, there are challenges in classifying the Neutral class, indicating the potential for further improvement. Nevertheless, the F1 score of 0.522 reflects a good balance between overall precision, recall so it can be concluded the naïve bayes algorithm is successful for performing sentiment analysis. Keywords: Sentiment analysis; naïve bayes algorithm; pegipegi.com; playstore     Abstract: Saat ini banyak pengguna platform online dibandingkan offline untuk pemesanan tiket, yang melibatkan berbagai layanan seperti penerbangan, hotel, kereta api, bus, dan hiburan. PegiPegi.com, sebagai salah satu agen perjalanan online yang berkembang pesat di Indonesia, menunjukkan keberhasilan dengan memahami nilai teknologi dan mempertahankan kemitraan yang kuat. Pengguna platform ini sering memberikan ulasan, melihat ulasan pengguna bisa saja dilakukan secara manual tetapi hal ini akan memberikan dampak yang kurang efektif, sehingga perlu dilakukan secara otomatis dengan analisis sentiment. Penelitian ini bertujuan untuk menerapkan metode klasifikasi Naïve Bayes dalam analisis sentimen ulasan PegiPegi.com, dengan fokus pada pemahaman kepuasan pelanggan dan peningkatan layanan. Dengan menggabungkan pendekatan ini, penelitian ini berkontribusi pada pemahaman yang lebih dalam tentang tanggapan pengguna terhadap layanan OTA dan menyajikan hasil evaluasi model klasifikasi Multinomial Naive Bayes dengan tingkat akurasi 89,5%. Presisi tinggi di kelas Negatif menunjukkan kemampuan model untuk mengidentifikasi ulasan negatif. Namun, ada tantangan dalam mengklasifikasikan kelas Netral, menunjukkan potensi untuk perbaikan lebih lanjut. Namun demikian, skor F1 0,522 mencerminkan keseimbangan yang baik antara presisi keseluruhan dan daya ingat sehingga dapat disimpulkan algoritma naïve bayes berhasil untuk melakukan analisis sentimen. Keywords: Analisis sentimen; naïve bayes; pegipegi.com; playstore

SUPPORT VECTOR MACHINE ANALYSIS FOR INTEREST AND TALENT CLASSIFICATION WITH PYTHON LIBRARY

Sartika, Devi, Elfaladonna, Febie, Putra, Andre Mariza
Abstract: Abstract: Recognizing one's interests and talents early on is crucial in guiding an individual toward a prosperous future. While distinct, interests and talents share a close relationship. Interest denotes a genuine attraction… action to something without external pressure, and when consistently nurtured, it evolves into a skill or talent. Machine learning, specifically utilizing the SVM algorithm with the RBF kernel, can be applied to categorize interests and talents. Prior to SVM modeling, conducting Exploratory Data Analysis (EDA) is imperative for scrutinizing interests and talents. This analysis facilitates the identification of variables, enabling the elimination of missing values and ensuring the selection of appropriate interest and talent variables. The primary objective is to achieve optimal accuracy in modeling the classification of interests and talents. The insights gained from this research contribute to the creation of an application designed for categorizing interests and talents within SDN XYZ school. This application is designed for student use, assisting them in making informed decisions about their future education and career paths             Keywords: exploratory data analysis; interests and talents; machine learning; SVM Algorithm     Abstrak: Mengenali minat dan bakat seseorang sejak dini sangat penting dalam membimbing individu menuju masa depan yang sukses. Meskipun berbeda, minat dan bakat memiliki hubungan yang erat. Minat mengindikasikan ketertarikan yang tulus terhadap sesuatu tanpa tekanan eksternal, dan ketika terus-menerus dibina, berkembang menjadi keterampilan atau bakat. Pembelajaran mesin, khususnya dengan menggunakan algoritma SVM dan kernel RBF, dapat digunakan untuk mengelompokkan minat dan bakat. Sebelum pemodelan SVM, melakukan Analisis Data Eksploratif (EDA) sangat penting untuk mengkaji minat dan bakat. Analisis ini memfasilitasi identifikasi variabel, memungkinkan penghilangan nilai yang hilang, dan memastikan pemilihan variabel minat dan bakat yang tepat. Tujuan utamanya adalah mencapai akurasi optimal dalam pemodelan klasifikasi minat dan bakat. Temuan dari penelitian ini berkontribusi pada pengembangan aplikasi yang ditujukan untuk mengkategorikan minat dan bakat di sekolah SDN XYZ. Aplikasi ini dirancang untuk digunakan oleh siswa, membantu mereka membuat keputusan yang terinformasi mengenai pendidikan dan karier masa depan mereka.   Kata kunci: Algoritma SVM; exploratory data analysis; machine learning; minat dan bakat 

DETECTION OF LEAF SPOT DISEASE IN OIL PALM SEEDLINGS USING CONVOLUTIONAL NEURAL NETWORK METHOD

Azhar, Yufis, Zulva, Muhammad Shalahuddin
Abstract: Abstract: This research aims to develop a method for detecting leaf spot disease in oil palm seedlings using Convolutional Neural Network (CNN). Leaf spot disease in oil palm seedlings can hinder growth and production. CNN… NN has proven effective in image processing and classification, particularly in plant disease detection. In this study, we utilized a dataset of images containing oil palm seedling leaves infected with leaf spot disease and healthy leaves. We performed data processing, built a CNN model, and conducted hyperparameter tuning. The test results demonstrate that the developed CNN model achieves high accuracy in recognizing and distinguishing between oil palm seedling leaves infected with leaf spot disease and healthy ones. This research contributes to the development of plant disease detection technology that can support economic growth in the oil palm plantation sector.   Keywords: Convolutional Neural Network, image processing, leaf spot disease detection, oil palm seedlings.   Abstrak: Penelitian ini bertujuan untuk mengembangkan metode deteksi penyakit bercak pada bibit kelapa sawit menggunakan Convolutional Neural Network (CNN). Bibit kelapa sawit yang terinfeksi penyakit bercak dapat menghambat pertumbuhan dan produksi kelapa sawit. Metode CNN telah terbukti efektif dalam pengolahan citra dan klasifikasi, khususnya dalam deteksi penyakit pada tanaman. Dalam penelitian ini, kami menggunakan dataset citra daun bibit kelapa sawit yang terinfeksi penyakit bercak dan yang normal. Kami melakukan processing data, membangun model CNN, dan melakukan tuning hyperparameter. Hasil pengujian menunjukkan bahwa model CNN yang dikembangkan memiliki akurasi yang tinggi dalam mengenali dan membedakan citra daun bibit kelapa sawit yang terinfeksi penyakit bercak dan yang normal. Penelitian ini memberikan kontribusi dalam pengembangan teknologi deteksi penyakit tanaman yang dapat mendukung pertumbuhan ekonomi di sektor perkebunan kelapa sawit.   Kata kunci: bibit kelapa sawit, Convolutional Neural Network, deteksi penyakit bercak,  pengolahan citra.

CLASSIFICATION OF FAKE NEWS IN INDONESIAN LANGUAGE USING SUPPORT VECTOR MACHINE METHOD

Tandiano, Andreas Halim, Jollyta, Deny
Abstract: Abstract: Since information and communication technology has become ingrained in our daily lives, it has become easier to access information. However, there are some concerns. One of them is about fake news. The aim of this… his study is to develop an Indonesian system for detecting false news by utilizing news headlines. The methods used are linear kernel support vector ma- chine and n-gram. According to the findings of the performance test that was carried out, the linear kernel support vector machine model employing the term frequency inverse document frequency unigram feature performs better than utilizing bigram. The precision value generated from the model performance test is 1.00. This means that the degree of accuracy in matching the requested information regarding fake news detection with the answers provided by the system is very good. Then the recall value generated is 0.99. This means the linear kernel support vector machine model using unigram news features is very effective for detecting fake news according to the text classification approach.             Keywords: classification; fake news; n-gram; support vector machine   Abstrak: Dengan adanya integrasi teknologi informasi dan komunikasi dalam kehidupan mem- buat kemudahan dalam mengakses informasi. Walaupun demikian, terdapat kekhawatiran akan beberapa hal. Salah satu di antaranya adalah berita palsu. Tujuan penelitian ini adalah merancang sistem deteksi berita palsu berbahasa Indonesia berdasarkan judul berita. Metode yang digunakan adalah Support Vector Machine kernel linier dan n-gram. Berdasarkan hasil uji performa, model Support Vector Machine kernel linier yang menggunakan fitur term frequency inverse document frequency unigram menunjukkan kinerja yang lebih baik dibandingkan bi- gram. Nilai precision yang dihasilkan dari uji performa model sebesar 1,00. Ini berarti derajat akurasi dalam mencocokkan informasi yang diminta mengenai deteksi berita palsu dengan ja- waban yang diberikan oleh sistem sangat baik. Kemudian nilai recall yang dihasilkan sebesar 0,99. Ini berarti model Support Vector Machine kernel linier dengan menggunakan fitur berita unigram sangat efektif untuk mendeteksi berita palsu menurut pendekatan teks klasifikasi.   Kata kunci: klasifikasi; berita palsu; n-gram; support vector machine  

IMPLEMENTATION OF DATA ANALYSIS HOTEL RATING LEVELS IN BALI USING THE K-MEANS ALGORITHM AND DECISION TREE

Hamdani, Hamdani, Hartama, Dedy
Abstract: Abstract: The service dramatically affects the number of guests staying at the hotel. Bali is the most visited tourist area by foreign tourists. Therefore, improved service is crucial for determining the rating level of… a hotel. This research aims to combine two data mining algorithms: clustering and classification. This research is expected to contribute to hospitality in improving the best services for tourists, especially in the City of Bali.  Clustering algorithms are used to group the best number of hotels based on the four clusters selected from the k-means clustering algorithm. The classification algorithm using C4.5 determines the factors most dominant in determining the hotel rating level based on the gain ratio. The data used in this study results from observations on the website agoda.com in Bali of 51 data. The results of this study explained that cluster_0 is the highest-rated cluster, with a total number of 19 hotels found in claster_0. Data cluster0 is used for classification analysis using a decision tree, and the most dominant factor is the service factor, with an accuracy of 80%.             Keywords: data mining; kmeans; decision tree; hotel; bali;     Abstrak: Pelayanan sangat mempengaruhi jumlah pengunjung yang menginap dihotel. Bali merupakan daerah wisata paling banyak dikunjungi oleh wisatawan mancanegara. Oleh karena itu, peningkatan pelayanan sangat penting untuk penentuan level rating dari hotel. Tujuan dari penelitian ini untuk menggabungkan dua algoritma data mining yaitu clustering dan klasifikasi. Dengan penelitian ini diharapkan dapat memberikan kontribusi bagi perhotelan dalam meningkatkan pelayanan yang terbaik bagi wisatawan khususnya di Kota Bali.  Algoritma Clustering digunakan untuk mengelompokkan dari jumlah hotel yang terbaik berdasarkan empat cluster yang dipilih dari algoritma clustering berupa k-means. Algoritma klasifikasi menggunakan C4.5 digunakan untuk mengetahui faktor apa yang paling dominan dalam menentukan level rating hotel berdasarkan gain ratio. Data yang digunakan dalam penelitian ini hasil observasi di website agoda.com di bali sebanyak 51 data. Hasil dari penelitian ini menjelaskan dataset cluster_0 merupakan cluster rating tertinggi dengan jumlah 19 hotel yang terdapat di cluster_0. Data cluster_0 digunakan untuk analisis klasifikasi menggunakan decesion tree, didapat faktor yang paling dominan adalah faktor layanan dengan nilai akurasi sebesar 80%.   Kata kunci: data mining; kmeans; decision tree; hotel; bali;  

SENTIMENT ANALYSIS OF PUBLIC OPINIONS TOWARDS TELKOM UNIVERSITY POST PANDEMIC

Djakaria, Anindya Prameswari Putri, Pratiwi, Oktariani Nurul, Fakhrurroja, Hanif
Abstract: Abstract: Twitter, as a social media platform, has rapidly grown as a means for people to express their opinions and thoughts on various topics, including education. The number of Twitter users surged to 10.645.000 in 2020,… 20, with a significant increase during the pandemic. Telkom University, as a private institution of higher education in Indonesia, has become one of the topics of discussion on Twitter. Users’ opinions about Telkom University vary, ranging from positive to negative. To gain deeper insights into public view, sentiment analysis is essential. The analysis follows the Knowledge Discovery in Databases (KDD) process, utilizing the Naive Bayes classification algorithm. The evaluation results indicate the best accuracy achieved with an 80:20 data split, resulting in an accuracy rate of 82.05%, precision of 82.3%, recall of 82.05%, and F1-Score of 82.08%. The Naïve Bayes model demonstrates good performance for sentiment analysis of public views regarding Telkom University on Twitter.             Keywords: naïve bayes; sentiment analysis; twitter; telkom university.     Abstrak: Media sosial Twitter berkembang pesat sebagai sarana masyarakat berekspresi untuk menuangkan opini dan pikiran mereka mengenai topik apapun, termasuk pendidikan. Pengguna Twitter meningkat tajam hingga 10.645.00 pengguna pada tahun 2020 dan terus meningkat selama pandemi. Telkom University sebagai perguruan tinggi menjadi salah satu topik yang dibicarakan yang berkaitan dengan pendidikan. Pendapat mengenai Telkom University yang diungkapkan oleh pengguna Twitter beragam, baik positif maupun negatif. Analisis sentimen diperlukan untuk memahami pandangan publik lebih mendalam. Digunakan tahapan Knowledge Discovery in Databases dan algoritma klasifikasi Naïve Bayes dalam analisis ini. Hasil evaluasi menunjukkan akurasi paling baik dicapai dengan rasio data 80:20, dengan nilai akurasi sebesar 82.05%, nilai presisi sebesar 82.3%, nilai recall sebesar 82.05%, dan nilai F1-Score sebesar 82.08%. Model klasifikasi Naïve Bayes memiliki performa baik untuk analisis sentimen pandangan publik di Twitter mengenai Telkom University.   Kata kunci: analisis sentimen; naïve bayes; twitter; telkom university.