Abstract:Abstract: The rapid growth of the cosmetics industry on e-commerce platforms has intensified competition, creating a critical need for effective, data-driven marketing strategies. This study aims to conduct a comparative…
analysis of machine learning algorithms to predict the sales categories (High, Medium, Low) of cosmetic products on the Tokopedia marketplace. Four classification models; Random Forest, XGBoost, Logistic Regression, and Naive Bayes were trained and evaluated on data collected via web scraping. The methodology incorporates the Synthetic Minority Over-sampling Technique (SMOTE) to address significant class imbalance and GridSearchCV for hyperparameter optimization to ensure a fair and robust comparison. The experimental results conclusively show that the Random Forest model achieved the best performance, yielding the highest F1-Score Macro Average of 0.75 and an accuracy of 85.3%. The superior model was subsequently implemented in a simple recommendation system to simulate optimal discount strategies, demonstrating its practical utility in providing actionable insights for business decisions.
Keywords: classification; comparative analysis; machine learning; sales prediction; SMOTE
Abstrak: Pertumbuhan pesat industri kosmetik pada platform e-commerce telah membuat persaingan ketat, sehingga menciptakan kebutuhan krusial akan strategi pemasaran yang efektif dan berbasis data. Penelitian ini bertujuan untuk melakukan analisis komparatif terhadap algoritma machine learning untuk memprediksi kategori penjualan (Tinggi, Sedang, Rendah) produk kosmetik di marketplace Tokopedia. Empat model klasifikasi, yaitu Random Forest, XGBoost, Regresi Logistik, dan Naive Bayes, dilatih dan dievaluasi menggunakan data yang dikumpulkan melalui web scraping. Metodologi penelitian ini menerapkan Synthetic Minority Over-sampling Technique (SMOTE) untuk mengatasi ketidakseimbangan kelas yang signifikan dan GridSearchCV untuk optimisasi hyperparameter guna memastikan perbandingan yang adil. Hasil eksperimen menunjukkan bahwa model Random Forest mencapai performa terbaik, dengan menghasilkan F1-Score Macro Average tertinggi sebesar 0,75 dan akurasi 85,3%. Model unggul ini kemudian diimplementasikan dalam sebuah sistem rekomendasi sederhana untuk menyimulasikan strategi diskon yang optimal, yang menunjukkan kegunaan praktisnya dalam memberikan wawasan yang dapat ditindaklanjuti untuk pengambilan keputusan bisnis.
Kata kunci: analisis komparatif; klasifikasi; machine learning; prediksi penjualan; SMOTE
Abstract:Abstract: Infectious diseases are one of the most common health problems in children because they have immature immune systems. Children are more susceptible to infections caused by bacteria, viruses, fungi, and protozoa.…
. Some common infectious diseases in children include fever, acute respiratory infections (ARI), pneumonia, acute gastroenteritis (GAE), measles, chickenpox, and diphtheria. The limited number of pediatricians and the difficulty of accessing health facilities in remote areas hinder children's health services. To overcome this, an Android-based expert system is needed using the Naïve Bayes method to help diagnose infectious diseases in children earlier. The research method used is the Software Development Life Cycle (SDLC), where Black Box is used for internal testing, and PSSUQ is used to measure user satisfaction. The data set used was 1320 taken from a local hospital. The test results show that all the main features work as expected without any errors. The implementation of the system in diagnosing diseases went well and based on end-user feedback from 74 respondents, the system obtained a user satisfaction score of 6.40, where users felt that the system was easy to use, efficient, and provided clear and useful information.
Keywords: expert system; infectious disease; naïve bayes; PSSUQ; SDLC
Abstrak: Penyakit menular merupakan salah satu masalah kesehatan yang paling umum terjadi pada anak-anak karena mereka memiliki sistem kekebalan tubuh yang belum matang. Anak-anak lebih rentan terhadap infeksi yang disebabkan oleh bakteri, virus, jamur, dan protozoa. Beberapa penyakit infeksi yang umum terjadi pada anak-anak antara lain demam, infeksi saluran pernapasan akut (ISPA), pneumonia, gastroenteritis akut (GEA), campak, cacar air, dan difteri. Keterbatasan jumlah dokter spesialis anak dan sulitnya akses ke fasilitas kesehatan di daerah terpencil, menjadi kendala pada pelayanan kesehatan anak. Untuk mengatasi hal tersebut, diperlukan sistem pakar berbasis Android menggunakan metode Naïve Bayes untuk membantu mendiagnosis penyakit infeksi pada anak-anak lebih dini. Metode penelitian yang digunakan adalah Software Development Life Cycle (SDLC), di mana Black Box untuk pengujian internal, dan PSSUQ untuk mengukur kepuasan pengguna. Data set yang digunakan adalah 1320 yang diambil dari rumah sakit setempat. Hasil pengujian menunjukkan bahwa seluruh fitur utama berjalan sesuai harapan tanpa kesalahan. Implementasi sistem dalam mendiagnosa penyakit berjalan dengan baik dan berdasarkan umpan balik pengguna akhir dari 74 responden, sistem memperoleh skor kepuasan pengguna sebesar 6,40, di mana pengguna merasa sistem ini mudah digunakan, efisien, serta menyediakan informasi yang jelas dan bermanfaat.
Kata kunci: naïve bayes; penyakit menular; PSSUQ; SDLC; sistem pakar
Abstract:Analisis sentimen adalah proses mengidentifikasi dan mengklasifikasikan opini dalam teks menjadi kategori tertentu seperti positif, negatif, atau netral. Penelitian ini bertujuan untuk menganalisis sentimen ulasan produk…
pada platform e-commerce menggunakan algoritma Naive Bayes. Dataset ulasan produk diambil dari Kaggle, terdiri dari ribuan ulasan dengan label sentimen. Metodologi mencakup tahap preprocessing teks, ekstraksi fitur menggunakan teknik TF-IDF, dan penerapan algoritma Naive Bayes untuk klasifikasi sentimen. Hasil penelitian menunjukkan bahwa algoritma Naive Bayes memberikan akurasi sebesar 94%, membuktikan kemampuannya dalam analisis sentimen dengan dataset teks pendek
Abstract:Abstract: Today, many users use online platforms rather than offline platforms for ticket bookings, involving a wide range of services such as flights, hotels, trains, buses, and entertainment. PegiPegi.com, as one of the…
e fastest growing online travel agencies in Indonesia, demonstrates success by understanding the value of technology and maintaining strong partnerships. Users of this platform often provide reviews, viewing user reviews can be done manually but this will have a less effective impact, so it needs to be done automatically with sentiment analysis. This research the Naïve Bayes method in sentiment analysis of PegiPegi.com reviews, with a focus on understanding customer satisfaction and service improvement. By combining these approaches, this research contributes to a deeper understanding of user responses to OTA services and presents the evaluation results of the Multinomial Naive Bayes classification model with an accuracy rate of 89.5%. The high precision in the Negative class demonstrates the model's ability to identify negative reviews. However, there are challenges in classifying the Neutral class, indicating the potential for further improvement. Nevertheless, the F1 score of 0.522 reflects a good balance between overall precision, recall so it can be concluded the naïve bayes algorithm is successful for performing sentiment analysis.
Keywords: Sentiment analysis; naïve bayes algorithm; pegipegi.com; playstore
Abstract: Saat ini banyak pengguna platform online dibandingkan offline untuk pemesanan tiket, yang melibatkan berbagai layanan seperti penerbangan, hotel, kereta api, bus, dan hiburan. PegiPegi.com, sebagai salah satu agen perjalanan online yang berkembang pesat di Indonesia, menunjukkan keberhasilan dengan memahami nilai teknologi dan mempertahankan kemitraan yang kuat. Pengguna platform ini sering memberikan ulasan, melihat ulasan pengguna bisa saja dilakukan secara manual tetapi hal ini akan memberikan dampak yang kurang efektif, sehingga perlu dilakukan secara otomatis dengan analisis sentiment. Penelitian ini bertujuan untuk menerapkan metode klasifikasi Naïve Bayes dalam analisis sentimen ulasan PegiPegi.com, dengan fokus pada pemahaman kepuasan pelanggan dan peningkatan layanan. Dengan menggabungkan pendekatan ini, penelitian ini berkontribusi pada pemahaman yang lebih dalam tentang tanggapan pengguna terhadap layanan OTA dan menyajikan hasil evaluasi model klasifikasi Multinomial Naive Bayes dengan tingkat akurasi 89,5%. Presisi tinggi di kelas Negatif menunjukkan kemampuan model untuk mengidentifikasi ulasan negatif. Namun, ada tantangan dalam mengklasifikasikan kelas Netral, menunjukkan potensi untuk perbaikan lebih lanjut. Namun demikian, skor F1 0,522 mencerminkan keseimbangan yang baik antara presisi keseluruhan dan daya ingat sehingga dapat disimpulkan algoritma naïve bayes berhasil untuk melakukan analisis sentimen.
Keywords: Analisis sentimen; naïve bayes; pegipegi.com; playstore
Abstract:Abstract: Twitter, as a social media platform, has rapidly grown as a means for people to express their opinions and thoughts on various topics, including education. The number of Twitter users surged to 10.645.000 in 2020,…
20, with a significant increase during the pandemic. Telkom University, as a private institution of higher education in Indonesia, has become one of the topics of discussion on Twitter. Users’ opinions about Telkom University vary, ranging from positive to negative. To gain deeper insights into public view, sentiment analysis is essential. The analysis follows the Knowledge Discovery in Databases (KDD) process, utilizing the Naive Bayes classification algorithm. The evaluation results indicate the best accuracy achieved with an 80:20 data split, resulting in an accuracy rate of 82.05%, precision of 82.3%, recall of 82.05%, and F1-Score of 82.08%. The Naïve Bayes model demonstrates good performance for sentiment analysis of public views regarding Telkom University on Twitter.
Keywords: naïve bayes; sentiment analysis; twitter; telkom university.
Abstrak: Media sosial Twitter berkembang pesat sebagai sarana masyarakat berekspresi untuk menuangkan opini dan pikiran mereka mengenai topik apapun, termasuk pendidikan. Pengguna Twitter meningkat tajam hingga 10.645.00 pengguna pada tahun 2020 dan terus meningkat selama pandemi. Telkom University sebagai perguruan tinggi menjadi salah satu topik yang dibicarakan yang berkaitan dengan pendidikan. Pendapat mengenai Telkom University yang diungkapkan oleh pengguna Twitter beragam, baik positif maupun negatif. Analisis sentimen diperlukan untuk memahami pandangan publik lebih mendalam. Digunakan tahapan Knowledge Discovery in Databases dan algoritma klasifikasi Naïve Bayes dalam analisis ini. Hasil evaluasi menunjukkan akurasi paling baik dicapai dengan rasio data 80:20, dengan nilai akurasi sebesar 82.05%, nilai presisi sebesar 82.3%, nilai recall sebesar 82.05%, dan nilai F1-Score sebesar 82.08%. Model klasifikasi Naïve Bayes memiliki performa baik untuk analisis sentimen pandangan publik di Twitter mengenai Telkom University.
Kata kunci: analisis sentimen; naïve bayes; twitter; telkom university.
Abstract:Abstract: The JKN Mobile application is a mobile application created to facilitate healthcare administration in Indonesia since 2017. The application has been downloaded by over 10 million users and has received 484,000…
diverse reviews, including positive, negative, and neutral feedback. The average rating given by users is 4.5 out of 5 stars. This research aims to perform sentiment analysis on user reviews found in the Google Play Store review column. The methods used for sentiment analysis are Naive Bayes, K-Nearest Neighbor (K-NN), and Support Vector Machine (SVM). The test results show that with a 10% test data and 90% training data proportion, the SVM method achieves the highest accuracy of 95%. Naive Bayes follows with an accuracy of 87%, and K-NN with an accuracy of 75%.
Keywords: JKN mobile application, sentiment analysis, naive bayes, k-nearest neighbor (K-NN), support vector machine (SVM).
Abstrak: Aplikasi Mobile JKN adalah sebuah aplikasi yang dibuat untuk mempermudah administrasi kesehatan di Indonesia sejak tahun 2017. Aplikasi ini telah diunduh lebih dari 10 juta pengguna dengan 484 ribu ulasan beragam positif, negatif, dan netral. Rata-rata rating yang diberikan pengguna adalah 4,5 bintang dari 5 bintang. Penelitian ini bertujuan untuk melakukan analisis sentimen terhadap ulasan pengguna yang terdapat di kolom review Google Play Store. Metode yang digunakan untuk analisis sentimen adalah Naive Bayes, K-Nearest Neighbor (K-NN), dan Support Vector Machine (SVM). Hasil pengujian menunjukkan bahwa dengan menggunakan proporsi data uji sebesar 10% dan data training sebesar 90%, metode SVM mencapai akurasi tertinggi sebesar 95%. Diikuti oleh Naive Bayes dengan akurasi 87%, dan K-NN dengan akurasi 75%.
Kata kunci: JKN mobile, analisis sentimen, naïve bayes, k-nearest neighbor (K-NN), support vector machine (SVM).
Abstract:Abstract: Superior students represent active and intelligent students who are able to make a direct contribution to the development of the nation. To produce excellent students, one of the efforts that can be made is to…
foster these superior students from the start by forming a superior class. In forming a superior class, an effective selection process is needed to select students who are truly superior. This research was conducted at STMIK Royal Kisaran where the object of research was prospective new students for the superior class and the related party was the Academic and Student Administration Bureau (BAAK) because the party was directly involved in selecting prospective new students for the superior class. But the problem is, BAAK does not yet have an election process based on a method and there is no information system for selecting prospective new students for superior classes. The purpose of this study is to make predictions related to the selection of prospective new students for superior classes in the future with the Naïve Bayes Algorithm and implemented in an information system.The results of this research are that the Naïve Bayes algorithm produces an accuracy rate of 60% in the good classification category which is measured by the level of accuracy using the confusion matrix, so that the information system produced in this study also has efficient prediction results and is expected to help BAAK.
Keywords: data mining; naïve bayes; superior student class.
Abstrak: Mahasiswa unggulan merupakan representasi dari mahasiswa aktif dan cerdas yang mampu memberikan kontribusi langsung terhadap perkembangan bangsa. Untuk menghasilkan mahasiswa yang benar unggul salah satu upaya yang dapat dilakukan adalah dengan membina mahasiswa unggulan tersebut sejak awal dengan membentuk kelas unggulan. Dalam membentuk kelas unggulan dibutuhkan proses seleksi yang benar efektif untuk memilih mahasiswa yang benar unggul. Penelitian ini dilakukan di STMIK Royal Kisaran dimana objek penelitian merupakan calon mahasiswa baru untuk kelas unggulan dan pihak yang terkait adalah Biro Administrasi Akademik dan Kemahasiswaan (BAAK) karena pihak tersebut terlibat langsung dalam pemilihan calon mahasiswa baru untuk kelas unggulan. Namun permasalahannya, pihak BAAK belum memiliki sebuah proses pemilihan berdasarkan metode dan belum adanya sistem informasi pemilihan calon mahasiswa baru kelas unggulan. Tujuan dari penelitian ini adalah melakukan prediksi terkait pemilihan calon mahasiswa baru kelas unggulan di masa mendatang dengan Algoritma Naïve Bayes dan diimplementasikan dalam sebuah sistem informasi. Hasil penelitian ini algoritma Naïve Bayes menghasilkan tingkat akurasi sebesar 60% dengan kategori good classification yang diukur tingkat akurasinya dengan menggunakan confusion matrix, sehingga sistem informasi yang dihasilkan dalam penelitian ini juga memiliki hasil prediksi yang efisien dan diharapkan dapat membantu pihak BAAK.
Kata Kunci: data mining; mahasiswa kelas unggulan; naïve bayes.
Abstract:• Abstract: The skin is an elastic wrapping that protects the body from environmental influences, the skin is the organ that is located on the outside and limits it from the human environment. Skin diseases can be caused…
ed by fungi, viruses, germs, animal parasites, bacterial infections and others. To identify skin diseases, we usually have to see a doctor, but we still experience problems in dealing with disease identification. This is sometimes influenced by the community, sometimes they feel embarrassed to consult their skin disease to a doctor because the signs of skin disease have started to appear, consultation fees and drugs are relatively expensive. Current technological developments are able to process knowledge with artificial intelligence techniques. Due to the many symptoms of disease nowadays, it is necessary to make a system application with artificial intelligence that can diagnose skin diseases and provide solutions for skin diseases using one of the methods, namely the Naïve Bayes Classifier. Naïve Bayes is a simple classification algorithm where each attribute is independent and may contribute to the final decision. The goal is to produce an expert system website that helps the general public in diagnosing skin diseases and providing solutions for detected skin diseases. The results of this study concluded that based on the application of skin cancer diagnosis can display the results of skin cancer diagnosis decisions.
Keywords: expert system; naïve bayes; skin disease
Abstrak: Kulit merupakan pembungkus yang elastis yang melindungi tubuh dari pengaruh lingkungan, kulit merupakan organ tubuh yang terletak paling luar dan membatasinya dari lingkungan hidup manusia. Penyakit kulit dapat disebabkan oleh jamur, virus, kuman, parasit hewani, infeksi bakteri dan lain-lain. Mengidentifikasi penyakit kulit biasanya kita harus ke dokter, namun masih mengalami kendala dalam menangani pengidentifikasi penyakit hal itu terkadang dipengarui oleh masyarakat terkadang merasa malu untuk mengkonsultasikan penyakit kulitnya ke dokter karena tanda-tanda penyakit kulit sudah mulai tampak, biaya konsultasi dan obat yang tergolong mahal. Perkembangan teknologi saat ini mampu mengolah pengetahuan dengan teknik kecerdasan buatan. Karena banyaknya gejala penyakit pada masa sekarang ini perlu dibuat aplikasi sistem dengan kecerdasan buatan yang dapat mendiagnosa penyakit kulit dan memberikan solusi dari penyakit kulit dengan salah satu metode yaitu Naïve Bayes Classifier. Naïve bayes merupakan algoritma klasifikasi yang sederhana dimana setiap atribut bersifat berdiri sendiri dan memungkinkan berkontribusi terhadap keputusan akhir. Tujuannya adalah menghasilkan website sistem pakar yang dan membantu masyarakat luas dalam mendiagnosa penyakit kulit dan memberikan solusi dari penyakit kulit yang terdeteksi. Hasil dari penelitian ini menyimpulkan bahwa berdasarkan aplikasi diagnosa penyakit kanker kulit dapat menampilkan hasil keputusan diagnosa penyakit kanker kulit.
Kata kunci: Naïve Bayes; Penyakit Kulit; sistem pakar
Abstract:Abstract: Non-performing loan (NPL) is a risk that credit unions must face and to avoid that, prospective debtors need to be surveyed. With previous loan data, support vector machine and naïve bayes can be used as classification…
ssification methods to give a decision about NPL. We use a data set with 61 data and process the data with orange 3.30 application to see the difference between SVM using linear (SVM-L), polynomial (SVM-P), RBF (SVM-R) and sigmoid (SVM-S) kernel with naïve bayes. We use a cross validation technique with various folds to measure the classification results and a convusion matrix to measure the data training classification results. Naïve bayes scores the highest in terms of accuracy and SVM-R scores the highest in terms of F1, precision and recall. SVM-P scores the lowest in terms of accuracy, F1, precision and recall. Naïve bayes scores the highest in terms of proportion of predicted for true negative class and proportion of actual for true positive class. SVM-S scores the highest in terms of proportion of predicted for true positive class and proportion of actual for true negative class. SVM-P scores the lowest in both proportion of predicted and proportion of actual.
Keywords: classification; naïve bayes; non-performing loan; support vector machine
Abstrak: Kredit macet merupakan resiko yang sering dialami koperasi simpan pinjam, sehingga perlu dilakukan survei terhadap calon debitur agar kredit menjadi sehat. Dengan menggunakan data pemberian kredit sebelumnya, support vector machine dan naïve bayes digunakan sebagai metode klasifikasi untuk memberikan keputusan macet atau tidaknya kredit anggota koperasi Mutiara Sejahtera. Data set yang berjumlah 61 data diolah menggunakan aplikasi Orange 3.30 dan dilihat perbandingan antara metode SVM dengan kernel linear, polynomial, RBF dan sigomoid dengan metode naïve bayes. Cross validation dengan jumlah fold bervariasi digunakan sebagai nilai ukur klasifikasi dan convusion matrix digunakan sebagai nilai ukur klasifikasi data training. Hasil yang diperoleh adalah naïve bayes memiliki nilai accuracy tertinggi dan SVM kernel RBF memiliki nilai F1, precision dan recall tertinggi. SVM kernel polynomial memiliki nilai terendah untuk accuracy, F1, precision dan recall. Naïve bayes memiliki nilai tertinggi untuk proportion of predicted (PoP) kelas true negative dan proportion of actual (PoA) kelas true positive. SVM kernel sigmoid memiliki nilai tertinggi untuk PoP kelas true positive dan PoA kelas true negative. SVM kernel polynomial memiliki nilai terendah baik untuk PoP maupun PoA true negative dan kelas true positive.
Kata kunci: klasifikasi; kredit macet; naive bayes; SVM
Abstract:Abstract: Question classification is a computer science system, which aims to analyze questions and can label each question based on existing categories. Questions can be collected from several materials or topics that are…
re many and different. Therefore, the researcher intends to create a classification system for quiz questions Data Warehouse and Business Intelligence which can be grouped into topics Data Warehouse, Business Intelligence, Data Analytics, and Performance Measurement. One way to solve this problem is by approach machine learning. In this study, researchers used a comparison of machine learning algorithms, namely the algorithm NaïveBayes and SupportVectorMachine using SMOTE and methods Cross-Validation The results of this study show the best accuracy results and are very helpful. The results obtained in the method cross-validation before SMOTE resulted in an accuracy rate of 82.02% for the results after going through the SMOTE stage of 94.79% on the algorithm Naïve Bayes, while the algorithm SupportVectorMachine get accuracy of 81.39% in the process before SMOTE for the results after going through SMOTE of 96.52%.
Keywords: Cross-Validation; Machine Learning; Naive Bayes; Support Vector Machine; Question Classification
Abstrak: Klasifikasi pertanyaan merupakan sebuah sistem ilmu komputer, yang bertujuan untuk menganalisis pertanyaan serta dapat memberi label pada setiap pertanyaan berdasarkan kategori yang ada. Pertanyaan soal dapat dikumpulkan dari beberapa materi atau topik yang banyak dan berbeda. Oleh karena itu, bermaksud untuk membuat sistem klasifikasi pertanyaan soal kuis Data Warehouse dan Business Intelligence yang dapat dikelompokkan menjadi topik Data Warehouse, Business Intelligence, Data Analitik, dan Pengukuran Kinerja. Cara yang dapat dilakukan untuk permasalahan ini dengan menggunakan pendekatan MachineLearning. Pada penelitian kali ini menggunakan perbandingan algoritma MachineLearning yaitu algoritma NaïveBayes dan SupportVectorMachine menggunakan metode SMOTE dan Cross-Validation. Hasil penelitian ini menunjukkan hasil akurasi yang terbaik dan sangat membantu. Hasil yang diperoleh pada metode cross-validation sebelum SMOTE menghasilkan tingkat akurasi sebesar 82.02% untuk hasil sesudah melalui tahap SMOTE sebesar 94.79 % pada algoritma Naïve Bayes, sedangkan pada algoritma Support Vector Machine menghasilkan akurasi sebesar pada proses sebelum SMOTE 81.39% untuk hasil sesudah melalui SMOTE sebesar 96.52%.
Kata kunci: Klasifikasi Pertanyaan; Pembelajaran Mesin; Naive Bayes; Support Vector Machine; Cross-Validation