Search Articles & Publications

Showing 381 articles found for "Character"

COMPARISON OF CLUSTERING MODELS FOR GROUPING LIFESTYLE PATTERNS AND OBESITY FACTORS

Al Mas Ud, Khalid, Fathoni, Fathoni, Muhammad Kurniawan, Hafiz
Abstract: Abstract: Obesity is an escalating global health concern, with unhealthy lifestyle patterns contributing significantly to its development. This study aims to evaluate and compare three clustering techniques for categorizing… ing lifestyle patterns and obesity-related factors: K-Means, Agglomerative Clustering, and Gaussian Mixture Model (GMM). The data used in this study is sourced from the Food Nutrition dataset, which includes variables such as dietary habits, physical activity, and socio-economic status. The three clustering methods were assessed using evaluation metrics such as Silhouette Score, Davies-Bouldin Index (DBI), and Calinski-Harabasz Index (CHI). The findings revealed that K-Means exhibited the best performance in terms of cluster separation with a Silhouette Score of 0.5559, while GMM showed better flexibility in handling more complex data. Although Agglomerative Clustering produced acceptable results, it had a higher overlap between clusters compared to the other methods. This study offers valuable insights into selecting the most appropriate clustering technique based on the data characteristics.             Keywords: agglomerative; clustering; GMM; k-means; lifestyle patterns; obesity   Abstrak: Obesitas menjadi masalah kesehatan yang semakin meningkat di seluruh dunia, dengan pola hidup yang tidak sehat berperan besar dalam perkembangannya. Penelitian ini bertujuan untuk membandingkan tiga metode clustering dalam mengelompokkan pola gaya hidup dan faktor yang memengaruhi obesitas, yaitu K-Means, Agglomerative Clustering, dan Gaussian Mixture Model (GMM). Data yang digunakan diperoleh dari dataset Food Nutrition yang mencakup informasi terkait pola makan, aktivitas fisik, serta faktor sosial-ekonomi. Ketiga metode tersebut diuji dengan menggunakan beberapa metrik evaluasi, seperti Silhouette Score, Davies-Bouldin Index (DBI), dan Calinski-Harabasz Index (CHI). Hasil penelitian menunjukkan bahwa K-Means memiliki kinerja terbaik dalam hal pemisahan klaster, dengan nilai Silhouette Score sebesar 0.5559, sementara GMM lebih fleksibel dalam menangani data yang lebih kompleks. Meskipun Agglomerative Clustering memberikan hasil yang dapat diterima, tumpang tindih antar klaster lebih besar dibandingkan dengan kedua metode lainnya. Penelitian ini memberikan pemahaman yang lebih baik mengenai pemilihan metode clustering yang tepat berdasarkan karakteristik data yang digunakan.   Kata kunci: agglomerative; clustering; GMM; k-means; obesitas; pola gaya hidup

A COMPARATIVE ANALYSIS OF OPTIMIZED NEURAL NETWORK AND LARGE-SCALE LANGUAGE MODELS FOR MUSIC GENRE CLASSIFICATION

Marzuqi, Ahmad Naufal Luthfan, Nastiti , Vinna Rahmayanti Setyaning
Abstract: Abstract: The rapid growth of the digital music industry requires accurate music genre classification systems to enhance user experience in streaming services. This study compares a domain-specific Long Short-Term Memory… (LSTM) network with three Large Language Models (LLMs)—HuBERT, WavLM, and WAV2Vec 2.0—for Music Genre Classification (MGC). The LSTM model was trained using Mel-spectrograms transformed from the GTZAN dataset, while the LLMs were fine-tuned using a smaller set of raw audio samples due to computational constraints. All models were tested on datasets with identical genre labels to ensure a fair evaluation. Results show that the LSTM model achieved the highest accuracy of 97.10%, outperforming HuBERT (86.00%), WavLM (83.00%), and WAV2Vec 2.0 (80.00%). The LSTM demonstrated superior generalization and stability without overfitting, while the LLMs struggled to differentiate between genres with similar acoustic characteristics. These findings indicate that general-purpose pre-trained models, although powerful, are less effective in music-specific tasks due to domain mismatch. Therefore, incorporating music-specific features and architectures remains essential for achieving higher accuracy and reliability in automatic genre classification systems. Keywords: audio large language models; comparative deep learning; music genre classification.   Abstrak: Pertumbuhan industri musik digital yang pesat menuntut sistem klasifikasi genre musik yang akurat untuk meningkatkan pengalaman pengguna dalam layanan streaming. Penelitian ini dilatarbelakangi oleh perkembangan pesat model pembelajaran mendalam, khususnya jaringan LSTM dan model bahasa berskala besar LLM seperti HuBERT, WavLM, dan WAV2Vec 2.0, yang telah menunjukkan kemampuan representasi audio yang kuat. Tujuan penelitian ini ini membandingkan jaringan Long Short-Term Memory (LSTM) khusus domain dengan tiga model Large Language Models (LLM)—HuBERT, WavLM, dan WAV2Vec 2.0—untuk tugas Klasifikasi Genre Musik (MGC). Metode penelitian melibatkan pelatihan LSTM menggunakan data Mel-spectrogram hasil transformasi dari dataset GTZAN, sementara LLM disesuaikan (fine-tuning) menggunakan data audio mentah dalam jumlah lebih kecil karena keterbatasan komputasi. Seluruh model diuji pada dataset dengan label genre yang sama untuk memastikan evaluasi yang adil. Hasil penelitian menunjukkan bahwa model LSTM mencapai akurasi tertinggi sebesar 97,10%, sedangkan model HuBERT, WavLM, dan WAV2Vec 2.0 masing-masing memperoleh 86,00%, 83,00%, dan 80,00%. Model LSTM menunjukkan kemampuan generalisasi yang lebih baik tanpa overfitting, sedangkan model LLM cenderung kesulitan membedakan genre dengan karakteristik akustik yang mirip. Kesimpulan penelitian ini adalah ketidaksesuaian domain secara signifikan membatasi performa model umum saat diterapkan pada tugas berbasis musik. Oleh karena itu, penggunaan fitur dan arsitektur khusus musik sangat penting dalam membangun sistem klasifikasi genre yang lebih akurat. Kata kunci: klasifikasi genre musik; model bahasa besar; perbandingan pembelajaran mendalam.

DEVELOPMENT OF A BLOCKCHAIN-BASED DECENTRALISED APPLICATION WITH NFT FOR LAND REGISTRATION

Gesang, Rahmat Nugrohoning, Teduh Dirgahayu, Raden
Abstract: Abstract: Land registration in Indonesia often encounters challenges in transparency, data integrity, and centralized bureaucracy. Manual and semi-digital systems remain vulnerable to manipulation and delays. The National… l Land Agency has initiated digitalization, but several challenges remain, particularly in ensuring transparency, efficiency, and security of land ownership data. Blockchain technology offers a potential solution through its decentralized and immutable characteristics. This study adopted a design and development method consisting of system analysis, requirements identification, architecture design, implementation, and black-box testing. The developed decentralized application (DApp) integrates smart contracts, NFTs, and IPFS to manage land certificates. Core functions such as minting, transfer, splitting, and self-custody were implemented and successfully tested, with all scenarios producing expected results. The findings demonstrate that blockchain integration can enhance security, reduce duplication, and streamline land administration. The study contributes a functional prototype with practical implications for modernizing land registration in Indonesia while identifying scalability and regulatory adaptation as areas for further research.             Keywords: blockchain; decentralized application; land registration; NFT; smart contract.

COMPARISON OF K-MEANS AND K-MEDOIDS FOR DRUG DATA CLUSTERING

Andika, Tripa, Kurniabudi, Sharipuddin
Abstract: Abstract: Ineffective drug demand management can lead to problems such as imbalanced drug distribution, excess stock, or shortages in community health centers. To address this, data mining can be utilized to support the… planning and control process of drug inventory. Clustering techniques were chosen because they are able to group drug data based on certain characteristics, thus identifying stable and unstable drug supply patterns. This study aims to group drug data at Simpang Kawat Community Health Center in Jambi City, which can be used as a reference in planning drug needs in the next period. Data grouping is divided into three categories: slow-moving, medium-moving, and fast-moving. The research data includes attributes of drug name, initial stock, receipt, inventory, usage, and final stock, with a total of 1758 data sets, which were processed using the CRISP-DM framework through the RapidMiner application. Cluster quality evaluation was carried out using the Davies-Bouldin Index (DBI). The results showed that the K-Means algorithm obtained a DBI value of 0.175, smaller than K-Medoids which obtained a value of 0.354. Because a smaller DBI value indicates better cluster quality, K-Means provides more optimal clustering results than K-Medoids. Through these clustering results, community health centers can utilize drug cluster information to support more efficient drug procurement planning, as well as reduce the risk of excess or shortage of stock.             Keywords: data mining; clustering; k-means; k-medoids; davies-bouldin index

CRITERIA ANALYSIS OF COURSE PARTICIPANTS USING K-MEANS: A CASE STUDY OF INET PALEMBANG

Muhammad Rasuandi Akbar, Agramanisti Azdy, Rezania, Novaria Kunang, Yesi, Adha Oktarini Saputri , Nurul
Abstract: Abstract: INET Computer Palembang, as a computer training institution, faces difficulties in understanding participant characteristics due to variations in age, educational background, and chosen course packages. This study… udy aims to analyze participant criteria and group them based on similarities using the K-Means Clustering algorithm. The data used were historical records of course participants from 2022 to 2025. The research process followed the CRISP-DM stages, starting from data cleaning and transformation, determining the optimal number of clusters using the Elbow Method, to evaluating cluster quality with the Davies-Bouldin Index. The implementation was carried out using Python and the scikit-learn library. The results show that the optimal number of clusters is k=5 with a Sum of Squared Errors (SSE) value of 1064.66 and a Davies-Bouldin Index (DBI) score of 0.820, indicating good cluster quality. The resulting clustering provides a structured profile of participants and demonstrates that K-Means is effective in segmenting course participants. These findings are expected to assist the institution in designing more targeted training programs. Keywords: clustering; data mining; elbow method; k-means; computer course

APPLICATION OF PARTICLE SWARM OPTIMIZATION SUPPORT VECTOR MACHINE FOR ELECTRICAL INSTALLATION CERTIFICATION PREDICTION

Priyono, Priyono, Panca Saputra, Elin, Suswandi, Suswandi, Rahman, Taufik
Abstract: Abstract: Feature selection is a crucial process that is very important to improve the performance of machine learning models, in accordance with data preprocessing. The feature selection process can be considered as a global… lobal combinatorial optimization problem in machine learning, which reduces the number of features, eliminates irrelevant data, and produces acceptable classification accuracy. The purpose of this study is to predict or determine the results of the electrical installation operation feasibility test based on data and obtain attribute selection features, and obtain accuracy level results. The Particle Swarm Optimization (PSO) approach is used to select the right characteristics to determine the results of the electrical installation operation feasibility test because attribute selection is needed in data analysis, because the PSO method will increase accuracy than just SVM in determining attribute selection. If SVM is used with PSO, the accuracy value is 96% and AUC is 0.994%, while the SVM method produces an accuracy level of 94.89% and AUC of 0.994%. With this finding, the accuracy value increases by 2%, making it a very good categorization category. It has been proven that the use of Particle Swarm Optimization (PSO) based algorithms can improve and improve results.             Keywords: PSO; SVM; Certification   Abstrak: Pemilihan fitur merupakan proses krusial yang sangat penting untuk meningkatkan kinerja model machine learning, sesuai dengan praproses data. Proses pemilihan fitur dapat dianggap sebagai masalah optimasi kombinatorial global dalam pembelajaran mesin, yang mengurangi jumlah fitur, menghilangkan data yang tidak relevan, dan menghasilkan akurasi klasifikasi yang dapat diterima. Tujuan dari penelitian ini adalah untuk memprediksi atau menentukan hasil uji kelayakan operasi instalasi listrik berdasarkan data dan memperoleh fitur pemilihan atribut, serta memperoleh hasil tingkat akurasi. Pendekatan Particle Swarm Optimization (PSO) digunakan untuk memilih karakteristik yang tepat untuk menentukan hasil uji kelayakan operasi instalasi listrik karena pemilihan atribut diperlukan dalam analisis data, karena metode PSO akan meningkatkan akurasi dari pada hanya SVM dalam menentukan pemilihan atribut. Jika SVM digunakan dengan PSO, nilai akurasinya adalah 96% dan AUC sebesar 0,994%, sedangkan metode SVM menghasilkan tingkat akurasi sebesar 94,89% dan AUC sebesar 0,994%. Dengan temuan ini, nilai akurasi meningkat sebesar 2%, menjadikannya kategori kategorisasi yang sangat baik. Telah terbukti bahwa penggunaan algoritma berbasis Particle Swarm Optimization (PSO) dapat meningkatkan dan memperbaiki hasil.   Kata kunci: PSO; SVM; Sertifikasi  

STUDENT CLUSTER ANALYSIS AS AN EFFORT TO OPTIMIZE CAMPUS PROMOTION

Aulia, Romy, Khomarudin, Agus Nur, Laksmana, Indra, Jamaluddin, Jamaluddin, Novita, Rina
Abstract: Abstract: This research tries to describe student cluster analysis, as an effort to optimize campus promotion to various schools and regions. It is known that every year, Politeknik Pertanian Negeri Payakumbuh, abbreviated… ed as PPNP, brings in students from various regions in Indonesia. Regarding the campus promotion strategy process, the PPNP promotion section has not been based or referred to the results of processing existing student data. So that the budget used by the campus promotion team has not been right on target with the results of students who can be brought to campus. In addition, the existing student database has not been processed or explored further, so that it has not produced knowledge that is very useful as material to support the decisions of the academic and student affairs department and the campus promotion team. The method used in this research is CRISP-DM which stands for Cross- Industry Standard Process for Data Mining. Based on the characteristics of each cluster, the PPNP Promotion Team in conducting the next socialization is advised to prioritize provinces such as West Sumatra and North Sumatra. Currently, managerial circles in this context, university leaders are expected to be able to make data-based decisions. Data-based decision making can foster a culture of sustainable innovation, produce customer-centric offerings and drive long-term business growth.             Keywords: cluster analysis; student data; k-means clustering; campus promotion     Abstrak: Penelitian ini mencoba untuk mendeskripsikan analisis cluster mahasiswa, sebagai upaya optimalisasi dalam melakukan promosi kampus ke berbagai sekolah dan daerah. Diketahui bahwa setiap tahunnya, Politeknik Pertanian Negeri Payakumbuh disingkat PPNP mendatangkan mahasiswa dari berbagai daerah di Indonesia. Terkait dengan proses strategi promosi kampus, bagian promosi PPNP belum didasarkan pada hasil pengolahan data mahasiswa yang ada. Sehingga anggaran yang digunakan tim promosi belum tepat sasaran dengan hasil mahasiswa yang dapat didatangkan ke kampus. Selain itu database mahasiswa yang ada selama ini belum diolah atau digali secara jauh, sehingga belum menghasilkan pengetahuan yang bermanfaat sebagai bahan untuk mendukung keputusan bagian akademik dan kemahasiswaan serta tim promosi kampus. Metode yang digunakan dalam penelitian ini yaitu CRISP-DM merupakan singkatan dari Cross-Industry Standart Process for Data Mining. Berdasarkan karakteristik setiap cluster, maka untuk Tim Promosi PPNP dalam melakukan sosialisasi berikutnya disarankan memprioritaskan pada provinsi seperti Sumatera Barat dan Sumatera Utara. Saat ini kalangan manajerial yaitu pimpinan perguruan tinggi diharapkan dapat melakukan pengambilan keputusan berbasis pada data. Pengambilan keputusan berbasis data dapat menumbuhkan  budaya  inovasi  yang berkelanjutan, menghasilkan penawaran yang berpusat pada pelanggan dan mendorong pertumbuhan bisnis jangka panjang.   Kata kunci: analisis cluster; data mahasiswa; k-means clustering, promosi kampus

IMPLEMENTATION OF K-NEAREST NEIGHBOR ALGORITHM FOR CLASSIFICATION OF LUNG CANCER CAUSES

Almeyda, Hanindiya Putri, Khoiri, Zidan Fathannul, Haris, M Sabirin, Alkaff, Nabilah Husen, Sukmadiningtyas, Sukmadiningtyas
Abstract: Abstract: Lung cancer is most deadly cancers in the world. Identification and classification of the causes of understanding lung cancer is essential for developing more effective prevention and treatment strategies. The… issue is that a lot of individuals are unaware about the characteristics and causes of lung cancer. The purpose of this study is to apply the K-Nearest Neighbor (K-NN) algorithm in the classification of the causes of lung cancer and provide education to the public must be aware of the traits of lung cancer patients and, to stay away from the causes of lung cancer. The dataset used consists of 309 samples with 16 relevant attributes. The K-NN algorithm was trained and tested to assess its ability to classify the factors that cause lung cancer. The results showed an accuracy of 90.32%, with a precision for the "YES" class of 96% and the "NO" class of 67%. The recall value for the "YES" class was 92% and for the "NO" class was 80%. The implementation of this algorithm gives good results in classification and can help in early detection and prevention of lung cancer which can be used in the development of more effective prevention and early diagnosis strategies. Keywords: lung cancer; k-nearest neighbor; classification; machine learning     Abstrak: Kanker paru-paru tergolong jenis penyakit kanker yang memperoleh angka kematian paling tinggi di dunia. Identifikasi dan klasifikasi penyebab kanker paru-paru sangat penting untuk pengembangan strategi pencegahan dan pengobatan yang lebih efektif. Masalah yang terjadi adalah banyak orang yang belum mengetahui tentang ciri-ciri dan penyebab-penyebab dari kangker paru tersebut. Tujuan penelitian ini adalah mengimplementasikan algoritma K-Nearest Neighbor (K-NN) dalam klasifikasi penyebab kanker paru-paru serta memberikan edukasi kepada masyarakat banyak agar mengetahui ciri-ciri orang yang mengidap kangker paru-paru dan tentunya untuk menghindari penyebab-penyebab dari kangker paru-paru tersebut. Dataset yang digunakan terdiri dari 309 sampel dengan 16 atribut yang relevan. Algoritma K-NN kemudian dilatih dan diuji untuk menilai kemampuannya dalam mengklasifikasikan faktor-faktor penyebab kanker paru-paru. Hasil penelitian menunjukkan akurasi sebesar 90.32%, dengan skor precision untuk kelas "YES" sebesar 96% dan kelas "NO" sebesar 67%. Nilai recall untuk kelas "YES" adalah 92% dan untuk kelas "NO" sebesar 80%. Implementasi algoritma ini memberikan hasil yang baik dalam klasifikasi dan dapat membantu dalam deteksi dini serta pencegahan kanker paru-paru yang dapat digunakan dalam pengembangan strategi pencegahan dan diagnosis dini yang lebih efektif.   Kata kunci: kanker paru-paru; k-nearest neighbor; klasifikasi; machine learning

CLUSTERING ROTATIONAL CHURN OF TELECOMMUNICATIONS CUSTOMERS USING A DATA-CENTRICAI APPROACH

Muttaqin, Widang, Lydia, Maya Silvi, Fahmi, Fahmi
Abstract: Abstract: In the current era of very fast technological development, customer churn is a serious challenge, especially in the competitive telecommunications industry. Churn refers to customers who stop using a service or… move to another provider, and can be categorized into three types: Active Churn, Passive Churn, and Rotational Churn. Rotational Churn, which is difficult to predict be- cause the reasons for stopping are unclear, is the main focus of this research. This research aims to group Rotational Churn customers using a Data-Centric AI approach. This approach emphasizes improving data quality through Confident Learning and Synthetic Data before being applied to the K-Means clustering algorithm. The data used in this research is customer churn data from one telecommunications company during 2023. The research results show that customer grouping using the K-Means algorithm can provide deep insight into the characteristics of customer churn. The application of Data-Centric AI is proven to be able to increase the accuracy of clustering models, which ultimately helps compa- nies optimize programs and services to minimize churn and retain customers.       Keywords: data-centric AI; clustering; K-means     Abstrak: Dalam era perkembangan teknologi yang sangat pesat saat ini, churn pelanggan menjadi tantangan serius, terutama dalam industri telekomunikasi yang sangat kompetitif. Churn mengacu pada pelanggan yang berhenti menggunakan layanan atau beralih ke penyedia lain, dan dapat dikategorikan menjadi tiga jenis: Churn Aktif, Churn Pasif, dan Churn Rotasional. Churn Rotasional, yang sulit diprediksi karena alasan penghentian layanan tidak jelas, menjadi fokus utama penelitian ini. Penelitian ini bertujuan untuk mengelompokkan pelanggan Churn Rotasional menggunakan pendekatan Data-Centric AI. Pendekatan ini menekankan pada peningkatan kualitas data melalui Confident Learning dan Synthetic Data sebelum diterapkan ke algoritma K-Means clustering. Data yang digunakan dalam penelitian ini adalah data churn pelanggan dari satu perusahaan telekomunikasi selama tahun 2023. Hasil penelitian menunjukkan bahwa pengelompokan pelanggan menggunakan algoritma K-Means dapat memberikan wawasan mendalam tentang karakteristik churn pelanggan. Penerapan Data-Centric AI terbukti mampu meningkatkan akurasi model klastering, yang pada akhirnya membantu perusahaan mengoptimalkan program dan layanan untuk meminimalkan churn serta mempertahankan pelanggan.   Kata kunci: data-Centric AI; klasterisasi; K-means  

COMPARATIVE ANALYSIS OF SOBEL AND CANNY METHOD IN BATIK KAWUNG IMAGE

Surmayanti, Surmayanti, Sumijan, Sumijan
Abstract: Abstract: Abstract: this study evaluates and compares the performance of two edge detection methods Sobel method and Canny method on batik image.  Batik images have unique characteristics and complex patterns, making it… difficult to analyze the edges.  This study presents a comparison of the results using sobel and canny edge detection methods on batik kawung images both from peak signal-to-noise reatio and from mean squared error. The results showed that canny edge detection was better than sobel method. This can be seen from the results of PSNR and MSE that is 100%. This analysis is determined by considering factors such as the accuracy of edge detection, sensitivity to noise, and the ability to handle the complexity of batik drawing patterns. The results of this study provide a detailed description of the advantages and disadvantages of each method in the image of batik kawung. The conclusions that can be drawn from this study can provide valuable guidance for choosing the optimal edge detection method in image analysis of batik kawung and others.      Keywords: batik kawung; canny; MSE; PSNR; sobel     Abstrak: Penelitian ini mengevaluasi dan membandingkan kinerja dua metode deteksi tepi metode Sobel dan metode Canny pada citra batik.  Gambar batik mempunyai ciri-ciri yang unik dan pola yang kompleks, sehingga menyulitkan analisis tepian.  Penelitian ini menyajikan perbandingan hasil menggunakan metode deteksi tepi sobel dan canny pada citra batik kawung baik dari peak signal-to-noise reatio maupun dari mean squared error. Hasil penelitian ini menunjukkan bahwa deteksi tepi canny lebih baik dibandingkan dari metode sobel. Hal ini dapat dilihat dari hasil PSNR dan MSE yang dihasilkan yaitu 100%. Analisis ini ditentukan dengan mempertimbangkan faktor-faktor seperti keakuratan deteksi tepi, kepekaan terhadap noise, dan kemampuan menangani kompleksitas pola gambar batik. Hasil penelitian ini memberikan gambaran secara detail mengenai kelebihan dan kekurangan masing-masing metode pada citra batik kawung. Kesimpulan yang dapat diambil dari penelitian ini dapat memberikan panduan berharga untuk memilih metode deteksi tepi yang optimal dalam analisis citra batik kawung dan yang lainnya.   Kata kunci: batik kawung; canny; MSE; PSNR; sobel