Abstract:Abstract: Choosing the right elementary school is a crucial milestone for a child's future. However, the large number of school options in Asahan Regency with diverse criteria such as accreditation, facilities, fees, curriculum,…
riculum, and accessibility often makes it difficult for parents. Decision-making tends to be based on subjective word-of-mouth recommendations, which risks triggering bias. This research aims to develop an adaptive and objective web-based elementary school selection Decision Support System (DSS) framework to minimize such bias. The system is designed using a hybrid model that combines the Simple Multi-Attribute Rating Technique (SMART) method as a dynamic criteria weighting engine based on parents' preferences, and the Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) method to rank ten alternative schools. System testing was conducted through black-box testing functionality and empirical accuracy testing using Spearman Rank Correlation involving 40 respondents. The black-box testing results confirmed that all main system operations ran perfectly with a 100% success rate. Validity testing demonstrated very high accuracy, with a Spearman correlation coefficient of 0.89, demonstrating significant alignment between the system's recommendations and actual choices on the ground. Thus, the SMART-TOPSIS framework has proven reliable as a data-driven approach to helping parents choose the best elementary school.
Keywords: decision support system; elementary school; SMART; TOPSIS
Abstrak: Memilih sekolah dasar yang tepat merupakan tonggak awal krusial bagi masa depan anak. Namun, banyaknya pilihan sekolah di Kabupaten Asahan dengan keberagaman kriteria seperti akreditasi, fasilitas, biaya, kurikulum, dan aksesibilitas sering kali menyulitkan orang tua. Pengambilan keputusan pun cenderung didasarkan pada rekomendasi subjektif dari mulut ke mulut, yang berisiko memicu bias. Penelitian ini bertujuan mengembangkan kerangka kerja Sistem Pendukung Keputusan (SPK) pemilihan sekolah dasar berbasis web yang adaptif dan objektif guna meminimalkan bias tersebut. Sistem dirancang menggunakan model hibrida yang mengombinasikan metode *Simple Multi-Attribute Rating Technique* (SMART) sebagai mesin pembobotan kriteria dinamis sesuai preferensi orang tua, serta metode *Technique for Order of Preference by Similarity to Ideal Solution* (TOPSIS) untuk memeringkat sepuluh sekolah alternatif. Pengujian sistem dilakukan melalui uji fungsionalitas *black-box testing* dan uji akurasi empiris menggunakan Korelasi Peringkat Spearman dengan melibatkan 40 responden. Hasil *black-box testing* mengonfirmasi seluruh operasi utama sistem berjalan sempurna dengan tingkat keberhasilan 100%. Uji validitas menunjukkan akurasi sangat tinggi dengan koefisien korelasi Spearman sebesar 0,89, membuktikan keselarasan signifikan antara rekomendasi sistem dan pilihan nyata di lapangan. Dengan demikian, kerangka SMART-TOPSIS ini terbukti andal sebagai pendekatan berbasis data untuk membantu orang tua memilih sekolah dasar terbaik.
Kata kunci: sekolah dasar; sistem pendukung keputusan; SMART; TOPSIS
Abstract:Abstract: Existing IoT anomaly detection studies have achieved high classification performance, but most focus on accuracy and F1-score without explicitly controlling the false positive rate (FPR). In addition, many approaches…
oaches rely on a single detection perspective, limiting their operational reliability. To address this gap, this study proposes a hybrid anomaly detection framework integrating Long Short-Term Memory (LSTM), Shannon entropy, and autoencoder reconstruction error. Shannon entropy is incorporated as an additional feature, while LSTM and the autoencoder capture temporal and reconstruction characteristics. The resulting hybrid representation is processed by a constraint-based threshold selection mechanism that enforces FPR . Experiments on the TON-IoT and Edge-IIoTset datasets achieved average F1-scores of 0.9250 and 0.9934, while maintaining average FPR values of 0.0091 and 0.0714, respectively. Analysis of entropy distributions showed consistent differences between normal and anomalous traffic across both datasets, indicating that Shannon entropy provides discriminative information for anomaly detection. These results demonstrate strong detection performance with controlled false alarms, while ablation studies confirm the significant contribution of Shannon entropy to overall model performance.
Keywords: false positive rate; hybrid deep learning; Internet of Things; network anomaly detection; Shannon entropy
Abstrak: Penelitian deteksi anomali Internet of Things (IoT) telah menunjukkan performa klasifikasi yang tinggi, namun sebagian besar masih berfokus pada accuracy dan F1-score tanpa mengendalikan false positive rate (FPR) secara eksplisit. Selain itu, banyak pendekatan hanya memanfaatkan satu perspektif deteksi sehingga reliabilitas operasionalnya masih terbatas. Untuk mengatasi kesenjangan tersebut, penelitian ini mengusulkan kerangka deteksi anomali hybrid yang mengintegrasikan Long Short-Term Memory (LSTM), Shannon entropy, dan autoencoder reconstruction error. Shannon entropy digunakan sebagai fitur tambahan, sedangkan LSTM dan autoencoder menangkap karakteristik temporal dan deviasi rekonstruksi. Representasi hybrid yang dihasilkan kemudian diproses melalui mekanisme constraint-based threshold selection dengan batas FPR . Hasil pengujian pada dataset TON-IoT dan Edge-IIoTset menghasilkan F1-score rata-rata sebesar 0,9250 dan 0,9934, dengan FPR rata-rata sebesar 0,0091 dan 0,0714. Perbedaan nilai entropy yang konsisten antara trafik normal dan anomali pada kedua dataset menunjukkan bahwa Shannon entropy menyediakan informasi diskriminatif untuk deteksi anomali. Hasil tersebut menunjukkan performa deteksi yang kuat dengan false alarm yang terkendali, sementara studi ablasi mengonfirmasi kontribusi signifikan Shannon entropy terhadap performa model.
Kata kunci: deteksi anomali jaringan; false positive rate; hybrid deep learning; Internet of Things; Shannon entropy
Abstract:Abstract: The advancement of digital technology has made it easier to create, process, and distribute files—using 317 files from the dataset https://www.kaggle.com/datasets/axon data/selfie-and-official-id-photo-dataset-18k…
t-18k images?select=metadata_image.csv has also introduced new challenges, such as the increasing practice of digital file manipulation that is difficult to detect visually. Therefore, an intelligent digital forensics system that can automatically and accurately detect file authenticity is required. This study aims to develop an intelligent digital forensics system for detecting file manipulation by leveraging metadata analysis and the Random Forest classification method. The methods used include extracting metadata from digital files—such as time information, device details, and processing history—followed by analysis to identify patterns of inconsistency that indicate manipulation. This data is then used as features in the classification process using the Random Forest algorithm to distinguish between original and manipulated files. The results of this study are expected to show that the use of metadata analysis combined with the Random Forest algorithm can improve accuracy in detecting digital file manipulation compared to conventional methods. The resulting system is expected to provide an effective, efficient, and integrated solution to support digital forensic investigations, Based on the test results, the system demonstrated good performance with an accuracy rate of 94%.
Keywords: Digital Forensics;File Manipulation;Metadata Analysis;Random Forest;Classification;Machine Learning
Abstrak:Perkembangan teknologi digital telah meningkatkan kemudahan dalam pembuatan, pengolahan,dan distribusi file sebanyak 317 file, sumber datasets https:// www.kaggle.com/datasets/axondata/selfie-and-official-id-photo-dataset-18k-images?select =metadata_image.csv, namun juga menimbulkan tantangan baru berupa meningkatnya praktik manipulasi file digital yang sulit dideteksi secara kasat mata. Oleh karena itu, diperlukan suatu sistem forensik digital yang cerdas dan mampu mendeteksi keaslian file secara otomatis dan akurat. Penelitian ini bertujuan untuk mengembangkan sistem forensik digital cerdas untuk deteksi manipulasi file dengan memanfaatkan analisis metadata dan metode klasifikasi Random Forest. Metode yang digunakan meliputi proses ekstraksi metadata dari file digital, seperti informasi waktu, perangkat, dan riwayat pengolahan, kemudian dilakukan analisis untuk menemukan pola ketidaksesuaian yang mengindikasikan adanya manipulasi. Selanjutnya, data tersebut digunakan sebagai fitur dalam proses klasifikasi menggunakan algoritma Random Forest untuk membedakan antara file asli dan file yang telah dimanipulasi. Hasil dari penelitian ini diharapkan menunjukkan bahwa penggunaan analisis metadata yang dikombinasikan dengan algoritma Random Forest mampu meningkatkan akurasi dalam mendeteksi manipulasi file digital dibandingkan metode konvensional. Sistem yang dihasilkan dapat memberikan solusi yang efektif, efisien, dan terintegrasi dalam mendukung proses investigasi forensik digital, Berdasarkan hasil pengujian, sistem menunjukkan performa yang baik dengan tingkat akurasi sebesar 94%.
Kata Kunci: Forensik Digital, Manipulasi File, Metadata, Random Forest, Klasifikasi, Machine Learning.
Abstract:Abstract: The use of QR Codes in academic settings has increased with the digitization of attendance systems, but it has also introduced potential abuse in the form of quishing attacks (QR phishing). Previous studies have…
e mainly focused on user behavior, while forensic analysis of digital artifacts as evidence is still limited. This study aims to conduct a forensic analysis of browser artifacts resulting from interactions with dangerous QR Codes at Aisyiyah University Yogyakarta using the framework of the National Justice Institute (NIJ). Six investigation parameters are defined: domain identification, endpoint identification, identification of supporting resources, visualization of image artifacts, timestamp correlation, and HTML reconstruction. Data is obtained from the Google Chrome profile directory and analyzed using Autopsy, focusing on Web Cache, Browser History, and Cookies artifacts. The results showed that five parameters were successfully identified with an investigation success rate of 83.3%, while HTML reconstruction could not be fully achieved due to cache limitations. These findings show that Web Cache artifacts provide evidentiary value in the forensic investigation of QR Code-based attacks. Future research should focus on improving full-page reconstruction techniques.
Keywords: browser forensics; digital artifacts; NIJ; quishing; Web Cache
Abstrak: Penggunaan Kode QR di lingkungan akademik telah meningkat seiring dengan digitalisasi sistem absensi, tetapi juga menimbulkan potensi penyalahgunaan dalam bentuk serangan phishing (QR phishing). Studi sebelumnya sebagian besar berfokus pada perilaku pengguna, sementara analisis forensik artefak digital sebagai bukti masih terbatas. Studi ini bertujuan untuk melakukan analisis forensik artefak browser yang dihasilkan dari interaksi dengan Kode QR berbahaya di Universitas 'Aisyiyah Yogyakarta menggunakan kerangka kerja Lembaga Kehakiman Nasional (NIJ). Enam parameter investigasi didefinisikan: identifikasi domain, identifikasi titik akhir, identifikasi sumber daya pendukung, visualisasi artefak gambar, korelasi stempel waktu, dan rekonstruksi HTML. Data diperoleh dari direktori profil Google Chrome dan dianalisis menggunakan Autopsy, dengan fokus pada artefak Cache Web, Riwayat Browser, dan Cookie. Hasil menunjukkan bahwa lima parameter berhasil diidentifikasi dengan tingkat keberhasilan investigasi sebesar 83,3%, sementara rekonstruksi HTML tidak dapat sepenuhnya dicapai karena keterbatasan cache. Temuan ini menunjukkan bahwa artefak Cache Web memberikan nilai bukti dalam investigasi forensik serangan berbasis Kode QR. Penelitian selanjutnya harus fokus pada peningkatan teknik rekonstruksi halaman penuh.
Kata kunci: forensik peramban; artefak digital; NIJ; quishing; web cache
Abstract:Abstract: Anomalous sound detection is essential for industrial predictive maintenance, as machine failures often originate from subtle acoustic changes during operation. However, high background noise and limitations of…
conventional Convolutional Neural Networks (CNN) reduce detection reliability. This study proposes a 1D-CNN-based anomaly detection framework with multi-view feature fusion and temporal segmentation to enhance detection performance. The approach combines MFCC, Log-Mel Spectrogram, and Chroma STFT features, while temporal segmentation divides audio signals into 5-second segments to better capture transient anomalies. Experiments on the MIMII dataset under varying Signal-to-Noise Ratio (SNR) conditions show that MFCC and Log-Mel fusion achieves the best performance, with 97.90% accuracy and ROC-AUC of 0.9789. The model maintains accuracy above 90% at −6 dB, demonstrating strong robustness in noisy industrial environments.
Keywords: industrial anomaly detection; 1D-CNN; multi-view feature fusion; temporal segmentation; MIMII dataset.
Abstrak: Deteksi anomali suara merupakan komponen penting dalam sistem pemeliharaan prediktif industri, karena kegagalan mesin sering diawali oleh perubahan akustik yang bersifat halus selama proses operasi. Namun, tingkat kebisingan yang tinggi serta keterbatasan arsitektur Convolutional Neural Network (CNN) konvensional dapat menurunkan keandalan deteksi. Penelitian ini bertujuan mengusulkan kerangka deteksi anomali berbasis 1D-CNN yang mengintegrasikan strategi fusi fitur multi-view dan segmentasi temporal untuk meningkatkan kinerja deteksi. Pendekatan yang digunakan menggabungkan fitur MFCC, Log-Mel Spectrogram dan Chroma STFT, sementara teknik temporal splitting membagi sinyal audio menjadi segmen berdurasi 5 detik untuk menangkap anomali yang bersifat sementara. Eksperimen menggunakan dataset MIMII pada berbagai kondisi Signal-to-Noise Ratio (SNR) menunjukkan bahwa kombinasi MFCC dan Log-Mel Spectrogram menghasilkan kinerja terbaik dengan akurasi 97,90% dan ROC-AUC sebesar 0,9789. Model juga mempertahankan akurasi di atas 90% pada kondisi kebisingan ekstrem (−6 dB) yang menunjukkan ketahanan yang baik dalam lingkungan industri yang bising.
Kata kunci: deteksi anomali industri; 1D-CNN; fusi fitur multi-view; segmentasi temporal; dataset MIMII
Abstract:Abstract: The rapid growth of e-commerce mobile applications has generated large volumes of user reviews, making manual sentiment analysis increasingly impractical. This study aims to compare the effectiveness of three machine…
achine learning algorithms Support Vector Machine (SVM), Random Forest, and Naive Bayes for automated sentiment classification of Indonesian-language mobile application reviews. A dataset of 3,000 user reviews from the RupaRupa application on the Google Play Store was collected and preprocessed through normalization, tokenization, stopword removal, and stemming. TF-IDF vectorization was applied for feature extraction, while the Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance across three sentiment categories: positive, negative, and neutral. The results show that SVM achieved the highest accuracy of 90.02%, while Random Forest obtained the best F1-score of 88.08% when sufficient training data were available. Naive Bayes demonstrated relatively stable performance across varying training data sizes. Furthermore, TF-IDF keyword analysis revealed that negative reviews were primarily associated with delivery issues, technical problems, and pricing concerns. These findings demonstrate the effectiveness of machine learning approaches for sentiment classification and provide practical insights for improving mobile application services.
Keywords: sentiment analysis; machine learning; SMOTE; TF-IDF; text classification
Abstrak: Pertumbuhan pesat aplikasi mobile e-commerce telah menghasilkan volume ulasan pengguna yang sangat besar, sehingga analisis sentimen secara manual menjadi semakin tidak praktis. Penelitian ini bertujuan untuk membandingkan efektivitas tiga algoritma machine learning Support Vector Machine (SVM), Random Forest, dan Naive Bayes dalam melakukan klasifikasi sentimen otomatis terhadap ulasan aplikasi mobile berbahasa Indonesia. Dataset yang digunakan terdiri dari 3.000 ulasan pengguna aplikasi RupaRupa yang dikumpulkan dari Google Play Store. Data kemudian diproses melalui tahapan preprocessing yang meliputi normalisasi, tokenisasi, penghapusan stopword, dan stemming. Ekstraksi fitur dilakukan menggunakan metode Term Frequency–Inverse Document Frequency (TF-IDF), sedangkan ketidakseimbangan kelas ditangani menggunakan Synthetic Minority Over-sampling Technique (SMOTE) pada tiga kategori sentimen, yaitu positif, negatif, dan netral. Hasil penelitian menunjukkan bahwa SVM mencapai tingkat akurasi tertinggi sebesar 90,02%, sementara Random Forest memperoleh nilai F1-score terbaik sebesar 88,08% ketika tersedia data pelatihan yang memadai. Naive Bayes menunjukkan performa yang relatif stabil pada berbagai ukuran data pelatihan. Selain itu, analisis kata kunci berbasis TF-IDF mengungkapkan bahwa ulasan negatif terutama berkaitan dengan masalah pengiriman, kendala teknis aplikasi, dan isu harga. Temuan ini menunjukkan bahwa pendekatan machine learning efektif untuk klasifikasi sentimen serta memberikan wawasan yang bermanfaat dalam meningkatkan kualitas layanan aplikasi mobile.
Kata Kunci: analisis sentimen; pembelajaran mesin; SMOTE; TF-IDF; klasifikasi teks.
Abstract:Abstract: Inventory management of consumable medical devices and drugs plays a crucial role in maintaining the continuity of healthcare operations. However, Jelita Dental Care still faces challenges in recording and controlling…
rolling stock due to manual procedures, which can lead to data inaccuracy, procurement delays, and the risk of stockouts. To address these issues, this study aims to develop a web-based Electronic Supply Chain Management (E-SCM) system that integrates stock monitoring and procurement processes. The Reorder Point (ROP) method is applied to determine the optimal reorder point based on average demand, lead time, and safety stock. This system was built using the PHP programming language and MySQL database. The results show that the JelitaMed system is able to improve the effectiveness and accuracy of inventory management, simplify the structured procurement submission process between the admin, owner, and supplier, and support decision-making in maintaining the availability of consumable medical devices and drugs. Thus, the implementation of E-SCM combined with the ROP method is a practical solution to improve inventory control in small-scale health clinics.
Keywords: e-scm; inventory; information system; medical supplies; reorder point.
Abstrak: Pengelolaan persediaan alat dan obat medis habis pakai memiliki peran penting dalam menjaga keberlangsungan operasional layanan kesehatan. Namun, Jelita Dental Care masih menghadapi kendala dalam pencatatan dan pengendalian stok akibat prosedur manual, yang dapat menyebabkan ketidaktepatan data, keterlambatan pengadaan, serta risiko kekurangan persediaan. Untuk mengatasi permasalahan tersebut, penelitian ini bertujuan mengembangkan sistem Electronic Supply Chain Management (E-SCM) berbasis web yang mengintegrasikan pemantauan stok dan proses pengadaan. Metode Reorder Point (ROP) diterapkan untuk menentukan waktu pemesanan ulang yang optimal berdasarkan permintaan rata-rata, lead time, dan safety stock. Sistem ini dibangun menggunakan bahasa pemrograman PHP dan database MySQL. Hasil penelitian menunjukkan bahwa sistem JelitaMed mampu meningkatkan efektivitas dan akurasi pengelolaan persediaan, mempermudah proses pengajuan pengadaan secara terstruktur antara admin, owner, dan supplier, serta mendukung pengambilan keputusan dalam menjaga ketersediaan alat dan obat medis habis pakai. Dengan demikian, penerapan E-SCM yang dikombinasikan dengan metode ROP menjadi solusi praktis untuk meningkatkan pengendalian persediaan pada klinik kesehatan skala kecil.
Kata kunci: alat medis; e-scm; persediaan; reorder point; sistem informasi
Abstract:Abstract: Academic achievement mapping is an important process in higher education to support effective academic monitoring and guidance. In practice, student grouping is often conducted manually by academic staff using…
simple criteria such as Grade Point Average (GPA) thresholds and subjective judgment, without systematic data analysis. This study aims to apply the Fuzzy C-Means (FCM) clustering algorithm to objectively group students based on their academic achievement levels. The dataset consists of academic records from 179 sixth-semester students of the Computer Science Study Program at Universitas Islam Negeri Sumatera Utara, where 160 eligible students are processed in the FCM calculation. Three variables are used: cumulative GPA, total completed credits, and the total number of low grades (D/E). The FCM algorithm automatically performs the mapping and groups students into three categories, namely excellent, stable, and at-risk students. Cluster quality is evaluated using the Silhouette Score and Davies–Bouldin Index, showing satisfactory clustering performance. The results indicate that the proposed approach provides a data-driven and objective basis for academic decision support.
Keywords: academic achievement; clustering; fuzzy c-means; student
Abstrak: Pemetaan pencapaian akademik mahasiswa merupakan proses penting dalam pendidikan tinggi untuk mendukung pemantauan dan pembinaan akademik yang tepat sasaran. Dalam praktiknya, pengelompokan mahasiswa masih sering dilakukan secara manual oleh pihak akademik berdasarkan kriteria sederhana, seperti batasan Indeks Prestasi Kumulatif (IPK) dan penilaian subjektif, tanpa analisis data yang sistematis. Penelitian ini bertujuan menerapkan algoritma Fuzzy C-Means (FCM) untuk mengelompokkan mahasiswa secara objektif berdasarkan tingkat pencapaian akademik. Data penelitian berasal dari 179 mahasiswa semester enam Program Studi Ilmu Komputer Universitas Islam Negeri Sumatera Utara, dengan 160 mahasiswa memenuhi kriteria dan diproses menggunakan algoritma FCM. Variabel yang digunakan meliputi IPK kumulatif, jumlah SKS yang telah ditempuh, dan total nilai rendah (D/E). Proses pemetaan sepenuhnya dilakukan oleh algoritma FCM dan menghasilkan tiga kategori mahasiswa, yaitu unggul, stabil, dan berisiko. Evaluasi menggunakan Silhouette Score dan Davies–Bouldin Index menunjukkan kualitas pengelompokan yang cukup baik.
Kata kunci: fuzzy c-means; clustering; mahasiswa; pencapaian akademik
Abstract:Abstract: This research is motivated by the problem of building material inventory management at Jaqfar Building Store, which is still done manually and based on subjective estimates. This often results in inaccuracies in…
n determining stock levels, either in the form of overstock or understock, which hinders operational effectiveness. The purpose of this study is to apply the Multiple Linear Regression method to analyze the relationship between incoming stock (X1) and outgoing stock (X2) variables with the ending stock variable (Y) to produce an optimal inventory prediction model. The research methodology used includes collecting historical transaction data for building materials such as cement, ceramics, zinc, plywood, and iron. This web-based prediction system was developed using the PHP programming language and a MySQL database. The analysis results show that the resulting regression model can provide a mathematical picture of future inventory patterns based on historical data. Implementation of this system is expected to assist the management of Jaqfar Building Materials Store in making strategic decisions regarding purchasing and sales in a more measured and efficient manner.
Keyword: building materials; data mining; inventory; multiple linear regression
Abstrak: Penelitian ini dilatarbelakangi oleh permasalahan pengelolaan persediaan bahan bangunan di Toko Bangunan Jaqfar yang masih dilakukan secara manual dan berdasarkan perkiraan subjektif. Hal ini menyebabkan sering terjadinya ketidaktepatan dalam menentukan jumlah stok, baik berupa kelebihan barang (overstock) maupun kekurangan barang (understock) yang menghambat efektivitas operasional. Tujuan dari penelitian ini adalah menerapkan metode Multiple Linear Regression (Regresi Linear Berganda) untuk menganalisis hubungan antara variabel stok masuk (X1) dan stok keluar (X2) terhadap variabel stok akhir (Y) guna menghasilkan model prediksi persediaan yang optimal. Metodologi penelitian yang digunakan mencakup pengumpulan data historis transaksi bahan bangunan seperti semen, keramik, seng, triplek, dan besi. Sistem prediksi ini dikembangkan berbasis web menggunakan bahasa pemrograman PHP dan basis data MySQL. Hasil analisis menunjukkan bahwa model regresi yang dihasilkan mampu memberikan gambaran matematis mengenai pola persediaan di masa mendatang berdasarkan data historis. Implementasi sistem ini diharapkan dapat membantu manajemen Toko Bangunan Jaqfar dalam mengambil keputusan strategis terkait pembelian dan penjualan secara lebih terukur serta efisien.
Kata kunci: bahan bangunan; data mining; persediaan; regresi linear berganda
Abstract:Abstract: The growing intensity of cyber attacks, marked by rapid, large-scale, automated, and adaptive execution, requires analytical methods that represent the diversity of network environments, including variations in…
target platforms such as IoT, traditional networks, and hybrid infrastructures. This study compares machine learning models for cyber attack classification under heterogeneous environmental conditions and formulates a conceptual optimization framework based on model performance. Four publicly available benchmark datasets were used, namely UNB CIC IoT 2023, UNB CIC IDS-2018, UNSW-NB15, and a Kaggle cyber security attacks dataset, comprising approximately 40,000 to over 3.6 million records and 25 to 80 features across IoT, conventional, and mixed network environments. Random Forest, XGBoost, Multilayer Perceptron, and Transformer were implemented within a unified pipeline involving preprocessing, feature selection, and Bayesian Optimization-based hyperparameter tuning. All models achieved F1-score and Cohen's Kappa above 96%, with XGBoost performing best (97.80%, 97.26%), followed by Random Forest (97.78%, 96.96%) and Transformer (97.44%, 96.82%), while MLP scored lowest (96.74%, 96.00%), a gap below one percentage point. Confusion matrix analysis revealed persistent misclassification in minority and overlapping attack classes, informing a proposed adaptive cyber attack simulation optimization framework.
Keywords: cyber attacks; optimization; machine learning; environmental variability.
Abstrak: Meningkatnya intensitas serangan siber yang berlangsung cepat, masif, otomatis, dan adaptif menuntut pendekatan analitis yang merepresentasikan keragaman lingkungan jaringan, termasuk perbedaan karakteristik platform sasaran seperti Internet of Things (IoT), jaringan konvensional, dan infrastruktur hibrida. Penelitian ini membandingkan model machine learning untuk klasifikasi serangan siber pada kondisi lingkungan heterogen, sekaligus menyusun kerangka optimasi konseptual berdasarkan performa model. Empat dataset benchmark publik digunakan, yaitu UNB CIC IoT 2023, UNB CIC IDS-2018, UNSW-NB15, serta dataset Kaggle cyber security attacks, dengan jumlah data berkisar 40.000 hingga lebih dari 3,6 juta rekaman dan 25 sampai 80 fitur, mewakili lingkungan IoT, konvensional, dan campuran. Random Forest, XGBoost, Multilayer Perceptron, dan Transformer diimplementasikan melalui pipeline terpadu mencakup pra-pemrosesan, seleksi fitur, dan optimasi hyperparameter berbasis Bayesian Optimization. Seluruh model mencapai F1-score dan Cohen's Kappa di atas 96%, dengan XGBoost menunjukkan performa terbaik (97,80%, 97,26%), diikuti Random Forest (97,78%, 96,96%) dan Transformer (97,44%, 96,82%), sementara MLP mencatat skor terendah (96,74%, 96,00%), dengan selisih kurang dari satu poin persentase. Analisis confusion matrix mengungkap misklasifikasi yang konsisten pada kelas minoritas dan serangan dengan karakteristik serupa, yang menjadi dasar kerangka optimasi simulasi serangan siber adaptif yang diusulkan.
Kata kunci: serangan siber; optimasi; machine learning; variabilitas lingkungan