Class Imbalance-Aware SVM Optimization to Improve Daily Rainfall Event Prediction at Tanjung Perak

  • Ary Pulung Baskoro Universitas Hang Tuah
Keywords: rainfall events, Support Vector Machine, Class Imbalance, decision threshold, port operations

Abstract

Daily rainfall events have significant operational meaning in the port environment because rain can
disrupt loading and unloading activities, reduce occupational safety, hinder ship movements, and affect
scheduling efficiency. This study aims to improve the prediction of daily rainfall events at Tanjung
Perak Port through an optimization that explicitly considers class imbalance within the existing SVM-
RBF framework without adding new predictor variables. The model uses the original predictor set
comprising Madura Strait sea surface temperature anomalies, the Oceanic Niño Index, and cyclic
seasonal features, with daily data for the period 2015 to 2024. The 2015–2022 period is used as the
model development set, with 2015–2021 for main training and 2022 for internal validation in decision
threshold selection. The 2023–2024 period is used as independent test data. Five model configurations
were evaluated: the baseline SVM-RBF model, a variant with class weighting, a variant with decision
threshold tuning, a variant with class weighting and threshold tuning, and a variant with
oversampling and threshold tuning. The results show that the rainfall dataset has a significant class
imbalance, with the non-rain class dominating both the development and test sets. Among the tested
strategies, class weighting provides the most balanced improvement, raising recall from 0.531 to 0.676,
F1-score from 0.581 to 0.640, and balanced accuracy from 0.704 to 0.748, while maintaining the overall
accuracy almost unchanged. Threshold-based optimization indeed increases recall substantially but
also causes excessive false alarms. These findings indicate that class imbalance-aware optimization can
operationally improve rainfall event detection without expanding the predictor space, and that the
class-weighted SVM-RBF model is the most appropriate default solution for rain risk mitigation at
Tanjung Perak Port.

Downloads

Download data is not yet available.

References

[1] P. Branco, L. Torgo, and R. P. Ribeiro, "A survey of predictive modelling on imbalanced domains," ACM Computing Surveys, vol. 49, no. 2, Article 31, 2016.
[2] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic Minority Over-sampling Technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.
[3] C. Cortes and V. Vapnik, "Support-vector networks," Machine Learning, vol. 20, pp. 273–297, 1995.
[4] G. Gunawan, W. Andriani, and A. A. Akbar, "Application of machine learning for short-term climate prediction in Indonesia," Mantik Journal, vol. 8, no. 1, pp. 828–837, 2024.
[5] H. He and E. A. Garcia, "Learning from imbalanced data," IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263–1284, 2009.
[6] M. F. R. Mahendra, N. L. Azizah, and Sumarno, "Implementasi machine learning untuk memprediksi cuaca menggunakan Support Vector Machine," Jurnal Ilmiah Komputasi, vol. 23, no. 1, pp. 45–50, 2024.
[7] S. E. Purwati and Y. Pristyanto, "Model Random Forest and Support Vector Machine for flood classification in Indonesia," Sinkron: Jurnal dan Penelitian Teknik Informatika, vol. 8, no. 4, pp. 2261–2268, 2024.
[8] W. P. Putra et al., "Energy-efficient rainfall prediction using Support Vector Machine on edge AI platforms," JOIV: International Journal on Informatics Visualization, vol. 8, no. 3-2, pp. 1686–1692, 2024.
[9] T. Saito and M. Rehmsmeier, "The Precision-Recall Plot is more informative than the ROC Plot when evaluating binary classifiers on imbalanced datasets," PLOS ONE, vol. 10, no. 3, p. e0118432, 2015.
[10] O. A. Wani et al., "Predicting rainfall using machine learning, deep learning, and time series models across an altitudinal gradient in the North-Western Himalayas," Scientific Reports, vol. 14, p. 27876, 2024.
[11] V. N. Vapnik, Statistical Learning Theory. New York, NY, USA: Wiley, 1998.
[12] B. Schölkopf and A. J. Smola, Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. Cambridge, MA, USA: MIT Press, 2002.
[13] C. C. Chang and C. J. Lin, "LIBSVM: A library for support vector machines," ACM Transactions on Intelligent Systems and Technology, vol. 2, no. 3, pp. 1–27, 2011.
[14] G. Haixiang, Y. Li, S. Shang, M. Guanjun, H. Yuanyue, and B. Bing, "Learning from class-imbalanced data: Review of methods and applications," Expert Systems with Applications, vol. 73, pp. 220–239, 2017.
[15] R. Batuwita and V. Palade, "Class imbalance learning methods for support vector machines," in Imbalanced Learning: Foundations, Algorithms, and Applications, H. He and Y. Ma, Eds. Hoboken, NJ, USA: Wiley, 2013, pp. 83–99.
[16] Y. Tang, Y. Q. Zhang, N. V. Chawla, and S. Krasser, "SVMs modeling for highly imbalanced classification," IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 39, no. 1, pp. 281–288, 2009.
[17] A. Ali, S. M. Shamsuddin, and A. L. Ralescu, "Classification with class imbalance problem: A review," International Journal of Advances in Soft Computing and its Applications, vol. 7, no. 3, pp. 176–204, 2015.
[18] A. Fernández, S. Garcia, F. Herrera, and N. V. Chawla, "SMOTE for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary," Journal of Artificial Intelligence Research, vol. 61, pp. 863–905, 2018.
[19] D. Chicco and G. Jurman, "The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation," BMC Genomics, vol. 21, no. 1, p. 6, 2020.
[20] K. H. Brodersen, C. S. Ong, K. E. Stephan, and J. M. Buhmann, "The balanced accuracy and its posterior distribution," in Proc. 20th International Conference on Pattern Recognition, Istanbul, Turkey, 2010, pp. 3121–3124.
Published
2026-09-30