PERFORMANCE COMPARISON OF MISSFOREST AND MICE IN HANDLING MISSING CATEGORICAL DATA

  • Nurhidayah Nurhidayah Statistics and Data Science Program, School of Data Science, Mathematics, and Informatics, IPB University, Indonesia https://orcid.org/0009-0001-1279-3971
  • Kusman Sadik Statistics and Data Science Program, School of Data Science, Mathematics, and Informatics, IPB University, Indonesia https://orcid.org/0000-0001-8361-8057
  • Aji Hamim Wigena Statistics and Data Science Program, School of Data Science, Mathematics, and Informatics, IPB University, Indonesia https://orcid.org/0009-0008-6723-2347
Keywords: Categorical Data, Imputation, Missing Data

Abstract

Missing data represent a common challenge in statistical modeling and can substantially reduce the performance of classification algorithms. This study examines the impact of missing values on the performance of the XGBoost model by considering different proportions of missingness (50% and 75%) and various combinations of affected variables under the Missing Completely at Random (MCAR) and Missing at Random (MAR) mechanisms. Two imputation methods, MissForest and Multiple imputation by chained equations (MICE), were compared with a baseline model without imputation. The analysis of variance revealed that the interaction between imputation method, missing data proportion, and variable combinations had a significant effect on accuracy and specificity, while sensitivity remained relatively stable across scenarios. Tukey tests confirmed that MissForest consistently outperformed the other approaches, producing the highest accuracy and specificity, especially at 50% missingness with three variables affected. Moreover, the evaluation of categorical distributions before and after imputation indicated that MissForest better preserved category balance compared to MICE. These findings highlight that the performance of imputation methods strongly depends on the characteristics of missing data. Overall, MissForest demonstrated clear superiority in handling missing categorical data, maintaining distributional integrity while enhancing the classification performance of the XGBoost model. This study advances statistical learning by giving empirical evidence and practical tips for choosing strong imputation strategies in categorical datasets. This improves the reliability of predictive modeling when data is missing.

 

Downloads

Download data is not yet available.

References

A. C. Miola, “COMPARING CATEGORICAL VARIABLES IN CLINICAL AND EXPERIMENTAL STUDIES,” vol. 7301, 2022.

N. Nezami, P. Haghighat, D. Gándara, and H. Anahideh, “ASSESSING DISPARITIES IN PREDICTIVE MODELING OUTCOMES FOR COLLEGE STUDENT SUCCESS: THE IMPACT OF IMPUTATION TECHNIQUES ON MODEL PERFORMANCE AND FAIRNESS,” Educ. Sci., vol. 14, no. 2, p. 136, Jan. 2024. doi: https://doi.org/10.3390/educsci14020136.

I. M. Sumertajaya, E. Rohaeti, A. H. Wigena, and K. Sadik, “VECTOR AUTOREGRESSIVE-MOVING AVERAGE IMPUTATION ALGORITHM FOR HANDLING MISSING DATA IN MULTIVARIATE TIME SERIES,” IAENG Int. J. Comput. Sci., vol. 50, no. 2, 2023, [Online]. Available: https://www.researchgate.net/profile/I-Made-Sumertajaya/publication/371322705_Vector_Autoregressive-Moving_Average_Imputation_Algorithm_for_Handling_Missing_Data_in_Multivariate_Time_Series/links/647f2db979a722376513992f/Vector-Autoregressive-Moving-Average-Imputation-Algorithm-for-Handling-Missing-Data-in-Multivariate-Time-Series.pdf

O. Akande, F. Li, and J. Reiter, “AN EMPIRICAL COMPARISON OF MULTIPLE IMPUTATION METHODS FOR CATEGORICAL DATA,” Am. Stat., vol. 71, no. 2, pp. 162–170, 2017. doi: https://doi.org/10.1080/00031305.2016.1277158.

Y. Sun, J. Li, Y. Xu, T. Zhang, and X. Wang, “DEEP LEARNING VERSUS CONVENTIONAL METHODS FOR MISSING DATA IMPUTATION: A REVIEW AND COMPARATIVE STUDY,” Expert Syst. Appl., vol. 227, no. August 2022, 2023. doi: https://doi.org/10.1016/j.eswa.2023.120201.

D. B. Little, Roderick J. A. ; Rubin, STATISTICAL ANALYSIS WITH MISSING DATA. 2019. doi: https://doi.org/10.1002/9781119482260

P. Ayokunle, J. Tapamo, and A. G. Honor, “EFFECTIVE AND EFFICIENT HANDLING OF MISSING DATA IN SUPERVISED MACHINE LEARNING,” vol. 8, no. December 2024, pp. 361–373, 2025. doi: https://doi.org/10.1016/j.dsm.2024.12.002.

D.-H. Lee, S.-E. Woo, M.-W. Jung, and T.-Y. Heo, “EVALUATION OF ODOR PREDICTION MODEL PERFORMANCE AND VARIABLE IMPORTANCE ACCORDING TO VARIOUS MISSING IMPUTATION METHODS,” Appl. Sci., vol. 12, no. 6, p. 2826, Mar. 2022. doi: https://doi.org/10.3390/app12062826.

T. Shadbahr et al., “THE IMPACT OF IMPUTATION QUALITY ON MACHINE LEARNING CLASSIFIERS FOR DATASETS WITH MISSING VALUES,” pp. 1–15, 2023. doi: https://doi.org/10.1038/s43856-023-00356-z.

I. Nirmala, H. Wijayanto, and K. A. Notodiputro, “PREDICTION OF UNDERGRADUATE STUDENT’S STUDY COMPLETION STATUS USING MISSFOREST IMPUTATION IN RANDOM FOREST AND XGBOOST MODELS,” ComTech Comput. Math. Eng. Appl., vol. 13, no. 1, pp. 53–62, Feb. 2022. doi: https://doi.org/10.21512/comtech.v13i1.7388.

L. O. Joel, W. Doorsamy, and B. S. Paul, “ON THE PERFORMANCE OF IMPUTATION TECHNIQUES FOR MISSING VALUES ON HEALTHCARE DATASETS,” pp. 1–20, 2024, [Online]. Available: http://arxiv.org/abs/2403.14687

J. Schwerter, K. Gurtskaia, A. Romero, B. Zeyer-Gliozzo, and M. Pauly, “EVALUATING TREE-BASED IMPUTATION METHODS AS AN ALTERNATIVE TO MICE PMM FOR DRAWING INFERENCE IN EMPIRICAL STUDIES,” 2024, [Online]. Available: http://arxiv.org/abs/2401.09602

Y.-H. Hu, R.-Y. Wu, Y.-C. Lin, and T.-Y. Lin, “A NOVEL MISSFOREST-BASED MISSING VALUES IMPUTATION APPROACH WITH RECURSIVE FEATURE ELIMINATION IN MEDICAL APPLICATIONS,” BMC Med. Res. Methodol., vol. 24, no. 1, p. 269, Nov. 2024. doi: https://doi.org/10.1186/s12874-024-02392-2.

P. C. Austin, I. R. White, D. S. Lee, and S. van Buuren, “MISSING DATA IN CLINICAL RESEARCH: A TUTORIAL ON MULTIPLE IMPUTATION,” Can. J. Cardiol., vol. 37, no. 9, pp. 1322–1331, Sep. 2021. doi: https://doi.org/10.1016/j.cjca.2020.11.010.

D. Dey, S. Haque, M. Islam, U. I. Aishi, and S. S. Shammy, “THE PROPER APPLICATION OF LOGISTIC REGRESSION MODEL IN COMPLEX SURVEY DATA : A SYSTEMATIC REVIEW,” BMC Med. Res. Methodol., vol. 5, 2025. doi: https://doi.org/10.1186/s12874-024-02454-5.

X. Zhang, “HOW TO GENERATE MISSING DATA FOR SIMULATION STUDIES,” vol. 19, no. 2, pp. 100–122, 2023. doi: https://doi.org/10.20982/tqmp.19.2.p100

S. Farhadpour, T. A. Warner, and A. E. Maxwell, “SELECTING AND INTERPRETING MULTICLASS LOSS AND ACCURACY ASSESSMENT METRICS FOR CLASSIFICATIONS WITH CLASS IMBALANCE : GUIDANCE AND BEST PRACTICES,” pp. 1–22, 2024. doi: https://doi.org/10.3390/rs16030533

P. Buczak, J.-J. Chen, and M. Pauly, “ANALYZING THE EFFECT OF IMPUTATION ON CLASSIFICATION PERFORMANCE UNDER MCAR AND MAR MISSING MECHANISMS,” Entropy, vol. 25, no. 3, p. 521, Mar. 2023. doi: https://doi.org/10.3390/e25030521.

N. B. Mendoza et al., “EVALUATING IMPUTATION METHODS TO IMPROVE PREDICTION ACCURACY FOR AN HIV STUDY IN UGANDA,” pp. 1405–1420, 2024. doi: https://doi.org/10.3390/stats7040082

M. Le Morvan and G. Varoquaux, “IMPUTATION FOR PREDICTION: BEWARE OF DIMINISHING RETURNS,” 13th Int. Conf. Learn. Represent. ICLR 2025, pp. 27597–27636, 2025.

T. Chen and C. Guestrin, “XGBOOST,” IN PROCEEDINGS OF THE 22ND ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, New York, NY, USA: ACM, Aug. 2016, pp. 785–794. doi: https://doi.org/10.1145/2939672.2939785.

D. Aulia and R. Hendri, “XGBOOST IN HANDLING MISSING VALUES FOR LIFE INSURANCE RISK PREDICTION,” SN Appl. Sci., vol. 2, no. 8, pp. 1–10, 2020. doi: https://doi.org/10.1007/s42452-020-3128-y.

Published
2026-08-24
How to Cite
[1]
N. Nurhidayah, K. Sadik, and A. H. Wigena, “PERFORMANCE COMPARISON OF MISSFOREST AND MICE IN HANDLING MISSING CATEGORICAL DATA”, BAREKENG: J. Math. & App., vol. 20, no. 4, pp. 3271-3282, Aug. 2026.