PERFORMANCE COMPARISON OF MISSFOREST AND MICE IN HANDLING MISSING CATEGORICAL DATA
Abstract
Missing data represent a common challenge in statistical modeling and can substantially reduce the performance of classification algorithms. This study examines the impact of missing values on the performance of the XGBoost model by considering different proportions of missingness (50% and 75%) and various combinations of affected variables under the Missing Completely at Random (MCAR) and Missing at Random (MAR) mechanisms. Two imputation methods, MissForest and Multiple imputation by chained equations (MICE), were compared with a baseline model without imputation. The analysis of variance revealed that the interaction between imputation method, missing data proportion, and variable combinations had a significant effect on accuracy and specificity, while sensitivity remained relatively stable across scenarios. Tukey tests confirmed that MissForest consistently outperformed the other approaches, producing the highest accuracy and specificity, especially at 50% missingness with three variables affected. Moreover, the evaluation of categorical distributions before and after imputation indicated that MissForest better preserved category balance compared to MICE. These findings highlight that the performance of imputation methods strongly depends on the characteristics of missing data. Overall, MissForest demonstrated clear superiority in handling missing categorical data, maintaining distributional integrity while enhancing the classification performance of the XGBoost model. This study advances statistical learning by giving empirical evidence and practical tips for choosing strong imputation strategies in categorical datasets. This improves the reliability of predictive modeling when data is missing.
Downloads
References
A. C. Miola, “COMPARING CATEGORICAL VARIABLES IN CLINICAL AND EXPERIMENTAL STUDIES,” vol. 7301, 2022.
N. Nezami, P. Haghighat, D. Gándara, and H. Anahideh, “ASSESSING DISPARITIES IN PREDICTIVE MODELING OUTCOMES FOR COLLEGE STUDENT SUCCESS: THE IMPACT OF IMPUTATION TECHNIQUES ON MODEL PERFORMANCE AND FAIRNESS,” Educ. Sci., vol. 14, no. 2, p. 136, Jan. 2024. doi: https://doi.org/10.3390/educsci14020136.
I. M. Sumertajaya, E. Rohaeti, A. H. Wigena, and K. Sadik, “VECTOR AUTOREGRESSIVE-MOVING AVERAGE IMPUTATION ALGORITHM FOR HANDLING MISSING DATA IN MULTIVARIATE TIME SERIES,” IAENG Int. J. Comput. Sci., vol. 50, no. 2, 2023, [Online]. Available: https://www.researchgate.net/profile/I-Made-Sumertajaya/publication/371322705_Vector_Autoregressive-Moving_Average_Imputation_Algorithm_for_Handling_Missing_Data_in_Multivariate_Time_Series/links/647f2db979a722376513992f/Vector-Autoregressive-Moving-Average-Imputation-Algorithm-for-Handling-Missing-Data-in-Multivariate-Time-Series.pdf
O. Akande, F. Li, and J. Reiter, “AN EMPIRICAL COMPARISON OF MULTIPLE IMPUTATION METHODS FOR CATEGORICAL DATA,” Am. Stat., vol. 71, no. 2, pp. 162–170, 2017. doi: https://doi.org/10.1080/00031305.2016.1277158.
Y. Sun, J. Li, Y. Xu, T. Zhang, and X. Wang, “DEEP LEARNING VERSUS CONVENTIONAL METHODS FOR MISSING DATA IMPUTATION: A REVIEW AND COMPARATIVE STUDY,” Expert Syst. Appl., vol. 227, no. August 2022, 2023. doi: https://doi.org/10.1016/j.eswa.2023.120201.
D. B. Little, Roderick J. A. ; Rubin, STATISTICAL ANALYSIS WITH MISSING DATA. 2019. doi: https://doi.org/10.1002/9781119482260
P. Ayokunle, J. Tapamo, and A. G. Honor, “EFFECTIVE AND EFFICIENT HANDLING OF MISSING DATA IN SUPERVISED MACHINE LEARNING,” vol. 8, no. December 2024, pp. 361–373, 2025. doi: https://doi.org/10.1016/j.dsm.2024.12.002.
D.-H. Lee, S.-E. Woo, M.-W. Jung, and T.-Y. Heo, “EVALUATION OF ODOR PREDICTION MODEL PERFORMANCE AND VARIABLE IMPORTANCE ACCORDING TO VARIOUS MISSING IMPUTATION METHODS,” Appl. Sci., vol. 12, no. 6, p. 2826, Mar. 2022. doi: https://doi.org/10.3390/app12062826.
T. Shadbahr et al., “THE IMPACT OF IMPUTATION QUALITY ON MACHINE LEARNING CLASSIFIERS FOR DATASETS WITH MISSING VALUES,” pp. 1–15, 2023. doi: https://doi.org/10.1038/s43856-023-00356-z.
I. Nirmala, H. Wijayanto, and K. A. Notodiputro, “PREDICTION OF UNDERGRADUATE STUDENT’S STUDY COMPLETION STATUS USING MISSFOREST IMPUTATION IN RANDOM FOREST AND XGBOOST MODELS,” ComTech Comput. Math. Eng. Appl., vol. 13, no. 1, pp. 53–62, Feb. 2022. doi: https://doi.org/10.21512/comtech.v13i1.7388.
L. O. Joel, W. Doorsamy, and B. S. Paul, “ON THE PERFORMANCE OF IMPUTATION TECHNIQUES FOR MISSING VALUES ON HEALTHCARE DATASETS,” pp. 1–20, 2024, [Online]. Available: http://arxiv.org/abs/2403.14687
J. Schwerter, K. Gurtskaia, A. Romero, B. Zeyer-Gliozzo, and M. Pauly, “EVALUATING TREE-BASED IMPUTATION METHODS AS AN ALTERNATIVE TO MICE PMM FOR DRAWING INFERENCE IN EMPIRICAL STUDIES,” 2024, [Online]. Available: http://arxiv.org/abs/2401.09602
Y.-H. Hu, R.-Y. Wu, Y.-C. Lin, and T.-Y. Lin, “A NOVEL MISSFOREST-BASED MISSING VALUES IMPUTATION APPROACH WITH RECURSIVE FEATURE ELIMINATION IN MEDICAL APPLICATIONS,” BMC Med. Res. Methodol., vol. 24, no. 1, p. 269, Nov. 2024. doi: https://doi.org/10.1186/s12874-024-02392-2.
P. C. Austin, I. R. White, D. S. Lee, and S. van Buuren, “MISSING DATA IN CLINICAL RESEARCH: A TUTORIAL ON MULTIPLE IMPUTATION,” Can. J. Cardiol., vol. 37, no. 9, pp. 1322–1331, Sep. 2021. doi: https://doi.org/10.1016/j.cjca.2020.11.010.
D. Dey, S. Haque, M. Islam, U. I. Aishi, and S. S. Shammy, “THE PROPER APPLICATION OF LOGISTIC REGRESSION MODEL IN COMPLEX SURVEY DATA : A SYSTEMATIC REVIEW,” BMC Med. Res. Methodol., vol. 5, 2025. doi: https://doi.org/10.1186/s12874-024-02454-5.
X. Zhang, “HOW TO GENERATE MISSING DATA FOR SIMULATION STUDIES,” vol. 19, no. 2, pp. 100–122, 2023. doi: https://doi.org/10.20982/tqmp.19.2.p100
S. Farhadpour, T. A. Warner, and A. E. Maxwell, “SELECTING AND INTERPRETING MULTICLASS LOSS AND ACCURACY ASSESSMENT METRICS FOR CLASSIFICATIONS WITH CLASS IMBALANCE : GUIDANCE AND BEST PRACTICES,” pp. 1–22, 2024. doi: https://doi.org/10.3390/rs16030533
P. Buczak, J.-J. Chen, and M. Pauly, “ANALYZING THE EFFECT OF IMPUTATION ON CLASSIFICATION PERFORMANCE UNDER MCAR AND MAR MISSING MECHANISMS,” Entropy, vol. 25, no. 3, p. 521, Mar. 2023. doi: https://doi.org/10.3390/e25030521.
N. B. Mendoza et al., “EVALUATING IMPUTATION METHODS TO IMPROVE PREDICTION ACCURACY FOR AN HIV STUDY IN UGANDA,” pp. 1405–1420, 2024. doi: https://doi.org/10.3390/stats7040082
M. Le Morvan and G. Varoquaux, “IMPUTATION FOR PREDICTION: BEWARE OF DIMINISHING RETURNS,” 13th Int. Conf. Learn. Represent. ICLR 2025, pp. 27597–27636, 2025.
T. Chen and C. Guestrin, “XGBOOST,” IN PROCEEDINGS OF THE 22ND ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, New York, NY, USA: ACM, Aug. 2016, pp. 785–794. doi: https://doi.org/10.1145/2939672.2939785.
D. Aulia and R. Hendri, “XGBOOST IN HANDLING MISSING VALUES FOR LIFE INSURANCE RISK PREDICTION,” SN Appl. Sci., vol. 2, no. 8, pp. 1–10, 2020. doi: https://doi.org/10.1007/s42452-020-3128-y.
Copyright (c) 2026 Nurhidayah Nurhidayah, Kusman Sadik, Aji Hamim Wigena

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this Journal agree to the following terms:
- Author retain copyright and grant the journal right of first publication with the work simultaneously licensed under a creative commons attribution license that allow others to share the work within an acknowledgement of the work’s authorship and initial publication of this journal.
- Authors are able to enter into separate, additional contractual arrangement for the non-exclusive distribution of the journal’s published version of the work (e.g. acknowledgement of its initial publication in this journal).
- Authors are permitted and encouraged to post their work online (e.g. in institutional repositories or on their websites) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published works.




1.gif)


