Abstract Accurate classification of imbalanced soil data remains a key challenge in Digital Soil Mapping (DSM). We propose a two-step preprocessing strategy that integrates feature selection with resampling to enhance soil class prediction. Using 140 soil profiles from the Darab region (Iran), four feature selection methods (VIF, PCA, Boruta, Expert Viewpoint) were combined with the Synthetic Minority Oversampling Technique (SMOTE) and evaluated with three classifiers (C5.0, Random Forest, XGBoost). Among the tested combinations, VIF + SMOTE with the Random Forest (RF) model achieved the best performance, increasing macro F1 by 15.14% and raising the number of detectable classes from four to five. These gains were class-dependent: SMOTE improved detection for classes with stable environmental signatures (e.g., Haplosalids) but failed for geopedologically complex classes (e.g., Calciusterts). This suggests that post-resampling sample size alone is insufficient and that covariate separability and environmental context are decisive factors. Balancing also shifted variable importance toward topographic and geomorphological predictors, while reducing reliance on weaker covariates. To better address hard-to-separate classes, we recommend boundary-aware resampling (e.g., Borderline-SMOTE, SMOTE-ENN, ADASYN) and covariate enrichment (e.g., horizon descriptors). The resulting soil maps revealed clearer spatial differentiation of rare but management-relevant classes, thereby supporting more targeted and sustainable land-use decisions.

