IntroductionThe flower of Camellia petelotii (Merr.) Sealy is highly valued in Traditional Chinese Medicine, making it susceptible to adulteration with its leaves or other lower-cost adulterants. This study aims to develop a robust authentication approach to distinguish different plant parts and identify adulterated mixtures using ATR-FTIR spectroscopy integrated with chemometrics and artificial intelligence.MethodsThe flowers, leaves, seeds, and mixed samples of C. petelotii were subjected to ATR-FTIR spectroscopy. The spectral data of pure plant parts were analyzed with Principal Component Analysis (PCA) and Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA). In the machine learning phase, a Convolutional Neural Network (CNN) model was trained on spectra of pure plant parts and laboratory-prepared mixed samples. The Synthetic Minority Oversampling Technique (SMOTE) was employed to address class imbalance and sample scarcity. Model robustness was validated via Repeated Random Subsampling Validation (RRSV).ResultsATR-FTIR analysis clearly differentiated the leaves from other plant parts, revealing distinct spectral characteristics. The PCA and OPLS-DA effectively classified the three distinct plant parts with high accuracy, sensitivity, and specificity scores exceeding 95%. The OPLS-DA model achieved high internal validity (R2X, R2Y, and Q2Y ≥ 0.738). In the machine learning workflow, the baseline CNN models suffered from minority-class collapse. Implementing SMOTE effectively resolved this issue. In multi-class configurations, the SMOTE-trained models showed high predictive precision and achieved high F1-scores for Seed (0.846) and Flower (0.931) classes, but moderate performance on Mix (0.593) and Leaf (0.657) classes. The binary classification model resolved the ambiguity, increasing the average F1-scores of the Mix and Leaf classes to 0.674 and 0.932, respectively. When validated against non-augmented data across both multi-class and binary-class configurations, SMOTE-trained models showed high stability for pure plant parts but remained sensitive to the Mix class across both multi-class (F1-score: 0.361) and binary-class (F1-score: 0.249) configurations.ConclusionWhile AI-driven data augmentation mitigates sample-size constraints, classification performance remains heavily influenced by spectral characteristics. The binary classification architecture offered distinct advantages when analyzing samples with heterogeneous spectral complexity. This integrated approach may serve as a reference for quality control and authentication of botanical products.
An integrative approach for rapid authentication of different parts of Camellia petelotii (Merr.) Sealy: combining ATR-FTIR spectroscopy with conventional chemometric analysis and a convolutional neural network
Mun Fei Yam

