IntroductionDyslexia affects around 10% of the global population, yet early screening depends heavily on resource-intensive clinical assessments that are difficult to scale. Gamified cognitive platforms offer a practical alternative, but most existing machine learning approaches rely on raw task metrics without exploiting the hierarchical cognitive structure of multi-task assessment batteries.MethodsNEUROCOG-DysNet is a multi-scale cognitive feature engineering and stacking ensemble framework for dyslexia screening using gamified assessment data from the Dytective platform (N = 3,644; 392 dyslexic, 3,252 non-dyslexic). The framework extracts 258 hierarchical features across five cognitive levels-task, phase, temporal, global, and cross-domain-from raw game interactions, supplemented by four demographic variables. LightGBM-based feature selection retains the 50 most discriminative features. Seven base classifiers, configured to emphasize high sensitivity, balanced performance, or high AUC, are combined using five ensemble strategies and evaluated under a leakage-safe nested cross-validation protocol.ResultsThe stacking meta-learner achieves the highest composite score and attains an AUC-ROC of 0.8924 (95% CI: 0.8543–0.9253), sensitivity of 0.8077, specificity of 0.8187, and balanced accuracy of 0.8132. DeLong's test (p = 0.6946) and McNemar's test (p = 0.1443) indicate that AUC gains over the strongest individual baseline (XGB_HighAUC) are not statistically significant. Nested cross-validation confirms a small optimism gap (+0.011 AUC). Clinical utility analysis yields a 5.25 × lift in the highest-risk decile and a negative predictive value of 0.9726. Ablation shows that demographic variables contribute more to discrimination than any engineered cognitive level; a demographics-free variant attains AUC 0.7948.DiscussionNEUROCOG-DysNet provides a tunable sensitivity-specificity balance and well-calibrated risk scores suitable for a first-stage screening aid. The ensemble's advantage over fixed-weight alternatives is marginal on this dataset, and the cross-domain feature level is effectively inert. External validation across independent language cohorts and prospective multi-center evaluation are required before clinical deployment.