Generative artificial intelligence (GenAI) has rapidly entered educational settings, yet a fundamental question remains unresolved: does GenAI enhance learning or does it improve immediate performance while reducing the cognitive activity on which durable learning depends? This Hypothesis and Theory article offers a theoretical reading of apparently contradictory findings—improved academic products alongside qualitative, neuroscientific, and behavioral signs of reduced cognitive engagement—which we term the critical-thinking paradox of GenAI-integrated learning. These contrasts may also reflect genuine heterogeneity across tasks, populations, and tools; the proposed convergence is treated as a testable interpretation, not a fact. Drawing on four theoretical traditions—levels of processing, desirable difficulties, cognitive load theory and Load Reduction Instruction, and cognitive offloading research—we propose a differentiated three-level framework that maps AI-integration strategies onto surface, intermediate, and deep cognitive processing, specifying level-appropriate AI roles, primary risks, and boundary conditions. We adopt the emerging construct of cognitive debt: a potential cumulative reduction in metacognitive calibration and unaided higher-order performance that persists beyond an AI-assisted episode. Our contribution is to distinguish episodic offloading (deliberate and task-specific) from habitual offloading (routine and weakly monitored) and to map both patterns onto the three cognitive levels. The framework generates falsifiable hypotheses, centrally that unrestricted AI use on deep-processing tasks may yield a product–process dissociation: higher-rated assignments but lower unaided delayed transfer. We specify developmental stage, prior knowledge, and metacognitive monitoring accuracy as preregistered boundary conditions and outline a research program combining confirmatory experiments, interaction telemetry, and longitudinal measurement.

