Scientific models are commonly evaluated by predictive performance, while many physical quantities of interest are inferred indirectly through multistage model architectures. This paper argues that these achievements require distinct robustness tests and that different kinds of model change must not be conflated. Admissible changes are classified as mathematical reformulations (R0), model-form or approximation choices (R1), or alternative physical hypotheses (R2). Model- specific internal parameters are compared only after mapping them to common physical target spaces. The benchmark distinguishes latent posterior predictive measures from experimentally observable predictive measures, so that physical model differences can be separated from limits of current measurement resolution. Inferential stability is defined primarily on a joint physical pushforward distribution, with marginal distances used to localize which quantities move. Held- out predictive adequacy is tested separately so that two equally poor models cannot count as robust merely because they agree with each other. The audit further requires controls for model discrepancy, discrepancy–parameter confounding and identifiability, discrepancy-prior sensitivity, prior pushforwards, numerical and emulator uncertainty, coordinate and metric choice, and model weighting. Published quark–gluon plasma analyses provide a retrospective stress test: some apparent parameter disagreements are not robust to explicit theory-uncertainty treatment; pre-hydrodynamic physical assumptions can alter extracted transport properties while several final-state observables remain only modestly affected after retuning; and other early-time hypotheses are empirically discriminated. A prospective test is specified for a harmonized Bayesian comparison of free streaming and QCD effective kinetic theory, explicitly treated as an R2 physical-hypothesis comparison and crossed, where feasible, with an independent R1 particlization choice. Harmonization is defined by common physical targets and pre-declared interface criteria rather than by forcing identical model-specific switching parameters. A brief reflexive extension asks whether the same methodological problem can be reconstructed from substantially different disciplinary starting points; such convergence is treated only as a check on path dependence, not as validation of the benchmark. The contribution is methodological: it makes the gap between predictive stability and physical-inference stability a controlled empirical object while preserving the physical status of the model changes being tested.

