The documented SmolLM3 fragility[3] (INT4 attack-success 34.5%→44.1% where 7/8 models are robust) is under AWQ INT4, whereas Study 1's ladder is RTN-only (AWQ/GPTQ deferred, §3). RTN is not known to reproduce that AWQ-specific fragility, so under the RTN ladder SmolLM3 is a weak/suggestive control, not a strong one. A second mismatch: the documented fragility is an attack-success-rate endpoint, while we read the control on E1 (exits) for comparability, so even the AWQ-arm control is indirect. Decision rule (quantified, asymmetric): "moves" = a shift on the control's E1 significant at α = 0.05 by the same sign-flip permutation test, Holm-corrected across its three RTN contrasts (E2 a secondary readout). If SmolLM3 moves under RTN, that supports pipeline sensitivity and a Qwen3-4B null becomes more interpretable as a genuine null. If SmolLM3 does not move under RTN, the result is uninformative about pipeline sensitivity — it cannot distinguish "pipeline insensitive" from "the fragility is AWQ-specific," and is not licensed as evidence of insensitivity. The strong form of this control requires SmolLM3 under its documented-fragile condition (AWQ w4), which becomes available with — and is registered alongside — the deferred GPTQ/AWQ method-comparison arm (see Design).
Does post-training quantization change welfare-relevant indicators in open-weight language models?
Topics: