Most clinical AI models are only tested when every scan is present. In the real world modalities go missing, and some models fail without ever raising a flag. This framework was built to catch those silent failures before deployment.