Cross Validated
2026-09-25 12:51 UTC
By Magnus
AI-113-20260925-social-media-84d207e9
Binary response variable with one continuous independent variable in small sample?
I'm writing my medical thesis and I want to test our research group's hypothesis that the width of the foramen magnum can predict response or non-response to a certain surgical intervention in a neurological disease. The sample is however quite small (only 23 patients) so any biological signal whould have to be quite strong to cut through the noise. I asked stats support who wanted me to "find some exact version of the ROC analysis, some analogue to Fisher's exact test", but I can't really find anything like that in the literature. Generative AI tends to point me towards Mann-Whitney U-tests, exact logistic regression and/or Firth regression, but cannot point me towards any supporting literature. Intuitively that's not a bad idea. I might test if the FM width is normally distributed with Shapiro-Wilk, then test for differences with t-test/U-test depending on distribution and only then proceed towards ROC-analyses/logistic regression if an interesting pattern emerges. But I will have to motivate the chosen approach formally. Is there an actual recommended approach? Is there any literature discussing the pros and cons of different strategies given this particular dilemma?
I'm writing my medical thesis and I want to test our research group's hypothesis that the width of the foramen magnum can predict response or non-response to a certain surgical intervention in a neurological disease. The sample is however quite small (only 23 patients) so any biological signal whould have to be quite strong to cut through the noise. I asked stats support who wanted me to "find some exact version of the ROC analysis, some analogue to Fisher's exact test", but I can't really find anything like that in the literature. Generative AI tends to point me towards Mann-Whitney U-tests, exact logistic regression and/or Firth regression, but cannot point me towards any supporting literature. Intuitively that's not a bad idea. I might test if the FM width is normally distributed with Shapiro-Wilk, then test for differences with t-test/U-test depending on distribution and only then proceed towards ROC-analyses/logistic regression if an interesting pattern emerges. But I will have to motivate the chosen approach formally. Is there an actual recommended approach? Is there any literature discussing the pros and cons of different strategies given this particular dilemma?
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com