Cross Validated
2026-09-10 19:13 UTC
By always.learning
AI-113-20260910-social-media-606adead
LASSO to identify which of ~50 screening variables relate to a biomarker, with n = 100?
I have baseline data on 100 participants from an ongoing longitudinal study. At the screening visit each participant answered sociodemographic, clinical and lifestyle questions and completed a cognitive screening battery — more than 50 candidate variables in total. A blood sample drawn at the same visit gives the serum concentration of a protein of interest, which is my (continuous) outcome. For a secondary, exploratory analysis I want to identify which of these screening items and test scores are more strongly associated with the biomarker concentration. However, this is not a prediction problem: I am not building a model to apply to new individuals, but trying to establish which of the screening variables are associated with the biomarker. Since the number of candidate variables is large relative to the sample size (n/p ≈ 2), an unpenalised model with all variables is not viable and stepwise selection is widely discouraged, so my first thought was LASSO with λ chosen by cross-validation. I am unsure this is defensible: selection at this ratio seems likely to be unstable, and inference on LASSO-selected variables is not valid. Is LASSO appropriate when the goal is identifying associated variables rather than predicting new observations? With n = 100 and ~50 candidate predictors, is the selected set stable enough to support any substantive association? If it is not, what would you recommend instead?
I have baseline data on 100 participants from an ongoing longitudinal study. At the screening visit each participant answered sociodemographic, clinical and lifestyle questions and completed a cognitive screening battery — more than 50 candidate variables in total. A blood sample drawn at the same visit gives the serum concentration of a protein of interest, which is my (continuous) outcome. For a secondary, exploratory analysis I want to identify which of these screening items and test scores are more strongly associated with the biomarker concentration. However, this is not a prediction problem: I am not building a model to apply to new individuals, but trying to establish which of the screening variables are associated with the biomarker. Since the number of candidate variables is large relative to the sample size (n/p ≈ 2), an unpenalised model with all variables is not viable and stepwise selection is widely discouraged, so my first thought was LASSO with λ chosen by cross-validation. I am unsure this is defensible: selection at this ratio seems likely to be unstable, and inference on LASSO-selected variables is not valid. Is LASSO appropriate when the goal is identifying associated variables rather than predicting new observations? With n = 100 and ~50 candidate predictors, is the selected set stable enough to support any substantive association? If it is not, what would you recommend instead?
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com