Cross Validated
2026-09-25 19:59 UTC
By adamkostanov
AI-113-20260925-social-media-f44d5e48
How to set up a simulation study?
Here is a hierarchical data generating process (DGP): (Population) Layer 1: $$\theta_i \overset{\text{iid}}{\sim} \text{Beta}(\alpha,\ \beta), \qquad i = 1, \dots, n$$ (Individual) Layer 2: $$y_{ij} \mid \theta_i \overset{\text{iid}}{\sim} \text{Bernoulli}(\theta_i), \qquad j = 1, \dots, k$$ $$\mu = E[\theta_i], \qquad \tau^2 = \operatorname{Var}(\theta_i)$$ Using a sample, my purpose is only to estimate the mean and the variance of the mean estimator. I want to know what values of $n$ and $k$ I should select to get good results (I know this hugely subjective) such that $nk$ is minimized (assume that increasing $n$ by 1 costs the same as increase $k$ by 1). After doing some research, it seems the best way to handle this question is by a simulation study. Assuming $\alpha,\ \beta$ are known, I could sample from the DGP for different combinations of $n$ and $k$ and record the average length of the confidence intervals and average coverage rate at each combination. Since the mean estimator is unbiased regardless of the choice in $n$ or $k$ , I could use average CI length and average coverage rate to make sure I am not getting a misleadingly good coverage rate at the expense of a large CI. Here is how I plan to do this in R (I wrote the code to focus on readability instead of speed - I use a moment based estimator and consider different true combinations of parameters): set.seed(2026) mu_vals $mu[s] tau2 tau2[s] alpha_plus_beta When visualized, the results look like this (direct…
Here is a hierarchical data generating process (DGP): (Population) Layer 1: $$\theta_i \overset{\text{iid}}{\sim} \text{Beta}(\alpha,\ \beta), \qquad i = 1, \dots, n$$ (Individual) Layer 2: $$y_{ij} \mid \theta_i \overset{\text{iid}}{\sim} \text{Bernoulli}(\theta_i), \qquad j = 1, \dots, k$$ $$\mu = E[\theta_i], \qquad \tau^2 = \operatorname{Var}(\theta_i)$$ Using a sample, my purpose is only to estimate the mean and the variance of the mean estimator. I want to know what values of $n$ and $k$ I should select to get good results (I know this hugely subjective) such that $nk$ is minimized (assume that increasing $n$ by 1 costs the same as increase $k$ by 1). After doing some research, it seems the best way to handle this question is by a simulation study. Assuming $\alpha,\ \beta$ are known, I could sample from the DGP for different combinations of $n$ and $k$ and record the average length of the confidence intervals and average coverage rate at each combination. Since the mean estimator is unbiased regardless of the choice in $n$ or $k$ , I could use average CI length and average coverage rate to make sure I am not getting a misleadingly good coverage rate at the expense of a large CI. Here is how I plan to do this in R (I wrote the code to focus on readability instead of speed - I use a moment based estimator and consider different true combinations of parameters): set.seed(2026) mu_vals $mu[s] tau2 tau2[s] alpha_plus_beta When visualized, the results look like this (direct…
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com