In my workplace, I came across an idea that OLS can be retooled to solve ratio metric inference problems. When I think of $k_i$ successes out of $m_i$ trials for individual $i$ , this seems to be a perfect use case for binomial regression. Where your chief interest is estimating the conditional probability of success. Its strength, natively handling non-linearity, is also its weakness: Extracting the marginal probability of success is a nontrivial operation; marginalization via G-computation would be needed: $$logit(K=k|M=m, X=x, d=1) - logit(K=k|M=m, X=x, d=0)$$ And this is a large computational burden to assume with millions or billions of observations. So, I've been pointed to the OLS solution, which I understand to be based on the "delta method", correcting the linear solution with gradient information to accommodate curvature in the nonlinearity (ratio function), through the Taylor Series Expansion. Naively, we have two options for OLS. First, infer in the ratio space directly. But this approach completely mutes the number of trials and biases inference when $corr(K, M)$ exists. $$ \frac{k_i}{m_i} = \alpha + \lambda d_i + \beta X +\epsilon $$ The second naive option is to infer the difference of global ratios directly where $D_j$ is the binary design vector for treatment exposure.This is equally problematic due to a sample size of one. $$ \frac{\sum_{j=1} K_j D_j}{\sum_{j=1} M_j D_j} - \frac{\sum_{j=1} K_j (1-D_j)}{\sum_{j=1} M_j (1-D_j)} = \alpha + \lambda d_i + \beta…

Full article content could not be extracted automatically. Read the original below.