Cross Validated
2026-08-22 10:36 UTC
By Malkoun
AI-113-20260822-social-media-5815969b
A possibly short proof of the bias-variance decomposition formula
I am reading the proof of the bias-variance decomposition formula on Wikipedia ( Bias-variance tradeoff ), and it appears to me that the derivation there is uselessly complicated. I think I can derive it in a much shorter way, but I am not sure if it is a 100% correct. Suppose that $y = f(x) + \epsilon$ , where $\epsilon$ is a random variable with mean equal to $0$ and variance $\sigma^2$ . Let $D$ be the training data, composed of $(x_i, y_i)$ , for, say, $i = 1, \dots, n$ , where the set of $(x, y)$ follows a joint density function $P(x, y)$ . We take a sample $(x, y)$ and denote by $\hat{f}(x; D)$ the predicted value corresponding to $x$ using some model trained on $D$ . The $\epsilon_i$ corresponding to the $i$ -th observation point is assumed to be independent of, say $\epsilon_j$ for $j \neq i$ ( $1 \leq i, j \leq n$ ) and also independent of $\epsilon$ . The bias-variance decomposition then says that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_D[\hat{f}(x; D)] + \sigma^2. $$ The way I would proceed to prove it is as follows. Given a random variable $Z$ , we have that $$ \mathbb{E}[Z^2] = \mathbb{E}[Z]^2 + \operatorname{Var}(Z).$$ We apply that with $Z = y - \hat{f}(x; D)$ . We obtain that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_{D, \epsilon}[y - \hat{f}(x; D)]. $$ Moreover if $U$ and $V$ are independent random variables and $\al…
I am reading the proof of the bias-variance decomposition formula on Wikipedia ( Bias-variance tradeoff ), and it appears to me that the derivation there is uselessly complicated. I think I can derive it in a much shorter way, but I am not sure if it is a 100% correct. Suppose that $y = f(x) + \epsilon$ , where $\epsilon$ is a random variable with mean equal to $0$ and variance $\sigma^2$ . Let $D$ be the training data, composed of $(x_i, y_i)$ , for, say, $i = 1, \dots, n$ , where the set of $(x, y)$ follows a joint density function $P(x, y)$ . We take a sample $(x, y)$ and denote by $\hat{f}(x; D)$ the predicted value corresponding to $x$ using some model trained on $D$ . The $\epsilon_i$ corresponding to the $i$ -th observation point is assumed to be independent of, say $\epsilon_j$ for $j \neq i$ ( $1 \leq i, j \leq n$ ) and also independent of $\epsilon$ . The bias-variance decomposition then says that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_D[\hat{f}(x; D)] + \sigma^2. $$ The way I would proceed to prove it is as follows. Given a random variable $Z$ , we have that $$ \mathbb{E}[Z^2] = \mathbb{E}[Z]^2 + \operatorname{Var}(Z).$$ We apply that with $Z = y - \hat{f}(x; D)$ . We obtain that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_{D, \epsilon}[y - \hat{f}(x; D)]. $$ Moreover if $U$ and $V$ are independent random variables and $\al…
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com