Cross Validated
2026-09-18 22:05 UTC
By amongus
AI-113-20260918-social-media-d215118f
"Conditional / Joint / Marginal" Likelihoods
I'm currently learning about the EM algorithm (in the context of filling in missing data). I'm trying to understand why specifying the "joint" likelihood is a fine thing to do. This has made me realize, I don't think I understand what exactly is the likelihood; I would love any clarification! I think my confusion comes from the following: In the context of regression . Suppose you have a dataset $D = \{(x_i, y_i)\}_{i=1}^n$ that captures input $x_i \in X$ and output $y_i \in Y$ relationships, and a data generating process $y_i = f_\theta(x_i) + \varepsilon_i$ (s.t. $f_\theta : X \to Y$ and $\theta \in \Theta$ ). If the noise $\varepsilon_i$ is drawn $\text{iid}$ , then $$ \begin{align} p(y\mid x, \theta) = \prod_{(x_i, y_i) \in D} p(y_i \mid x_i, \theta) \end{align} $$ We would call the (conditional) likelihood the function $L_\text{cond}: \Theta \to \mathbb R$ s.t. $L_\text{cond}(\theta) = p(y\mid x, \theta)$ with the points in the dataset $D$ plugged into the arguments. In the context of missing data . Now suppose you have an additional $D_\text{missing} = \{y_i\}_{i=1}^m$ . The conditional likelihood is misspecified here because of the unpaired data in $D_\text{missing}$ . I'm told we instead specify we inspect the joint distribution. Once again, assuming the noise is $\text{iid}$ , then $$ \begin{align} p(x,y\mid \theta) = \left(\prod_{(x_i, y_i) \in D}p(x_i, y_i \mid \theta)\right) \left(\prod_{y_i \in D_\text{missing}} p(y_i \mid \theta) \right) \end{align} $$ The EM a…
I'm currently learning about the EM algorithm (in the context of filling in missing data). I'm trying to understand why specifying the "joint" likelihood is a fine thing to do. This has made me realize, I don't think I understand what exactly is the likelihood; I would love any clarification! I think my confusion comes from the following: In the context of regression . Suppose you have a dataset $D = \{(x_i, y_i)\}_{i=1}^n$ that captures input $x_i \in X$ and output $y_i \in Y$ relationships, and a data generating process $y_i = f_\theta(x_i) + \varepsilon_i$ (s.t. $f_\theta : X \to Y$ and $\theta \in \Theta$ ). If the noise $\varepsilon_i$ is drawn $\text{iid}$ , then $$ \begin{align} p(y\mid x, \theta) = \prod_{(x_i, y_i) \in D} p(y_i \mid x_i, \theta) \end{align} $$ We would call the (conditional) likelihood the function $L_\text{cond}: \Theta \to \mathbb R$ s.t. $L_\text{cond}(\theta) = p(y\mid x, \theta)$ with the points in the dataset $D$ plugged into the arguments. In the context of missing data . Now suppose you have an additional $D_\text{missing} = \{y_i\}_{i=1}^m$ . The conditional likelihood is misspecified here because of the unpaired data in $D_\text{missing}$ . I'm told we instead specify we inspect the joint distribution. Once again, assuming the noise is $\text{iid}$ , then $$ \begin{align} p(x,y\mid \theta) = \left(\prod_{(x_i, y_i) \in D}p(x_i, y_i \mid \theta)\right) \left(\prod_{y_i \in D_\text{missing}} p(y_i \mid \theta) \right) \end{align} $$ The EM a…
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com