Cross Validated
2023-08-03 15:27 UTC
By taellipsis
AI-113-20230803-social-media-b61f5c94
Correlation vs Euclidean distance as measures of similarity or closeness between data points with an outlier
I am interested in the comparison of Pearson correlation and Euclidean distance as measures of similarity between data points. Suppose I have 4 data points, w, x, y, z , in a multidimensional space, where w is a very extreme outlier and x, y, z are highly similar to each other . For the correlation measure, I assume that x, y, z have high correlation coefficients (positive) with each other, but not equal to 1. For the euclidean measure, I assume that x, y, z have small euclidean distances with each other, but not equal to 0. Now, if I take the arithmetic mean of w, x, y, z and call it m , how does the similarity between m and x, y, z change depending on the measure I use? Which measure is more robust to the presence of the outlier w and can still capture the similarity or closeness of x, y, z?
I am interested in the comparison of Pearson correlation and Euclidean distance as measures of similarity between data points. Suppose I have 4 data points, w, x, y, z , in a multidimensional space, where w is a very extreme outlier and x, y, z are highly similar to each other . For the correlation measure, I assume that x, y, z have high correlation coefficients (positive) with each other, but not equal to 1. For the euclidean measure, I assume that x, y, z have small euclidean distances with each other, but not equal to 0. Now, if I take the arithmetic mean of w, x, y, z and call it m , how does the similarity between m and x, y, z change depending on the measure I use? Which measure is more robust to the presence of the outlier w and can still capture the similarity or closeness of x, y, z?
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com