Cross Validated
2018-04-11 12:17 UTC
By DaveTheAl
AI-113-20180411-social-media-44a6b990
Correctness of a skewed cosine similarity graph
I am currently implementing a word2vec model that uses the cosine similarity to determine the similarity between two vectors. When plotting all the possible cosine similarities, I get the following graph: The graph is pretty skewed towards the 'most similar value'. However, I am not sure if this is an ok thing, which can just be corrected by normalizing the data, or if this skewness indicates that something is terribly wrong with the model. I am not an expert on the topic of natural language processing. Can you guys provide me with some intuition what if it is ok to normalize, or if this indicates if the model is terribly wrong? Any ideas and insights are appreciated.
I am currently implementing a word2vec model that uses the cosine similarity to determine the similarity between two vectors. When plotting all the possible cosine similarities, I get the following graph: The graph is pretty skewed towards the 'most similar value'. However, I am not sure if this is an ok thing, which can just be corrected by normalizing the data, or if this skewness indicates that something is terribly wrong with the model. I am not an expert on the topic of natural language processing. Can you guys provide me with some intuition what if it is ok to normalize, or if this indicates if the model is terribly wrong? Any ideas and insights are appreciated.
Full article content could not be extracted automatically. Read the original below.
Source:
Cross Validated
· stats.stackexchange.com