I’m working on a computational social science project using comments from Weibo and Douyin (Chinese TikTok) for text and sentiment analysis. My study looks at how online discourse around a female public figure changed before and after a particular event. I first define several keywords related to the person and the event, search for relevant posts/videos, and then collect comments from selected posts. The problem I’m struggling with is: how should I systematically decide which posts to select? Taking the first 10 or 20 search results does not seem reliable, because search results on these platforms are algorithmically ranked and change between searches. The posts ranked at the top are also not necessarily the ones with the most discussion. Selecting posts purely by comment count is also problematic. Some videos have thousands of comments, but many are just emojis, users tagging friends, repeated/copied comments, or spam. Meanwhile, another video with fewer comments may contain much more substantive discussion relevant to my research. So in practice, I inevitably have to make a human judgment about whether a comment section contains enough meaningful, relevant text to analyze. But this creates the methodological problem I’m most worried about: Why these particular posts? If I selected another set of relevant posts, would I still get the same result? Even if I eventually collect millions of comments, I feel that a large number of comments does not solve the problem if I cannot…

Full article content could not be extracted automatically. Read the original below.