LessWrong AI
2026-09-19 14:27 UTC
By OscarGilg
USR-0152-20260919-community-fo-2902b41a
CommentBench: Can Models Match Human Comments on AI Safety Posts?
TL;DR We measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms. We built a pipeline that goes from a corpus of conceptual documents with comments to a set of target human points. Fable 5 performs best, matching 8.3% of targets, followed by Fable 5.1 (7.5%). We find that performance across models is highly correlated across different settings (LW posts, drafts, shortforms, replies). We checked whether memorisation explained performance. We found no consistent performance advantage on posts published before model’s knowledge cutoffs. All public documents postdate the top-performing model’s knowledge cutoff (Fable 5). CommentBench performance by number of comments from three of the four settings: forum posts, shortforms and Google Doc research drafts. The reply setting is excluded because comments are not ordered. For each document we compute the share of its human target points matched by at least one model-written comment. Each line is the mean of that share across documents, averaged over four samples. Introduction As progress in AI speeds up, we want to make sure that AI labour is used effectively to also differentially accelerate AI safety (we do not argue for this in depth, see Joe Carlsmith and related discussions here , here and here ). One of the capabilities we think is important to accelerate is conceptual reasoning about how to mitigate risks from transformative AI. Comments on blog posts and research dra…
TL;DR We measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms. We built a pipeline that goes from a corpus of conceptual documents with comments to a set of target human points. Fable 5 performs best, matching 8.3% of targets, followed by Fable 5.1 (7.5%). We find that performance across models is highly correlated across different settings (LW posts, drafts, shortforms, replies). We checked whether memorisation explained performance. We found no consistent performance advantage on posts published before model’s knowledge cutoffs. All public documents postdate the top-performing model’s knowledge cutoff (Fable 5). CommentBench performance by number of comments from three of the four settings: forum posts, shortforms and Google Doc research drafts. The reply setting is excluded because comments are not ordered. For each document we compute the share of its human target points matched by at least one model-written comment. Each line is the mean of that share across documents, averaged over four samples. Introduction As progress in AI speeds up, we want to make sure that AI labour is used effectively to also differentially accelerate AI safety (we do not argue for this in depth, see Joe Carlsmith and related discussions here , here and here ). One of the capabilities we think is important to accelerate is conceptual reasoning about how to mitigate risks from transformative AI. Comments on blog posts and research dra…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com