LessWrong AI
2026-07-14 18:06 UTC
By Andy Arditi
USR-0152-20260714-community-fo-b179a397
An analysis of AI-generated content at the Mechanistic Interpretability Workshop
Introduction Over the past few years, AI tools have become useful for conducting technical AI research. In the early ChatGPT era (~2023–2024), chat assistants were maybe useful as sounding boards for research ideas, or as editors for polishing a paper draft. In the more recent Claude Code era, competent coding agents can code up and run experiments; with enough direction, they can do much of the technical heavy lifting on a PhD-level research project ( Schwartz, 2026 ); and in well-defined settings, they can autonomously try new things and iterate without a human in the loop ( Karpathy, 2026 ). While these tools can unlock researchers to be more productive and widen their ambitions, they can also be abused – anyone can now hand an agent a research prompt, tell it to run the experiments and write up the results in a LaTeX document, and get back an artifact that roughly resembles a conference paper in form. With this steep of a change in the research process, it seems important to study how it is affecting technical research and peer review. The Mechanistic Interpretability Workshop has run three times in the past two years – at ICML 2024, NeurIPS 2025, and ICML 2026. The most recent iteration felt quite different from the first two – submissions more than doubled from the previous edition (which was only seven months earlier), and a noticeable share of the submissions seemed to resemble “AI slop” . We, the workshop’s program chairs, collectively read through hundreds of abstr…
Introduction Over the past few years, AI tools have become useful for conducting technical AI research. In the early ChatGPT era (~2023–2024), chat assistants were maybe useful as sounding boards for research ideas, or as editors for polishing a paper draft. In the more recent Claude Code era, competent coding agents can code up and run experiments; with enough direction, they can do much of the technical heavy lifting on a PhD-level research project ( Schwartz, 2026 ); and in well-defined settings, they can autonomously try new things and iterate without a human in the loop ( Karpathy, 2026 ). While these tools can unlock researchers to be more productive and widen their ambitions, they can also be abused – anyone can now hand an agent a research prompt, tell it to run the experiments and write up the results in a LaTeX document, and get back an artifact that roughly resembles a conference paper in form. With this steep of a change in the research process, it seems important to study how it is affecting technical research and peer review. The Mechanistic Interpretability Workshop has run three times in the past two years – at ICML 2024, NeurIPS 2025, and ICML 2026. The most recent iteration felt quite different from the first two – submissions more than doubled from the previous edition (which was only seven months earlier), and a noticeable share of the submissions seemed to resemble “AI slop” . We, the workshop’s program chairs, collectively read through hundreds of abstr…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com