LessWrong AI
2026-09-22 13:33 UTC
By invertedpassion
USR-0152-20260922-community-fo-48ef8789
Modern LLMs have tiny GPTs hidden inside them
Experiments into predicting GPT2 completions via Qwen models This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk , where we're investigating meta-cognition in LLMs as one of the projects. ---- Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls. Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams. In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs. This is an intriguing hypothesis. So I decided to do a quick-and-dirty exploratory study to investigate. The Experiment I selected two models for the experiment: GPT2-medium and Qwen3 base model (4bn variant). What I did was the following: Input: Take 12 headlines from the Internet for Sept 18 2026 This is to ensure models don’t just output memorized text as the starting pro…
Experiments into predicting GPT2 completions via Qwen models This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk , where we're investigating meta-cognition in LLMs as one of the projects. ---- Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls. Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams. In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs. This is an intriguing hypothesis. So I decided to do a quick-and-dirty exploratory study to investigate. The Experiment I selected two models for the experiment: GPT2-medium and Qwen3 base model (4bn variant). What I did was the following: Input: Take 12 headlines from the Internet for Sept 18 2026 This is to ensure models don’t just output memorized text as the starting pro…
Full article content could not be extracted automatically. Read the original below.