Introduction Small simple synthetic natural language datasets suitable for end-to-end training of tiny LLMs serve as an important resource for LLM interpretability researchers. Some well know examples include roneneldan/TinyStories , SimpleStories/SimpleStories , and klusai/ds-tf1-en-3m (TinyFabulist). This post solves key problems that degrade the quality of these datasets, while also offering an efficient accessible pipeline that can be run locally on an NVIDIA 5060 Ti (16GB) graphics card. The core problems this post solves, include: True Small Vocabulary. The aforementioned datasets attempt to produce a corpus with a small vocabulary, but arguably fall a bit short of that goal. E.g., TinyStories has 49,187 unique words, SimpleStories has 40,567, and TinyFabulist has 41,502. Guaranteed minimal word frequencies. In the aforementioned datasets, 15 to 24 percent of the unique words occur less than 2 times, while between around 43 to 54 percent occur less than 8 times. This means that most of the unique words are likely not learnable, and mostly contribute to noise and vocabulary bloat. Error free text. The aforementioned datasets, include lots of errors, such as misspelled and mangled words. Reliable Name Disambiguation and Stratification. The aforementioned datasets, have various name management issues, ranging from collision with existing words (e.g., May vs may), name bloat, and no control over gender balance, or bias (e.g., certain names may be more likely to co-occur wi…

Full article content could not be extracted automatically. Read the original below.