
A 7-day camp to reproduce GPT-2 from OpenWebText.
LLMs are becoming a new kind of apparatus: they quietly reshape how we write, learn, create, and even imagine intelligence. If we only use them, we become functionaries of their defaults. So we do something different: we open the black box and learn it as a craft.
Over 7 days, in a mixed audience of artists, hackers, researchers, and skeptics, we will reproduce a GPT-2 training pipeline on the OpenWebText dataset—from raw text to tokens, from tokens to gradients, from gradients to generation. We will walk the whole loop: dataset preparation, tokenization, Transformer forward pass, backpropagation, optimization, loss curves, sampling, and evaluation. No mysticism. No priesthood. Just careful steps and shared curiosity.
Why here, why now: this is not about treating LLMs as a productivity tool. It is about agency. When you can train even a small GPT-2 yourself, the spell breaks: you see what the model learns, what it cannot learn, how data shapes behavior, where failure modes come from, and why scale changes everything. That literacy won’t remove the social shock of automation—but it gives you a way to respond with clarity instead of helplessness.
Bring: a laptop, curiosity, and a willingness to experiment. If you have GPU access, great. If you don’t, also great: we will share runs, compare results, and make the system legible together. We are not only users of models—we are builders who can play with GPUs.
Class schedule: See the following event page
Repo: https://github.com/catslovefish1/catgpt
Email: afcionadofly@gmail.com, ashincherry23@gmail.com
Telegram: @takshire, @catslovefish