Built a 350M nanoGPT-style trainer for erotica corpus, ready for Colab/Kaggle
Takeaway: Next time I'd specify upfront that I want the split and training pipeline to preserve the conversation metadata (system/human/gpt) rather than letting the agent default to the flattened corpus.txt — that correction is what forced the whole pipeline to be rebuilt properly.