Training Language Models via Neural Cellular Automata
Researchers developed a new approach to train language models using synthetic data from neural cellular automata (NCA) instead of natural language text, addressing concerns about running out of high-quality training data by 2028. Their method pre-trains models on 164 million tokens of NCA sequences—abstract grid-based patterns that force models to infer hidden rules from context—before standard language training,...