Software 2.0 writes Software 1.0
As coding models become increasingly good, I think we will see some surprising examples of 'codification'. Instead of training a custom deep learning model to predict something, we can use 'just code'.
A clear example is shown below, where paintings in the style of famous painters are generated purely by code. There is no generative image model here - it's simply a pixel space created with a lot of rules about what pixels should look like.
Every image here is a Python program generated pixel by pixel. There is no image model, and no off-the-shelf art software. Instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. The agents don't use any pictures as reference, instead working only from what they know about each painter.
This might look useless at first. But here we see a more practical use case, where we can prompt a model to codify the rules that allow one to complete a robot task :
Coding models have become so good at understanding the world that they can often describe its underlying rules, given enough time, in code. Humans would not reasonably do this by hand. The code would be far too verbose and hard to maintain. The limited number of rules we could come up with would result in a brittle system. But with competent coding models, and with code maintenance handed off to coding models as well, what would have been brittle, unmaintainable code when written by humans might, in some cases, now be a valid solution to a problem.
In the robot example, the result is not learned model weights (although this can be part of the solution), but 10-100K lines of code. That can have many benefits:
- Updates are extremely easy; simply ask the coding model to change something, and it can update the code in minutes. As a result, updates become very cheap. No GPU is required for the end user, just a Codex/Claude Code subscription. No large training dataset is required. Continual learning becomes incredibly simple and fast!
- Updates can be isolated. Catastrophic forgetting might happen less often, since models can write good test cases that reduce the risk of regressions.
- Code can hardcode important rules we don't want to approximate but want set in stone.
- Out-of-distribution generalisation can come for free for any dimension that the code treats as a variable. A learned grasping policy might fail when an object is placed in a position it never saw in training; code that computes the grasp from a detected position doesn't care where the object is.
- The code is, to a certain extent, more auditable than model weights - the code can be reviewed by other language models.
- Running the code is cheap and usually deterministic; no GPU is required at inference time.
Funnily enough, this can be seen as the inverse of 'Software 2.0' as popularised by Karpathy. Yes, gradient descent can write better code than you. Now, the product of gradient descent, a coding model, writes actual code again: Software 2.0 writes Software 1.0 .
Is this an anti-bitter lesson? Not necessarily. A general model is still doing the work, and we add very little of our own knowledge. Only the output is different: code, not model weights.
Pim de Witte from General Intuition makes a similar point:
Anthropic bet really really hard on coding space. They did an excellent job at building really good frontier coding models. And then they bet that a sufficiently good general model would encourage people to bring problems back into coding space. So you start solving more problems because this general model in coding space is so incredibly capable.
Of course, codification doesn't apply to many parts of the world where things are very unstructured, such as natural language or raw perception. In practice, code will call learned components for those situations. However, this does not mean codification only applies when humans know the exact rules that underlie structure in advance. Few know the exact rules of a painter's style, yet a model can find rules that approximate it closely. I expect we will find many more, often surprising, use cases where codification is a useful way to solve problems.
Footnotes
This builds on Code as Policies (2022). The difference is that we can now describe more complex, long-horizon tasks.
Karpathy later used Software 3.0 for programming with natural language prompts. All three generations are at play here, but the surprising part is that there are examples where weights hand the work back to code, not that we prompt the weights.