AI-assisted coding
A daily practice since May 2024, followed along the whole curve: autocompletion, then agents, then custom layers on top of them.
- Ongoing
- Claude Code
This is not a project, it is the way the others are made. Since May 2024, nearly everything the neighbouring pages describe has gone through it, and the progression was followed from the start: first autocompletion, then agents able to edit a whole repository, and today two custom-built layers because the tool alone was no longer enough.
The starting point is older and much more modest. In 2023, a repository named ChessGPT held a chess engine entirely dictated by a model, at a time when that kind of exercise was more curiosity than method; another, left as a sketch, was already trying to give a model a project memory and an autonomous development loop. Both are attempts, not results, and they mark the start of the curve.
The first current layer, claude-remote, orchestrates several agents at once: parallel sessions on different projects, typed sub-agents, coder, architect, reviewer, each on the model that suits it, isolated workspaces so they do not overwrite each other, and the whole thing drivable from a phone. The second, local-factory, runs quantised open-weight models on a local graphics card and only sends the remote model what deserves it. Alongside, a second reviewer taken from another model family, because a model reviews its own work poorly.
Review does not follow the same regime on both sides. Code carried for years is reviewed by hand: its history says whether a change is right. What comes out of an agent goes through layered automated review, with a different provider for review and for writing. Having a model review itself amounts to asking it to find the error it just failed to see.
Two lessons, and neither is a speed gain. The first: a small model only produces usable code inside a loop that frames it, a specification written before, a diff read after, a regression suite, a review. Outside that loop, it produces text that looks like code. All the work is in the loop, not in the model.
The second is less pleasant. The custom verification harness contained bugs that made failed tasks look successful. All of them, without exception, leaned the same way: towards success. A verification system whose flaws are biased towards passing is worse than no verification at all, because it manufactures confidence. Since then, the harness reads the disk itself rather than taking the agent’s word for it.
This page gives no percentage. No protocol clean enough for a number to mean anything. The qualitative finding: the machine is better on volume and consistency, worse on what has never been written before, and it fails while giving the impression of having succeeded. That last property is what decides how it gets tooled.
