Local Factory
A software workshop running on a local GPU: a specification goes in, a diff comes out, the tests decide, without calling a single online service.
A harness that puts locally installed open-weight models to work. A task descends a ladder of models: the smallest tries first, only failures climb one rung. Each attempt is checked against a specification and a test suite before an automatic reviewer accepts or rejects it.
Around the harness, a few smaller experiments served as test benches. An agent that rewrites its
own code by following metrics computed on the syntax tree (self-improving-AI); an agent
orchestrator with a project memory, left as a sketch (GPT_code); a set of coding and logic
tasks for comparing local models (local-LLM); team roles, product manager, developer and
architect, defined as Ollama models (quiz-app); and a small Python project used as the target
of the first attempts (factory-demo).
What matters here is not the demonstration but the measurement. The project keeps a register of optimisations where ideas are marked falsified when the numbers say so: two of them are, and they stay there. Another time, the harness reported five steps out of five successful in front of an empty diff, which revealed three distinct flaws, all leaning the same way, the agent’s.
The same mistake happened in the other direction, on the human side. A thirty-five-billion parameter model seemed to have hit a reasoning ceiling; the real cause was a misconfigured context length. Since then the rule sits at the top of the project instructions: never conclude on a failure without an autopsy. Later, a tiered benchmark toppled a belief held for weeks, that the big model was only useful as an escalation. Two bugs in the harness were skewing every recorded verdict. Once fixed, the tier supposedly reserved for escalation became the best, and the previous belief went into the register marked falsified, not deleted.
The guardrails are executable rather than written on a wall, and that rule is a scar: an agent
ran git stash and wiped work in progress. The answer was not to write the instruction better,
since an instruction can be worked around, but to make it unenforceable: git stash,
reset --hard, push --force and any deletion under .git/ are refused at the tool level, not
the prompt level. On top of that, the two rules that protect the repository rather than the work
in progress: no delivery on the main branch, no push while tests are red.
The conclusion, after all this, is not comfortable: the limit is not the harness, it is in the models a personal machine can run.
Private project, no public link
