Skip to content
Martin Poiroux
Language : Français
Projects · CodeCode · AI tooling · 2026

Local Factory

A software workshop running on a local GPU: a specification goes in, a diff comes out, the tests decide, without calling a single online service.

A harness that puts locally installed open-weight models to work. A task descends a ladder of models: the smallest tries first, only failures climb one rung. Each attempt is checked against a specification and a test suite before an automatic reviewer accepts or rejects it.

Animated diagram: a task only ships after passing the tests and the review. A red test or a rejected review sends it back to the diff.

Around the harness, a few smaller experiments served as test benches. An agent that rewrites its own code by following metrics computed on the syntax tree (self-improving-AI); an agent orchestrator with a project memory, left as a sketch (GPT_code); a set of coding and logic tasks for comparing local models (local-LLM); team roles, product manager, developer and architect, defined as Ollama models (quiz-app); and a small Python project used as the target of the first attempts (factory-demo).

The model catalogue: local on the GPU, by subscription, by key or free, each with its context size and price.

What matters here is not the demonstration but the measurement. The project keeps a register of optimisations where ideas are marked falsified when the numbers say so: two of them are, and they stay there. Another time, the harness reported five steps out of five successful in front of an empty diff, which revealed three distinct flaws, all leaning the same way, the agent’s.

The same mistake happened in the other direction, on the human side. A thirty-five-billion parameter model seemed to have hit a reasoning ceiling; the real cause was a misconfigured context length. Since then the rule sits at the top of the project instructions: never conclude on a failure without an autopsy. Later, a tiered benchmark toppled a belief held for weeks, that the big model was only useful as an escalation. Two bugs in the harness were skewing every recorded verdict. Once fixed, the tier supposedly reserved for escalation became the best, and the previous belief went into the register marked falsified, not deleted.

The dashboard: jobs run, successes, success rate per model and latest failures.

The guardrails are executable rather than written on a wall, and that rule is a scar: an agent ran git stash and wiped work in progress. The answer was not to write the instruction better, since an instruction can be worked around, but to make it unenforceable: git stash, reset --hard, push --force and any deletion under .git/ are refused at the tool level, not the prompt level. On top of that, the two rules that protect the repository rather than the work in progress: no delivery on the main branch, no push while tests are red.

The conclusion, after all this, is not comfortable: the limit is not the harness, it is in the models a personal machine can run.

Private project, no public link