Autonomy is a dial, not a goal. The spec is the control artifact the loop converges on. Done is when code meets spec and the human says yes.
The run outputs software. The end-of-cycle step amends the spec and hardens the instructions based on automated tests, identified gaps, and user input. Then it starts again.
The human is at the amendment boundary. You are not patching code between agent turns. The spec stays authoritative. The code never becomes the source of truth.
| Code meets spec | Human test | Meaning | Action |
|---|---|---|---|
| Yes | Yes | Done | Release |
| Yes | No | Spec was wrong or incomplete | Amend spec |
| No | - | Spec not met | Harden instructions or fix spec |
| No | Yes | Spec too rigid, or test wrong | Amend spec |
The row that matters is Yes / No. Automated tests passed. The spec said the right mechanical thing. The human test fails because intent was missing.
That is not a code failure. It is a spec failure. It is why the human sits at the boundary, and why you amend the spec rather than patch the code.
| Layer | Spec maturity achievable | Automated verification | Human test | Agentic value |
|---|---|---|---|---|
| API / contracts | High | Strong | Light | High |
| Business logic | High if rules clear | Strong | Light-medium | High |
| Integration / data | Medium-high | Medium-strong | Medium | Medium-high |
| UI behaviour | Medium | Partial | Heavy | Medium |
| Visual / UX | Low-medium | Weak | Essential | Low without skills |
APIs are easy. The spec is the contract: endpoints, methods, status codes, schemas, error shapes, auth rules, pagination, rate limits. OpenAPI, types, mocks, contract tests. The pipeline verifies “code meets spec” almost completely. The human test is light, because the contract is the intent.
UI is harder. “Looks right,” “feels right,” “flows right,” “matches the brand,” “isn’t confusing” are not fully formalisable. You can test rendering, accessibility, DOM structure, snapshots. You cannot easily test intent. The human test is the real acceptance gate.
UI still needs human interaction and close guidance. That is not a defeat. It is correct division of labour.
Autonomy is a dial, not a goal.
A watertight spec does not make the agent useless. It makes it reliable.
The framework is model-agnostic. Each stage has its own selection.
| Stage | What it needs | Suits |
|---|---|---|
| Audit | Reading, extracting, summarising | Cheap / local: high volume, low reasoning |
| Architect | Judgement, constraints, trade-offs | Strong model: reasoning-heavy |
| Review | Checking against spec | Strong model, ideally different from the builder |
| Builder | High-volume code generation | Local: cheap, fast, bounded by spec |
| Adversarial review | Finding faults, different perspective | Different model from the builder |
| Completion gate | Independent verification | Deterministic tests plus a model that has not seen the build |
Using the same model to build and to review gives you correlated blind spots. A different model catches what the builder’s own reasoning would rationalise away. That is an architectural property, not a cost trick.
Swapping a model does not change the framework. Changing a cost envelope does not change the framework. The framework is the architecture. The configuration belongs to whoever owns the constraint.
Roles and gates, not the model. Every stage has a defined input, a defined output, and a gate that can reject. Every step leaves evidence. Every decision is attributed to a stage, not to “the AI.”
The spec boundary. Spec vN and Instructions vM are the control artifacts. Amendments are explicit, versioned, attributable. The spec drives the process, not the code, and not the model.
None of these have a correct answer in the abstract. They have a correct answer for a given organisation, workload, and constraint set. The framework’s job is to make the choice explicit and the consequences visible. It does not make the choice.
I am not recommending local over frontier, or frontier over local. I am not making absolute cost claims. Local brings privacy, data sovereignty, model version control, and cost predictability where those matter. Frontier brings capability where the judgement is hard enough to justify it. Both are valid. That is a per-stage, per-organisation decision.
I tested this on a single RTX 5090 with Qwen3.8-27B at Q6, using temperature and top-p profiles to move between precision and exploration across stages.
| Benchmark | Qwen3.8-27B (local) | GPT-6 Astra | Claude Opus 5 | Claude Fable 5.1 |
|---|---|---|---|---|
| SWE-bench Pro | 61.7% (vendor) | - | 79.2% | 81.2% |
| Terminal-Bench 4.0 | 73.0 (TB 2.1) | 57.9% | 52.6% | 55.8% |
| DeepSWE v1.1 | 42.2 | 74.1% | 73.7% | 67.4% |
| OSWorld 2.0 | 84.3 (Verified) | 72.6% | 70.6% | - |
| Cost / M output tokens | ~$0 marginal | $50 | $25 | $50 |
The frontier is ahead on absolute scores. The gap is real. But three things matter more than the gap.
Three different scores claim to be the best SWE-bench Pro result: 61.5%, 80.0%, 51.5%, all real, all measuring different things through different scaffolding and data splits. A benchmark saying an agent solved a GitHub issue does not tell you whether it built your UI correctly. Your acceptance criterion belongs to you, not to the benchmark.
A score inside someone else’s orchestration is not a prediction of what a model does inside yours. The pipeline is your harness. It constrains the model and defines the gates.
A model that needs one run at frontier prices is not cheaper than a model that needs two runs at zero marginal cost, if time-to-approved-release is acceptable.
The honest position is not “local matches the frontier.” It is: the pipeline changes what the model needs to be. For bounded, spec-driven work where the spec can be made watertight, the framework is viable on whatever model the constraints select. For open-ended greenfield work where the spec is vague, no model rescues you. You are doing requirements inference, and that is a spec problem, not a model problem.
The spec must be watertight for autonomous delivery. That has not changed.
What is worth adding: the spec is not the input. It is the control artifact the loop converges on. The pipeline produces software and reveals gaps. Automated tests and the human test feed back into the spec and the instructions. The spec matures through the loop, not before it.
APIs are the easy case. UI is the hard case. That is fine.
Agentic coding is not a chatbot. It is a controlled development pipeline with a spec-refinement feedback loop. The software is the output. The spec is the control artifact. The human owns the boundary. Autonomy is a dial. Execution is the constant. The model is a per-stage configuration.
Done is when code meets spec and the human says yes.
Benchmark figures are indicative and drawn from a mix of vendor-reported and independently scaffolded sources. They are not directly comparable across rows. Treat them as directional, not definitive. Your pipeline’s external validation is the only benchmark that matters for your acceptance criterion.
Framework notes · September 2026