Agentic Coding Is Not a Chatbot.
It’s a Pipeline With a Spec Boundary.

Autonomy is a dial, not a goal. The spec is the control artifact the loop converges on. Done is when code meets spec and the human says yes.

The points

  1. The spec must be watertight for the output to be correct. If it isn’t, the agent invents requirements, substitutes its own judgement, and produces something generally correct but wrong to your intent. A chatbot that writes plausible code is not a development team. That position holds.
  2. Agentic coding is a controlled pipeline, not a conversation. Roles, gates, findings sent back on failure, a completion gate that re-verifies independently, and validation outside the container.
  3. The human sits at the spec boundary, not in the inner loop. You do not review agent turns. You observe the validated outcome, amend the spec, harden the instructions, and rerun.
  4. Done is a conjunction. Code meets spec and the human test passes. Not one or the other.
  5. The spec is the control artifact the loop converges on. It is not the input. It matures through cycles.
  6. APIs are an easy agentic task to spec. UI is not. UI still needs human interaction and close guidance. That is where the boundary belongs.
  7. Autonomy is a dial, not a goal. It peaks at medium spec maturity. Execution usefulness stays high all the way to the end.
  8. Each stage of the pipeline has its own model selection. This is a framework, not a recommendation. The user configures it to their constraints: privacy, sovereignty, regulation, capability, cost envelope.

01The pipeline

Spec vN + Instructions vM control artifacts
Agentic pipeline
Audit
Architect
Review
Build
Adversarial review
Completion gate
Software candidate
Automated tests / gates outside validation
PASS release
FAIL / gaps amend spec vN+1
harden instructions vM+1
↻ next run, back to the top

The run outputs software. The end-of-cycle step amends the spec and hardens the instructions based on automated tests, identified gaps, and user input. Then it starts again.

The human is at the amendment boundary. You are not patching code between agent turns. The spec stays authoritative. The code never becomes the source of truth.


02Done

The terminal condition
Code meets specHuman testMeaningAction
YesYesDoneRelease
YesNoSpec was wrong or incompleteAmend spec
No-Spec not metHarden instructions or fix spec
NoYesSpec too rigid, or test wrongAmend spec

The row that matters is Yes / No. Automated tests passed. The spec said the right mechanical thing. The human test fails because intent was missing.

That is not a code failure. It is a spec failure. It is why the human sits at the boundary, and why you amend the spec rather than patch the code.


03Not all code is equally agentic

Weight of the human test by layer
LayerSpec maturity achievableAutomated verificationHuman testAgentic value
API / contractsHighStrongLightHigh
Business logicHigh if rules clearStrongLight-mediumHigh
Integration / dataMedium-highMedium-strongMediumMedium-high
UI behaviourMediumPartialHeavyMedium
Visual / UXLow-mediumWeakEssentialLow without skills

APIs are easy. The spec is the contract: endpoints, methods, status codes, schemas, error shapes, auth rules, pagination, rate limits. OpenAPI, types, mocks, contract tests. The pipeline verifies “code meets spec” almost completely. The human test is light, because the contract is the intent.

UI is harder. “Looks right,” “feels right,” “flows right,” “matches the brand,” “isn’t confusing” are not fully formalisable. You can test rendering, accessibility, DOM structure, snapshots. You cannot easily test intent. The human test is the real acceptance gate.

UI still needs human interaction and close guidance. That is not a defeat. It is correct division of labour.


04Where autonomy sits

Autonomy is a dial, not a goal.

Agentic coding usefulness and authority versus specification maturity Three conceptual curves. Autonomy you should grant rises from low at zero percent specification maturity, peaks around fifty-five percent, then falls towards one hundred percent. Execution usefulness rises with specification maturity and plateaus high. Risk of silent reinterpretation is high at low specification maturity, dips in the middle, and rises again at very high specification maturity if too much autonomy is granted. EXPLORE ARCHITECT + IMPLEMENT EXECUTE high med low 0% 20% 40% 60% 80% 100% Specification maturity: evidence only → fully specified peak autonomy
Autonomy you should grant Execution usefulness Risk of silent reinterpretation
Illustrative, not empirical. Two curves, not one. Autonomy peaks in the middle. Execution usefulness stays high to the end. Risk falls in the middle and rises at both extremes: at low maturity because the agent is inventing requirements, at high maturity because too much autonomy lets it substitute what makes sense for what you specified.

A watertight spec does not make the agent useless. It makes it reliable.


05Per-stage model selection

The framework is model-agnostic. Each stage has its own selection.

Stage requirements and candidate models
StageWhat it needsSuits
AuditReading, extracting, summarisingCheap / local: high volume, low reasoning
ArchitectJudgement, constraints, trade-offsStrong model: reasoning-heavy
ReviewChecking against specStrong model, ideally different from the builder
BuilderHigh-volume code generationLocal: cheap, fast, bounded by spec
Adversarial reviewFinding faults, different perspectiveDifferent model from the builder
Completion gateIndependent verificationDeterministic tests plus a model that has not seen the build
Architectural property

Using the same model to build and to review gives you correlated blind spots. A different model catches what the builder’s own reasoning would rationalise away. That is an architectural property, not a cost trick.

Swapping a model does not change the framework. Changing a cost envelope does not change the framework. The framework is the architecture. The configuration belongs to whoever owns the constraint.


06The framework sits on top of

Auditability

Roles and gates, not the model. Every stage has a defined input, a defined output, and a gate that can reject. Every step leaves evidence. Every decision is attributed to a stage, not to “the AI.”

Authority

The spec boundary. Spec vN and Instructions vM are the control artifacts. Amendments are explicit, versioned, attributable. The spec drives the process, not the code, and not the model.

Configurability

None of these have a correct answer in the abstract. They have a correct answer for a given organisation, workload, and constraint set. The framework’s job is to make the choice explicit and the consequences visible. It does not make the choice.

Not a recommendation

I am not recommending local over frontier, or frontier over local. I am not making absolute cost claims. Local brings privacy, data sovereignty, model version control, and cost predictability where those matter. Frontier brings capability where the judgement is hard enough to justify it. Both are valid. That is a per-stage, per-organisation decision.


07Support: what the comparison looks like

I tested this on a single RTX 5090 with Qwen3.8-27B at Q6, using temperature and top-p profiles to move between precision and exploration across stages.

Indicative benchmark positions, September 2026
Benchmark Qwen3.8-27B (local) GPT-6 Astra Claude Opus 5 Claude Fable 5.1
SWE-bench Pro61.7% (vendor)-79.2%81.2%
Terminal-Bench 4.073.0 (TB 2.1)57.9%52.6%55.8%
DeepSWE v1.142.274.1%73.7%67.4%
OSWorld 2.084.3 (Verified)72.6%70.6%-
Cost / M output tokens~$0 marginal$50$25$50

The frontier is ahead on absolute scores. The gap is real. But three things matter more than the gap.

Benchmarks are unreliable at the top

Three different scores claim to be the best SWE-bench Pro result: 61.5%, 80.0%, 51.5%, all real, all measuring different things through different scaffolding and data splits. A benchmark saying an agent solved a GitHub issue does not tell you whether it built your UI correctly. Your acceptance criterion belongs to you, not to the benchmark.

The harness is part of the system

A score inside someone else’s orchestration is not a prediction of what a model does inside yours. The pipeline is your harness. It constrains the model and defines the gates.

The metric is cost per approved release, not per token

A model that needs one run at frontier prices is not cheaper than a model that needs two runs at zero marginal cost, if time-to-approved-release is acceptable.

The honest position is not “local matches the frontier.” It is: the pipeline changes what the model needs to be. For bounded, spec-driven work where the spec can be made watertight, the framework is viable on whatever model the constraints select. For open-ended greenfield work where the spec is vague, no model rescues you. You are doing requirements inference, and that is a spec problem, not a model problem.


08Where this lands

The spec must be watertight for autonomous delivery. That has not changed.

What is worth adding: the spec is not the input. It is the control artifact the loop converges on. The pipeline produces software and reveals gaps. Automated tests and the human test feed back into the spec and the instructions. The spec matures through the loop, not before it.

APIs are the easy case. UI is the hard case. That is fine.

Agentic coding is not a chatbot. It is a controlled development pipeline with a spec-refinement feedback loop. The software is the output. The spec is the control artifact. The human owns the boundary. Autonomy is a dial. Execution is the constant. The model is a per-stage configuration.

Done is when code meets spec and the human says yes.

Benchmark figures are indicative and drawn from a mix of vendor-reported and independently scaffolded sources. They are not directly comparable across rows. Treat them as directional, not definitive. Your pipeline’s external validation is the only benchmark that matters for your acceptance criterion.

Framework notes · September 2026