Ageiro ARK, the Project Intelligence Platform for AECO
Back to Blog
Thought Leadership9 min read

Betting five years on an AI vendor? Look at how they build

Phil Johnson

Phil Johnson

Head of Engineering

October 1, 2026

Ageiro’s Head of Engineering on how we hand the repetitive loops of software delivery to AI agents, while humans stay accountable and in command of every plan and every merge. And why, in an industry built on provenance and sign-off, that makes us easier to trust rather than harder.

Every AECO organization evaluating an AI vendor right now is asking a bigger question than the one written on the RFP. The stated question is whether the software solves a problem they have this quarter. The real question is whether they trust this company enough to build the next five years of how they work around it.

A feature list will not tell you. How a company behaves when nobody is watching will: what discipline it holds itself to, what evidence it can produce, and whether the standards it asks its customers to rely on are the standards it applies to its own work.

This post is about how we build ARK. Which loops in our software development lifecycle we have deliberately handed to AI agents, where we have deliberately kept a human in charge, and why building this way produces a stronger record than the way most software gets built.

The discipline AECO already runs on

Nothing in the process below would be unfamiliar to anyone who has worked on a project. Provenance. Sign-off. A clear record of who decided what, when, and on what basis. A stage gate, a technical submittal, a valuation, an RFI closed out with a name against it.

AECO has run on documented accountability for a very long time, because the cost of getting it wrong is measured in more than a rollback.

Software has mostly not worked that way. The evidence of how a change came to exist tends to be reconstructed after the fact, usually under audit pressure, usually by someone reading commit history and trying to remember what happened in a meeting eight months ago.

At Ageiro the standard is the same. What changed is that the evidence is produced automatically, as a by-product of the work, rather than assembled afterward.

That is also why we can say something specific about ARK rather than something general about AI. ARK is the Project Intelligence Platform for AECO, and its value rests on a single claim: that an answer can be traced back to the source it came from. A company that cannot trace its own changes back to the decision that authorized them has no business making that claim about anyone else’s project data.

The toil got automated. The craft got enforced.

I have spent my career on clean code, clean architecture, SOLID, and the discipline of test-driven development. I still believe those are what produce software that performs, scales and survives contact with the next engineer.

But before generative AI, building software was more toil than craft. Keeping a branch in sync. Resolving the same merge conflicts again. A small interface change taking most of a day to build and test properly, then sitting in review for another two. Tickets drifting from reality. Documentation written under duress, usually after someone went on leave and something broke.

None of that reflects badly on engineers. It reflects how much of software delivery is process overhead rather than the part that creates value.

The obvious reading is that agents take the toil off your plate. The second effect matters more. When your architectural standards, coding practices and review criteria are written into the codebase as agent files that the AI reads before it plans anything, those standards are applied to every single change. Including at six in the evening on the day of a release, which is exactly when a team under deadline pressure starts making exceptions.

The toil gets automated. The craft does not get automated away. It gets enforced.

How a change actually gets made

We run our process as a set of Agent Developer Workflows. A single command run in Claude Code chains together the steps that take a ticket from To Do on the Jira board to a pull request ready for a human to merge. Every step in between is deliberate.

Two human decisions bracket everything the agents do. The Agent Developer Workflow runs from Jira ticket to merge in four stages: plan agreed, where an engineer approves what gets built; agent builds, where it implements, tests and reviews its own work; agent reviews, where a separate agent reviews the code, looping back to refine and re-review; and merge approved, where an engineer approves what ships. Every branch, commit and pull request carries its Jira ticket number, and an AI-flagged change cannot open a pull request until the evaluation suite has run clean.
Agent Developer Workflow · 4 stages, 2 human gatesNothing crosses either gate without a named engineer approving it.
  1. A human agrees the plan. An agent proposes an implementation approach against the ticket. An engineer works through it with the agent and approves it. Nothing gets built until a named person has agreed what is being built. This is the first checkpoint.
  2. Implementation and first-pass review run unattended. Once the plan is agreed, the agent implements the change, writes and runs the tests, and reviews its own work against our standards. Nobody needs to sit and watch.
  3. An independent agent reviews the pull request. A separate agent in GitHub, with no involvement in writing the code, reviews the pull request against our code review standards. The implementing agent responds to the feedback, refines, and the two converge over several rounds.
  4. A human approves the merge. An experienced engineer reviews the result and approves it. Nothing ships without that. This is the second checkpoint.

The two human checkpoints are decisions rather than rubber stamps. What gets built, and what ships. A named person owns each one, and the record of who they were outlives the change. Today that workflow runs on Claude Code, Jira and GitHub. The process is designed so the tools can change without the controls changing.

The evidence is a by-product, not paperwork

Every branch, commit and pull request carries its Jira ticket number. That gives each ticket a visible golden thread: from the moment the work was requested, through the plan that was agreed and the person who agreed it, to the code that was written and the engineer who approved it shipping.

Separating the person who authorizes work from the person who approves its release is the kind of segregation-of-duties control an ISO 27001 auditor expects to see. In our process it is structural rather than procedural. Nobody has to remember to do it, and there is no path through the workflow that skips it.

The second piece matters more for an AI product. Changes that touch AI functionality, a model change or a change to the retrieval pipeline behind ARK, are flagged for stricter handling. Classification works two ways, and both are recorded on the ticket. An engineer adds an ai-change label whenever they judge a change touches a model or the retrieval pipeline. The agent that picks up the ticket makes the same judgment independently, and adds the label if it is missing. Neither is the sole gatekeeper. Two independent judgments have to miss it for a change to reach review unflagged.

Once a change is flagged, our ship command will not open the pull request until the evaluation suite has run clean, with the results posted back to the ticket. That is auditable evidence, generated while the work happens, that the change does not regress the benchmarks we hold ourselves to. It is the kind of record an ISO 42001 audit asks for, and it exists weeks before anyone asks for it.

Grounded in truth is the promise we make about ARK. This is what it looks like when we apply it to ourselves.

What has been harder than we expected

The automated review loop works better than I expected. It behaves like a clean pair of eyes every time. No fatigue, no deference to whoever wrote the code, no decision that it is probably fine. By the time a human picks up a pull request for the final look, it has usually had a more thorough review than a routine change would ever otherwise receive.

The difficulty is what to do with everything it finds. A reviewer that catches more than a person would can pull a single change in too many directions at once. The blast radius grows and the change stops being reviewable, or mergeable, as one coherent unit. Deciding which suggestions belong in this change, which belong in a new ticket, and which are not worth tracking at all has been more of a judgment call than I expected going in.

The second thing is convergence. Two agents can oscillate, fixing one thing while breaking another. We review for that after three rounds. It is not a hard stop, it is the point at which a human checks whether the two are converging. Most changes settle within about five rounds. When one does not, that is usually a signal the ticket was scoped wrong, and we break it up rather than let the loop keep running.

Here’s an example of something the automated review caught that a human reviewer would likely miss. A check was being added within ARK to prove ARK never answers a factual question from memory, rather than content it retrieved. It works by matching a turn’s tool-call activity in the logs to that turn. The automated reviewer found the matching window was wide enough that a turn with zero tool calls of its own could inherit the previous turn’s tool names and pass anyway. It is not something you would find reading the code. You would only find it by lining up two turns’ log timestamps side by side, which is not something a person reviewing a diff has reason to do. Caught in the first round of automated review, before a human ever looked at the change.

Why this is the more trustworthy way to build, not the riskier one

Risk in software delivery is not about who typed the code. It is about whether standards were applied consistently, whether an independent review actually happened rather than nominally happened, whether a named person took responsibility, and whether any of that can be evidenced afterward.

On every one of those, this process is stronger than the one it replaced. The standards are applied to every change instead of most changes. The independent review is independent, and it is never skipped because the sprint is ending. The accountability sits with named people at two defined points. And the evidence exists by default rather than on request.

Every step an agent takes here runs inside a boundary a person set and ends at a gate a person holds. Autonomy without that is a bet.

That is what autonomy you can trust means here. It is not a claim we would make about our product if we were not willing to make it about our own engineering first.

What comes next

The models, the agents and the tooling around them keep improving. Nobody knows how much better they will be in a year. What I am confident about is that the process we have built absorbs those improvements rather than needing re-architecting each time one arrives. A sharper planning step, a stronger reviewing agent, a better evaluation suite all drop into the same workflow behind the same two human checkpoints.

That compounding is what our customers and partners experience. A steady improvement in the pace and quality of what we ship, without a corresponding loosening of the controls around it.

In a future post I will go into how our evaluation suite works, including how we use AI as a judge to benchmark answer quality for an AECO-specific retrieval pipeline. It is the mechanism behind a claim we want to be held to: AI functionality in ARK is measured against a benchmark before it ships, not assumed to be working.

Running a vendor evaluation? We are happy to walk an evaluation team through this workflow in detail, including the gates, the labelling and the evidence each change leaves behind. To see what the same standard of traceability looks like applied to your project data, talk to our team.


Phil Johnson

Phil Johnson

Head of Engineering, Ageiro

Phil leads engineering at Ageiro. He has spent his career on clean architecture and test-driven development, and now builds the agent developer workflows that keep Ageiro’s own delivery auditable, with a named engineer accountable at every gate.

Get Started

Ready to see ARK on your own project?

Join the AECO firms building the future on project intelligence