ETABLIX · INTEGRATED SITE SERVICES · PART OF GROUPE NSEYA

AI agents in construction: what they do, what they must never do, and who is accountable

We run thirteen AI agents inside a live construction site-services business. What they are refused matters more than what they do — and what a client pays for is neither.

Most writing about AI in construction is written by people who have not shipped anything. This is not that. We run thirteen agents inside a live site-services business, on documents a contractor prices against and a client relies on. The useful thing to report is not what they can do. It is where they are stopped, who stops them, and why none of it changes what the work is worth.

The interesting question in 2026 is no longer can a model write a specification. It can. The question is what has to be true around the model before the specification is safe to issue with your name on it, and that turns out to be an engineering problem rather than a prompting one.

What the fee is actually for

Worth answering before anything else, because it is the first thing a client thinks once they know we use these at all. If a machine drafted it, what exactly is being charged for?

You are not buying a document. You are buying a judgement, and somebody accountable for it.

The fee buys a named competent person who has read the output and put their name on it, an opinion you can rely on and challenge, a date we are held to, professional indemnity standing behind it, and an organisation that carries the consequence when it is wrong. None of that is affected by how the first draft came to exist, and none of it is something a model can hold.

What the agents change is how much of the work gets the same attention. A person writing a twelve-package requirements set has eleven other things to get out this week. The first four packages get scrutiny and the last eight get pattern-matching. That is not a criticism of anybody; it is what a deadline does to a human being. An agent does not arrive at the twelfth package tired. So the fee does not buy the hours we saved. It buys a document where the twelfth row was treated like the first, and a person who checked that it was.

And the hours saved are real but narrow. They come out of first-draft production and out of reading six documents against each other — never out of the review, which is where the liability sits and which does not compress at all. If we priced this as machine output we would be selling you the one part of it that carries no responsibility. The market rate for that is nothing, correctly.

Which is also why the rest of this piece is mostly about refusals. What a system will not do, and what enforces that, is the part you are actually buying.

An agent is a boundary, not a model

The word "agent" has been stretched until it means nothing. In our system it has one meaning: a defined job, with defined inputs, a defined output, and a defined thing it is not permitted to do. The model is the least interesting part. Two agents can run on the same model and be entirely different products because their boundaries differ.

Take the two that sit next to each other in our procurement chain. One assembles a tender pack from an approved requirements package. The other evaluates the returns when they come back. Same model. The first may not add a requirement that is not in the approved document. The second may not adjust one tenderer's price on an assumption it has not applied to the others. Those two sentences are the products. Everything else is plumbing.

This matters commercially, not just philosophically. When a client asks whether we use AI, the answer that reassures them is never "yes, the latest one". It is "here is the list of things it is refused, and here is the mechanism that refuses it".

The eight jobs they genuinely do

Across our three delivery models the agents earn their place in eight distinct kinds of work. Every one of them is a job where a competent person reading sequentially is at a structural disadvantage.

1. Reading documents against each other

This is the one that surprises people, and it is the most valuable. A client hands over a programme, a workforce forecast, a layout, a logistics plan, a set of planning conditions and a procurement schedule. A person reads them one at a time, because that is how reading works. Almost every finding worth having lives between two of those documents: a shift pattern that a planning condition prohibits, a generator enquiry sized against a cabin schedule that has since been superseded, a bed count that the local market cannot supply.

Nobody misses these because they are careless. They miss them because the contradiction is never on one page. Our entry engagement is built entirely around this: its first pass produces nothing a client sees, and instead builds a working paper of facts and contradictions with a source against every row. The twelve deliverables are then written from that paper rather than from the documents.

2. Working dates backwards from a fixed one

Consents, connections and long-lead items have lead times. An access date has a date. The arithmetic between them is trivial and almost nobody does it, because it requires holding twenty chains in your head at once. An agent does not get bored on the nineteenth one.

The output that earns the fee is not a list of consents. It is a column headed latest responsible start date, with the ones that have already passed at the top.

3. Turning an obligation into something verifiable

"Adequate welfare" is not a requirement. It cannot be priced, delivered against, inspected or enforced. "Twenty-two WCs, twenty-six washbasins, fourteen showers, cleaned twice per shift, verified by weekly inspection against the schedule" is a requirement. Converting the first kind into the second, hundreds of times, across a dozen packages, is exactly the work that gets abandoned at four in the afternoon.

4. Assembling documents from an approved source

Once the requirements are settled, the files that go to market are an assembly job: instructions to tenderers, conditions of tendering, a scope sheet per package, the blank pricing schedule, the return forms, the form of tender, the issue register. Each is a separate document because a tenderer receives them separately — their estimator opens the pricing schedule, their bid manager opens the instructions, their commercial lead opens the form of tender.

5. Normalising returns onto one basis

Three tenderers price one enquiry on three different bases. One includes fuel and two do not. One has priced a superseded revision. One has assumed a thirty-four month term and another thirty. Added up as returned, the cheapest is whoever excluded most. Making the returns comparable before anybody compares them is where a procurement desk earns its money, and it is patient, mechanical, unglamorous work.

6. Finding every qualification and saying what it costs

A qualification nobody read is a variation with a date on it. The job is to find each one, quote the tenderer's own wording where the wording matters, and state what accepting that return as written costs beyond its price.

7. Evidence completeness

What certificates exist, what the contract requires, and the gap between them. Sorted by expiry date. This is a database query wearing a report's clothes, and it is still the reason handovers slip.

8. First drafts of anything structured

Registers, matrices, schedules, checklists, minutes. Not because the draft is good, but because arguing with a draft is faster than facing a blank page, and the argument is where the expertise actually enters the document.

The four they must never hold

This is the part we would want a client to read first. Our agents are hard-coded to refuse four categories, and we would be sceptical of any supplier whose list is shorter.

They never commit money or award anything

No agent awards a contract, places an order, appoints a supplier or approves a payment. Every recommendation goes to a named human with delegated authority who accepts or rejects it. The tender evaluation report we produce says so on its own face, because the person reading it in six months during a dispute needs to see who decided.

They never accept work or close a defect

Acceptance is a legal act with consequences for payment, for defects liability and for insurance. An agent may report that something appears acceptable, and must say that it is only a report.

They never hold a statutory duty

CDM 2015 allocates duties to people and organisations. The HSE's own summary names seven: commercial clients, domestic clients, designers, principal designers, principal contractors, contractors and workers. Software is not on that list and cannot be added to it.

But "software is not a duty holder" is not the interesting part, because nobody thinks it is. Two things about that list matter far more, and most writing on this misses both.

First, only two of those roles are appointed — principal designer and principal contractor, one organisation at a time. The rest attach to what you actually do. Carry out, manage or control construction work and you are a contractor under the regulations whether anybody appointed you or not. Prepare or modify a design, or instruct somebody who does, and you are a designer. Nobody has to hand you a letter.

Second, and this is the real exposure: a document written in the voice of a duty holder is a representation about who is discharging that duty. It does not matter who or what typed it. A fluent model asked to produce a fire strategy will produce something that reads exactly like a competent person's determination, and a reader is entitled to treat it as one. That is the risk, and it is not solved by anybody understanding that a model is not a legal person.

So the agents write in the voice of the party who will actually hold it. They state the requirement and name who determines it. They do not determine it. The phrasings that would breach that are blocked at the point a document is generated, not written down in a style guide, and a test checks what is actually published rather than what we intended to publish.

What none of that tells you is which duties we hold. That is a separate question, it is not answered by anything about the software, and it has its own section below.

They never resolve a life-safety question

Fire strategy, means of escape, compartmentation, alarm category, escape widths, structural loads. Our workforce village agent is required to state the requirement and refer it — every such row carries a marker naming a competent fire engineer and the fire and rescue authority as the people who determine it. It is not permitted to propose a travel distance. People sleep in these buildings.

There is a version of this refusal that is just a disclaimer at the bottom of page forty. Ours is a marker on the row itself, printed in the client's document, next to the thing it applies to. A boundary a reader has to go looking for is decoration.

Which duties we do hold

The four refusals above are about the software. None of them says anything about ETABLIX's own position, and a company that answered the duty-holder question by talking about its tooling would be dodging it. So, plainly.

It depends on the appointment, and the appointment says. That is not evasion — it is what the regulations do. Principal designer and principal contractor are appointed roles, one organisation at a time, so nobody holds either by accident. The others attach to conduct: manage or control construction work and you are a contractor whether anybody appointed you or not.

Which means claiming to hold no statutory duties would be the more dangerous overclaim, and it is the one a supplier is tempted into. A business that manages construction work on a site holds contractor duties by operation of the regulations, whatever its marketing says. We would rather write that down than have a client discover we thought otherwise.

Two positions follow from it, and both are commercial rather than technical.

ETABLIX is not automatically the Principal Contractor. It is an appointment carrying specific health-and-safety duties, and where a client wants us to hold it, that is an explicit, priced and insured decision written into the appointment. It is available. What it is not is an inference somebody draws from how a document was worded, which is why the phrasing is blocked at generation. Under Model 03 the term is "Prime Service Contractor", where prime means prime for the site-services system and nothing else.

Where a role is genuinely arguable, we would rather establish it before the engagement than after an incident. Advisory work writes output specifications, and a designer under CDM 2015 is anyone who prepares or modifies a design in the course of business — so whether specifying brings designer duties is a real question with real insurance consequences. It belongs to a construction solicitor and an underwriter, not to a website, and we would put it to them on a given engagement rather than assert an answer here. A supplier who has never considered the question is the one to worry about.

The difference between a control and a promise

Here is the distinction that separates a system you can rely on from a demonstration.

A promise is an instruction in the prompt. "Trace every requirement to its source." "Never invent a quantity." "Keep the pricing schedule aligned with the scope." These are worth writing and they work most of the time. Most of the time is not a standard you can issue documents against.

A control is a mechanism that checks the finished work and refuses it. The difference is that a control does not care how convincing the output looks.

Our clearest example is the tender pack. A tender pack fails in exactly one way that nobody notices until the returns are in: the scope sheet says one thing and the pricing schedule asks for another.

  • A scope item with no priced line is work the tenderer has been instructed to do and given nowhere to price. It returns after award as a variation, at their rate rather than a tendered one.
  • A priced line with no scope item is a price for something never specified. Every tenderer prices it on a different assumption, and no two assumptions match — which is precisely the condition that makes returns incomparable.

So the scope sheets and the pricing schedule are written in two separate passes, so the second is written against the first rather than alongside it. Every scope item carries a reference. Every priced line names the scope reference it prices, in a column headed exactly Scope ref, and a unit of measurement from a stated vocabulary — because a line priced in "as required" comes back priced in whatever unit each tenderer chose.

Then the two sets of references are compared by a function, on every run, before anybody can approve it. If they do not reconcile, the pack does not issue, and the exceptions are named with the reference and the consequence. When it does reconcile, the count is printed on the issue certificate the client receives, so the check is visible rather than claimed.

That is a control. It is unglamorous and entirely deterministic, and it is worth more than any amount of prompt engineering, because it holds when the model has a bad day, when the pack is unusually large, and when whoever is reviewing it is tired.

If you take one thing from this piece: ask any supplier what their system refuses, and ask what performs the refusing. If the answer is a sentence in a prompt, it is a promise.

What it does not save you

We would rather be believed than impressive, so here is the other side.

It does not save the review. The hours come out of first-draft production and cross-reading, not out of the competent person who reads the output before it is issued. That review is where the liability sits and it does not compress. A document issued without it is cheaper only until it is priced.

It does not know your site. Nothing in a document set tells you the north-east corner holds standing water, that the neighbour's access agreement is verbal, or that the client's project director has already decided. Our only engagement with a site visit — the mobilisation review — exists precisely because of this, and its central distinction is the evidence class of every statement: observed, evidenced, asserted, or unknown. A verdict resting substantially on assertion has to say so.

It does not resolve silence. Where the client's information does not say, the honest output is to name what is missing and what it prevents. The failure mode of a fluent model is a plausible assumption written in the same confident register as a sourced fact. Every load, ratio, rate and duration we publish is labelled as a first-pass planning figure for validation by a competent person, against the table rather than once at the end.

It does not carry the relationship. Contract negotiation, client leadership, incident command, engineering approval. These are not automation-resistant because the technology is immature. They are automation-resistant because someone has to be accountable, and accountability is a property of persons.

Eight questions to put to anyone selling you AI in construction

Use these. They are the ones we would want to be asked, and most of them are uncomfortable.

  1. What does it refuse, and what performs the refusal? A prompt instruction is a promise. A function that inspects the output is a control.
  2. Who approves before anything leaves? Get a name and a role, not "a human in the loop".
  3. Where does a number come from? Ask them to trace one figure in a sample output back to the client document that mandates it. If they cannot, nothing in the document can be relied on.
  4. What happens when the source is silent? The right answer is an open item. The wrong answer is a plausible figure.
  5. What is marked as a proposal rather than a requirement? Anything the supplier invented should be visible as theirs and awaiting your approval.
  6. How are life-safety matters handled? The only acceptable answer is that they are referred to a competent person, on the row, in the document.
  7. Does anything they issue imply a statutory role? Check the wording against the CDM 2015 duty holder definitions yourself.
  8. Can you see a real output, not a demonstration? Ask for a specimen of the actual deliverable. A slide about capability is not a capability.

A worked example of the level of detail this changes

One concrete illustration, because the argument is otherwise abstract.

Welfare sizing is a standard construction task. The common approach is to apply a ratio from memory — one WC per seven workers is the figure most people carry — and size the compound from the peak headcount.

Two things are wrong with that, and they are the kind of thing a system reading against the source catches. First, Schedule 2 of CDM 2015 sets no numeric ratios at all. It requires facilities that are suitable and sufficient, readily accessible, and maintained. The familiar 1:7 comes from the Approved Code of Practice to different regulations — the workplace health, safety and welfare ACOP — and citing CDM for it, as tender documents routinely do, is citing the wrong instrument for your own requirement.

Second, peak headcount is usually the wrong basis. What sizes the facilities is the headcount at shift overlap, and what sizes the dining provision is sittings within the shift pattern rather than the total on site. A compound sized on the annual peak is expensive and still queues at seven in the morning.

Neither point is clever. Both are the sort of thing that is obvious once written down and routinely wrong in issued documents, because the person writing them had eleven other packages to get out. That is the actual shape of the opportunity, and it is what the fee is for: not replacing judgement, but making sure judgement reaches every row rather than the first four — and having somebody put their name to the result.

Where we think this goes

Two predictions, offered as ours rather than as fact.

The differentiator stops being the model and becomes the refusal set. Everyone will have access to comparable capability. What will separate suppliers is the discipline of what they will not let it do, and whether that discipline is mechanical or aspirational.

Traceability becomes a procurement requirement. Public buyers already work under the Procurement Act 2023 with obligations around transparency and fair treatment. It is a short step from there to a client asking, of a document in their tender pack, which source mandated a given requirement — and expecting an answer. We built for that assuming it arrives, because the cost of building for it afterwards is a rewrite.

If you want to see the shape of the output rather than read about it, there is a specimen extract of the real deliverable, and the way an engagement actually runs from first enquiry to issued document. If you would rather just ask us the eight questions above, that is what the contact page is for.

Questions people ask

If AI drafts the document, what am I paying a consultant for?

For a judgement and somebody accountable for it, which is the part no model can hold. The fee buys a named competent person who has read the output and put their name on it, an opinion you can rely on and challenge, a date we are held to, professional indemnity behind it, and an organisation that carries the consequence when it is wrong. What the agents change is how much of the work gets the same attention: the twelfth package is treated like the first rather than pattern-matched at the end of a long week. The hours saved come out of first-draft production and cross-reading, never out of the review, because the review is where the liability sits and it does not compress.

Can an AI agent write a tender pack for a construction project?

It can assemble one from an approved requirements package, and that distinction carries the whole risk. An agent that assembles takes obligations a competent person has already signed off and turns them into the separate files a tenderer receives. An agent that authors invents requirements nobody approved, which then sit in a contract. The control is that the assembler cannot start until a human has approved the document it works from, and that any gap it meets becomes an open item rather than an answer.

What should an AI agent never be allowed to do on a construction project?

Four things, on our reading. It must never award a contract, place an order or commit money. It must never accept work or close a defect. It must never hold a statutory duty, because the HSE names seven CDM 2015 duty holders and all of them are people or organisations. And it must never resolve a life-safety question such as a fire strategy, means of escape or a load; those are referred to a competent person and, where relevant, to the fire and rescue authority. None of that says anything about which duties the supplier itself holds, which is a separate question you should ask separately.

Does ETABLIX hold CDM 2015 duties, and is it the Principal Contractor?

It depends on the appointment, and the appointment says so in writing. Principal designer and principal contractor are appointed roles under CDM 2015, one organisation at a time, so nobody holds either by accident and ETABLIX is not automatically either of them. Where a client wants us to hold Principal Contractor that is an explicit, priced and insured decision written into the appointment, and it is available. The other duty-holder roles attach to conduct rather than paperwork: a business that manages or controls construction work holds contractor duties whatever its marketing says, so claiming to hold none at all would be the more dangerous answer. Under Model 03 the term is Prime Service Contractor, where prime means prime for the site-services system and nothing else.

How do you stop an AI agent inventing requirements?

By making the check mechanical rather than instructional. Telling a model to trace every requirement to its source is a hope. Extracting the references from the finished document and comparing the two sets by machine, on every run, before anybody can approve it, is a control. Ours refuses to issue a tender pack when a scope item has no priced line or a priced line names a scope item that does not exist.

Does AI reduce the cost of construction procurement documents?

It reduces the hours, which is not the same thing. The saving is in first-draft production and in cross-reading documents against each other, where a person reads sequentially and misses contradictions between inputs. It does not reduce the review, and the review is where the liability sits. A document issued without a competent person reading it is cheaper only until it is priced.

Is AI-generated content penalised by Google?

Not for being AI-generated. Google's own guidance on generative AI features says its AI answers are drawn from the same index and the same ranking systems as ordinary search, and rewards content that is helpful, reliable and carries a point of view the reader cannot get elsewhere. What gets penalised is commodity content, whoever or whatever wrote it.

Where this connects to the work

Sources

Primary sources only. Where this piece states a position rather than a fact, it says so on the line.

  1. The Construction (Design and Management) Regulations 2015 — legislation.gov.uk
  2. Summary of duties under the Construction (Design and Management) Regulations 2015 — Health and Safety Executive
  3. CDM 2015, Schedule 2 — Minimum welfare facilities required for construction sites — legislation.gov.uk
  4. Housing Grants, Construction and Regeneration Act 1996, Part II — legislation.gov.uk
  5. L24 — Workplace health, safety and welfare: Approved Code of Practice — Health and Safety Executive
  6. Optimizing your website for generative AI features on Google Search — Google Search Central

About the author

Justin Ngolu Nseya — Founder and Managing Director, ETABLIX. ETABLIX is one accountable partner for the temporary site environment and workforce accommodation around the permanent works. It is not a main contractor: it does not build, design or commission the permanent asset. More about the business, or connect on LinkedIn.

Share LinkedIn Email RSS

Read next

Ask us the eight questions.

We would rather be asked what our system refuses than what it can do. Start with a the Site Systems Diagnostic, or just put the questions to us directly.

Discuss your requirement