Buying change, not a plan

Estimated read time: 8 minutes

A vibe-coded world doesn’t only change how we build. It changes what it means to fund and procure delivery responsibly. Much of our commercial machinery still assumes that working software is expensive and slow, so it asks for confidence up front, manufactured through documents.

AI-assisted delivery collapses that premise. When real progress can be shown in weeks, and when the right answer may only emerge through iteration, the old model starts demanding false certainty. If funding and procurement stay locked to rhythms measured in months while delivery gets measured in days, the result is not control. It is either paralysis, or teams sprinting ahead and leaving legitimacy to be rebuilt later.

So the shift is simple to state and hard to do well: move from front-loaded certainty to ongoing validation. This is not less rigour, it is different rigour. A rigour based on proof, not fortune-telling.

Fund in tranches, release against evidence

The most practical move is evidence-based tranche funding. Not a single £5m bet placed on a plan, but a sequence of smaller commitments released against what the work has actually shown.

An initial tranche funds an Alpha mission with a clear expectation: show something working, show what it changed, show what you learned, show what risks remain. That “something” might be a prototype tested with users. It might be a policy engine exercised against edge cases. It might be an operational walkthrough run with frontline staff. It might be code in production behind a feature flag with early metrics. The point is that the funding decision is anchored in reality, not rhetoric.

The next tranche is released when the evidence justifies it: not because the paperwork is immaculate, but because the team can credibly say, this is working, this is safe to scale, and here is what we will prove next. And if it isn’t working, tranche funding gives you a dignified off-ramp: you stop early, capture the lessons, and avoid “completing” a programme that no longer deserves to exist.

This can look like stage-gating. The difference is that the gate is not a document. It is evidence: user outcomes, operational impact, risk reduction, and observable service performance.

But there is room for caution: tranche funding can easily become micro-gatekeeping – death by a thousand finance reviews. Adopting this model requires clear delegations, predictable and lighter-touch cadences, and a bias towards reusing assurance once it has been earned. Otherwise you simply swap one slow system for a slower one. 

Value for money becomes easier to show (and harder to fake)

In a vibe-coded world, “value for money” should become easier to demonstrate. And harder to bluff.

When change is cheap, the question is no longer “can we afford to build it?” but “are we moving the numbers that matter?” Completion rates, processing time, avoidable contact, cost-to-serve, staff handling time, errors, appeals, complaints – and, crucially, equity of access and differential outcomes. These are not afterthoughts. They are the public value of the service, made visible.

This is where finance and product language meet. A good tranche review is a short story, told with evidence:

  • What outcome did we target?
  • What changed for users and staff?
  • What did it cost to achieve that change?
  • What risks did we remove, and what risks did we discover?
  • What do we believe now that we didn’t believe six weeks ago?

That last question matters. In government, “learning” is often treated as an indulgence. In this model, it is the mechanism by which you protect the public purse.

Buy capability and partnership, not monolithic deliverables

Procurement has to change shape as well. Large contracts that specify a fixed scope years out don’t fit a world where the right answer is discovered in contact with reality. The stance becomes: buy small missions, buy specialist help in bursts, and keep ownership of the evolving service.

That can mean more use of flexible call-offs, outcome-shaped statements of work, and multi-supplier arrangements that allow you to pull in expertise when you need it – accessibility, security, content, data engineering, service design, performance analysis.

AI assistance also changes who can credibly deliver. It can level the playing field for smaller suppliers with deep domain expertise and strong engineering discipline, because the productivity gap between a “big team” and a “small expert team” narrows. Procurement should be open to that: favour agility and demonstrable competence, not just scale and familiar logos.

And if you want a cleaner test of capability, you can increasingly ask for working evidence early: not “a beautiful proposal”, but a short, time-boxed mission (that you fund) which produces a prototype, an analysis, or a working integration in a sandbox environment, so claims meet reality quickly. 

Make “the safe path” a contractual requirement

If more of the multidisciplinary system can contribute to runnable artefacts, then standards cannot live only in people’s heads. They have to live in defaults, and you can insist on those defaults commercially.

Contracts should require modern delivery practices and shared visibility, so government can see what is being built as it is built. This is the shift from black-box delivery to partnership-based work that can be safe at pace.

At minimum, suppliers should be required to:

  • work in the open with shared code and shared environments where appropriate
  • provide the artefacts that make services operable (runbooks, monitoring, alerts, incident patterns)
  • treat accessibility, security and data protection as non-negotiable “definition of done”, not bolt-ons
  • support feature-flagged release and safe rollback
  • provide evidence of performance and user outcomes, not just “scope delivered”

In other words: if delivery can go faster, the bar for operability needs to go up.

Treat AI as part of the supply chain

If AI tools are woven into delivery, they become part of the supply chain, and therefore part of the commercial and assurance landscape. It would be easy to ban them, but that would be counterproductive. We need to be explicit about:

  • what tools are used, and in what contexts
  • what data can and cannot be shared with them
  • how provenance is managed (what generated what, when, and under which constraints)
  • how AI-produced code is held to the same standards as human-written code
  • how IP, reuse and exit are handled so government is not locked out of its own service

In a vibe-coded world, prompts, policies and agent guardrails are important artefacts. As accountability is non-negotiable, then traceability is part of the price of admission. Not to punish teams, but so that when something goes wrong we can understand what happened, and fix it properly.

There’s also a sovereignty question hiding in the plumbing. If our ability to deliver depends on tools we can’t explain, can’t exit, or can’t run without a single external choke-point, we have moved risk, not removed it. In some cases that may still be acceptable but it has to be named, governed, and designed around, not discovered during an incident.

We will not certify the whole upstream story through procurement alone, but we can insist on a floor: clarity about what tools are in use, what data they touch, what gets retained, and what the exit looks like if we need to change course. If we rely on it, we should be able to describe it plainly to an auditor, a minister, and ourselves.

Rethink the on-ramps: procuring the road, not just the vehicle

All of this collapses if the on-ramps are broken.

There is no point having AI-assisted delivery if the route to a safe environment takes three months; if security review is a queue rather than a collaboration; if teams can generate change quickly but can’t deploy it safely. In a vibe-coded world, the bottleneck shifts from “writing code” to “getting onto the road”.

That means investing in enabling infrastructure: secure-by-default environments, one-click preview deployments, identity and access patterns, logging and monitoring, test automation, component libraries, feature flagging, and the guardrails that let more people contribute without lowering the bar.

Cross-government procurement and strategy needs to reflect that reality: less bespoke infrastructure per programme, more shared platform capability that makes the safe path the easiest path. This is not only a technical procurement question. It is organisational: who owns the platform, how it is funded, and how it is governed so it serves teams rather than slowing them down.

Bringing it together

The thread running through all of this is that evidence becomes the mechanism by which pace is earned. In a vibe-coded world, we should walk into funding and assurance conversations with something real: “try it”, “see what changed”, “here’s what users did”, “here’s what staff said”, “here’s what the metrics show”, “here’s the risk we removed”. The story stops being “trust us”. It becomes “come and see”. The strategy is delivery.

Funding, commercial and evidence practices have to evolve to keep pace with a vibe-coded world. The goal is not more speed for its own sake. It is faster learning, earlier proof, cleaner off-ramps, and a stronger ability to scale what works. 

The shape of it is familiar by now: fund in tranches, procure for partnership and outcomes, insist that standards live in defaults, treat AI as part of the supply chain, and invest in the on-ramps.

Download the full paper as a PDF.