From tasks to outcomes

Estimated read time: 8 minutes

‘Outcomes over outputs’ is what good product has always been trying to do. The problem is that our day-to-day mechanics – long backlogs, granular user stories, estimation rituals, ticket queues – have a way of pulling teams back into task management, even when everyone agrees that impact is the goal.

AI-assisted delivery doesn’t create a new philosophy. It changes the economics of execution. When small fixes and even more substantial features can be delivered almost as soon as they’re identified, the effort of tracking every task starts to look disproportionate. What remains hard, and therefore valuable, is clarity: what problem are we solving, what would success look like, and what evidence will tell us whether we’re helping or harming?

The backlog as a register of outcomes

A traditional backlog tries to be a list of everything: features, bugs, chores, refinements. In practice it becomes either endless or obsolete, and often both. When a typo, misaligned button, or small content defect can be fixed immediately, the ‘log a ticket and wait’ threshold rises. A lot of work that used to sit in Jira for weeks can simply get done.

What belongs in the backlog is what is enhanced by shared focus and deeper thought: problems to solve, hypotheses to test, outcomes to pursue. So the backlog can become a register of opportunity statements rather than a laundry list. Items like:

  • Improve accessibility of the appointment booking flow
  • Reduce drop-off at step three of the claim journey
  • Reduce average handling time for caseworkers processing complex cases

These aren’t bigger tickets. They’re better prompts. They invite the team to decide how to pursue them and keep the conversation anchored in purpose: if a change doesn’t plausibly move an outcome, why are we doing it?

This also clarifies the craft of product and delivery leadership. The job isn’t to manage a queue; it’s to cultivate and prune: keep the most important problems visible, retire what no longer matters, and stop the backlog becoming a museum of stale intentions.

A learning log matters as much as a to-do list

Speed increases the risk of organisational amnesia: repeating experiments, re-litigating decisions, rebuilding solutions that failed last year. Documenting, preferably in the open like this excellent work from NHS England, what was tried and what happened prevents teams from unknowingly revisiting dead ends and helps new joiners understand why the service is the way it is.

Something as simple as: ‘We tried X to address Y; it didn’t move the needle; here’s what we learned; here’s what we’d do differently next time’ can save weeks of repeated work and turn speed into cumulative progress rather than churn.

Planning becomes mission-shaped

Whether you keep two-week sprints or move towards continuous flow, planning works best when it is anchored in outcomes, not task lists. If the team can accomplish in two days what used to take two weeks, then a sprint becomes less a commitment to a static list of tasks and more a time-bounded mission within which the exact work will evolve as the team learns.

Stand-ups and boards still matter, but they shift purpose: not defending estimates, but synchronising attention – what changed, what we learned, what we’re doing next.

This makes it easier to behave the way agile always said we should: respond to change over following a plan. AI-assisted delivery doesn’t invent that principle, but it does make deviation cheaper and therefore more likely, so the discipline has to sit at the level of outcomes and evidence.

Roadmaps become hypotheses with targets

Outcome-oriented roadmaps are also not new. Instead of committing to a sequence of features, the roadmap becomes a sequence of value targets and questions. For example:

  • Increase completion rate of online claims from 70% to 85%
  • Reduce average processing time from five days to two
  • Extend reach to a new user group, with explicit equity and access measures

Under each outcome, you might hold provisional approaches: simplify the form, add guided support, automate a verification step, redesign a back-office workflow, but the roadmap is not a promise to ship a specific solution on a specific date. It is a promise to tackle the problem, to measure honestly, and to keep trying until the evidence says you’ve made meaningful progress.

This aligns well with an environment where if one idea doesn’t work, you can try another without paying a huge switching cost. But it does require leadership that is comfortable with adaptive delivery: trust that the team will deliver value by the end of the quarter even if the exact outputs evolve. In exchange, there will be more value, not less: teams that remove small impediments continuously and test multiple approaches empirically without clinging to pet features just because they were presented in slides weeks ago.

Reporting becomes a narrative of impact

In faster environments it’s easy to confuse activity with impact. This is where communication rituals like weeknotes become more valuable, not less, but their content shifts. The purpose isn’t to list what was built; it’s to capture what has changed.

A weeknote might say: ‘We set out to reduce confusion in error messaging. We released two iterations, and calls to the help line dropped by 10%. Next week we’ll test whether that holds, and then pivot to the next source of avoidable contact.’ That kind of narrative keeps everyone oriented around outcomes and forces the team to articulate what it’s learning.

Metrics-driven development becomes more practical

To cement ‘outcomes over tasks’, teams can organise work directly around the measures that matter to them: uptake, completion, avoidable contact, processing time, cost per transaction – with ideas for experiments under each of them. For public services we can think more expansively about how to include things like equity of access, differential outcomes for different groups, and operational sustainability.

The point isn’t to reduce everything to numbers. It’s to ensure that every meaningful piece of work can link to public value, and that you can quickly identify when that is absent.

Measures are a sharp tool. They can illuminate reality, and they can be weaponised. In public services, numbers are often harder than they look: baselines shift, user behaviour changes seasonally, policy changes contaminate results, and proxy metrics invite gaming. It is risky to hand a brittle metric to a difficult stakeholder who will treat it as a performance target rather than a learning signal. The discipline is not to avoid numbers; it is to treat them with the same care we treat code: define them, version them, attach confidence, and pair them with narrative and context. A good metric is a learning instrument first, and a performance instrument second.

The deeper prize: removing work from the system

If you care about deeper efficiency, and we should, the goal isn’t “ship more”. It’s to remove work from the system: fewer steps, fewer handoffs, fewer avoidable contacts, fewer appeals, fewer workarounds, fewer duplicate tools doing the same job. AI makes it easier to change services. The real prize is needing fewer changes because the service is simpler, and because the underlying estate is being paid down rather than patched around.

Managing the anxiety of “what exactly will you deliver?”

Some stakeholders will feel a loss of predictability. ‘What exactly will you deliver, and when?’ is not an unreasonable question, especially in environments shaped by business cases, programme controls, and supplier contracts. The answer is not to pretend certainty; it’s to be explicit about your confidence in the outcome you’re pursuing, the evidence you’ll use and the transparency in how you tell the story.

There is more integrity in saying, ‘By June we aim to halve the backlog of pending cases; we will try a number of approaches and publish what we learn,’ than in promising a specific set of features – a promise history tells us is often late, brittle, or not successfully solving the real problem.

Bringing it together

This is not a new doctrine. The above are tried and tested practices that lead to good outcomes for users – and in a vibe-coded world we have fewer excuses not to do them. Shifting from tasks to outcomes means reimagining the artefacts and conversations that organise the work: the backlog fulfilling its role as a strategic tool; sprints as missions; roadmaps articulating value targets and learning questions; reporting that tells stories about impact, evidence and the value being added. 

This is not a more casual approach. AI-assisted delivery makes it easier to live up to the ambitions of agile more fully: focus on the problem, iterate rapidly, measure honestly, and judge success by what improved for users and the wider system, not by how many tickets got closed.

Download the full paper as a PDF.