public digitalThe public digital logo
View all insights

3 lessons from 3 years spent building AI inside the UK government

Our new Senior Director and former Director of the UK Government's AI Incubator, Alex Jones, reflects on three core learnings from his time from testing and building AI in government.

Last week, the IGC published a playbook written by myself and my former colleagues at the UK's AI Incubator (i.AI), where I spent a year as Director. 

It is an honest account of what we built, what we got wrong, and what we learned prototyping and piloting AI across government.  

Three years ago, i.AI was three or four people spun out of the No.10 data science team. It's now around 60. In that time we shipped tools like Consult, Extract and Redbox, and killed several projects that weren't working, including a medication-safety tool. In those instances, we published the failure modes in full

It is in the spirit of that openness that we felt our wider experiences were worth writing up. 

Here are the three lessons I think matter most for anyone building and testing AI for public services.

1. Agree evaluation criteria before you build. 

Consult is i.AI’s tool for analysing the public’s responses to formal government consultations on matters such as online safety and education reforms. 

For Consult, we agreed evaluation criteria with departments before development started. It was tested on 26 live consultations, and we got it to a point where human reviewers agreed with Consult more than they agreed with each other. That finding was credible and powerful precisely because the criteria was agreed before we saw what the model could and couldn’t do, and was about the real-world quality of the analysis it produced.

2. High alignment, high autonomy, and rapid iteration matter more than the tech. 

We ran i.AI on high alignment and high autonomy. Through quarterly portfolio reviews we set and revisited the problem statements, and teams decided how to solve them. Every product passed through stage gates with explicit kill criteria, and where we sunsetted a project, we treated that as the process working - as success, not failure (an approach that PD has long advocated for). 

We kept feedback loops as short as possible. Products were demoed live to real users on real data, and we took pains to avoid vanity metrics, looking instead for the signals that mattered: whether teams asked to use the tool again afterwards; and whether it improved the quality of their work. 

But in government, none of that survives without a sponsor. This is one of the core lessons from any successful government AI initiative, as my colleague Alex MacEachern wrote about recently in relation to the Government of Alberta’s PRISM initiative

Our first recommendation for anyone starting out is to secure a named minister or senior official with the authority to give the team air cover. 

3. Many AI projects in government are data infrastructure projects

Extract, i.AI's planning-data tool, exists because most English local authorities have not digitised their planning records into a single asset. No AI application can work effectively if it sits on top of data that is unusable, contradictory or incomplete. Often, we needed to build the data layer first, and treat that as the actual product.

Digitisation is only the first step. Our medication-safety work was only possible because we built a secure data asset with two million NHS patient records. And our education programme requires a common data architecture and digital curriculum so AI tutors know what a pupil should be learning, has already learnt, and how they're being taught. The pattern holds internationally: Ukraine built Diia before Diia.AI, and India built Aadhaar and UPI before its AI mission. Countries getting value from government AI are those that have the data foundations in place. 

A further lesson sits underneath all these which may be obvious but warrants stating: governments will not get any value from AI without permanent, in-house people who stay long enough to build and hold the knowledge. 

Government therefore needs to put significant effort into attracting and retaining the talent to do this, because the cost of getting this wrong (such as through poor procurement) outstrips the investment needed to get this right. 

As PD's CEO Ben Terrett recently argued in City AM: government's consultancy habit isn't really a spending problem, it's a skills problem, and the civil service has let go of the wrong people. 

My experience at i.AI bears that out. Where governments need external support, they need support that builds capability rather than breeds dependency. That's the model which PD builds into every engagement, working alongside teams long enough that they don't need us for the same problem twice.

That's what I’m taking from my time leading i.AI, and that I'm bringing to PD’s work helping governments and multilateral institutions invest in and use AI for public good. 

Read the full reportHow to build a high-impact AI team inside the state

Written by