HR AI Strategy

Why HR AI Pilots Stall: The Four Places It Actually Happens

Why HR AI Pilots Stall: The Four Places It Actually Happens
Contents
  1. The shape of a pilot that stopped moving
  2. Stall one: the security review that nobody wrote the answers to
  3. Stall two: it works on the export and dies on the live system
  4. Stall three: shipped, announced, unused
  5. Stall four: nobody settled what working would look like
  6. When the answer is that it shouldn’t have started
  7. What to do differently on the next one
  8. Start with the one that’s already stuck
  9. Frequently asked questions

Month six, and the pilot is neither live nor dead.

It’s still on the roadmap. The status field says in progress, because it said that in March and nobody’s changed it since. When your CEO asks how the AI work is going you say the pilot’s going well, and the sentence comes out fine, which is the part that bothers you.

Picking the right idea doesn’t protect you from this. Ordering the queue is separate work, done earlier, and three ideas taken through it end to end shows what it produces. A well-chosen HR AI pilot stops in the same four places a badly chosen one does, and two of the four sit with your CIO.

The shape of a pilot that stopped moving

Stalled pilots don’t announce themselves. They look like projects having a slow quarter.

Nobody kills it, because killing it means the choice was wrong and the choice was defended in front of the exec team. Nobody ships it either, because whatever’s in the way lives in somebody else’s quarter.

The budget is the smallest thing this spends. The larger bill lands on your next proposal, which has to clear a room that remembers the last one, and on the people who tried the pilot, watched it go quiet, and drew their own conclusion about what happens to AI here. That shows up nowhere you can see, and it’s the part that moves employee lifetime value the wrong way.

Four places, from the inside.

Stall one: the security review that nobody wrote the answers to

Your pilot went to security in week two and it’s still there.

Almost nobody gets refused. What happens is a sequence of questions, each one sending somebody away to find out, and each round trip costing a fortnight. Which account does the model run in, and which region. What’s kept once a session ends, by the provider and by you. Does anything you send train anything, and is that in the contract. Who at the firm building it can open an employee record, and what happens to that access at handover. What’s the lawful basis, and does this need a data protection impact assessment. Does the works council see it before anything gets connected.

Every one of those is fair. What makes HR heavier than the last project your security team reviewed is the material. Your function holds the most personal records in the company, and some of it lands in the GDPR’s special categories without anyone deciding to collect it: health data through sickness and accommodations, union membership through payroll deductions. Consent is a weak basis when the party asking is the employer, so the lawful basis has to sit on something sturdier. In Europe, the EU AI Act classes AI used for recruitment, promotion, termination and task allocation as high-risk, and Article 22 of the GDPR already restricted decisions taken solely by automated means where the effect on a person is significant. A slow review on HR AI is a review working correctly.

So arrive with the answers already written: an architecture document, a data-flow diagram, a retention table, and a page listing what the system will never decide. Answering against a published standard gives the conversation a shape, whether that’s the NIST AI Risk Management Framework or ISO/IEC 42001.

Two moves change the arithmetic. Name the reviewer in week one and put them in the room while the idea’s still soft: security asked at the end is an audit, security asked at the start is a design input that costs one meeting. Then run the proof on synthetic or de-identified data in a sandbox they sign off, so your team gets a fortnight with a working system while the production questions get answered in parallel. That’s the proof step, scoped on its own.

Stall two: it works on the export and dies on the live system

Somebody pulled a file on a Thursday. The pilot read it, did something impressive, and everyone in the demo agreed.

Production needs different things. A service account that outlives whoever created it. A refresh that runs without anyone remembering to run it. An owner for the join between two systems that disagree about what a manager is. Definitions that hold when a field gets renamed. And a system of record whose only real interface, on your contract, is that same export.

None of that is technically interesting. It’s hard because it’s real work nobody allocated. It reaches your IT team as an unticketed request from a project whose status is already green, behind everything carrying a date. That’s the whole stall: a queue nobody has been asked to make room in. Ownership belongs to whoever sponsored the pilot. Your CIO can schedule the work. Only you can make the case that this quarter is when it happens.

Prevention costs an afternoon. Before anything gets built, write the production path for every field the system reads: which system of record, who owns the definition, how often it refreshes, whose name is on the pipeline, and what happens the day the schema changes. Any line that reads “we’d export it” can’t ship. Then pick one of two roads, both fine: fund the path with its own date and owner, or reshape the pilot onto data that already moves.

And that export is a copy of employee records in a folder with no expiry, still there long after the pilot is forgotten. Delete it on a date somebody owns. That’s the readiness lens on the framework, in one line.

Stall three: shipped, announced, unused

The build finished. There was an email, a slide at the all-hands, a training session with forty people on the call, and a channel for questions. By week four the users are the project team and two enthusiasts.

The announcement is usually where it went. A tool introduced as an HR rollout reaches a manager as one more thing the function wants from them, and the first question anybody asks about software HR built is who can see what they type into it. Answer it before it’s asked, in plain words, and make it true, because people test it. Risely’s coach, Merlin, is used by 5,000+ people from 40+ companies, and usage there tracks one thing: whether the privacy line is real. The week anyone suspects it isn’t, they stop typing. That’s most of what running our own taught us.

Five things earn use, and none is a training session. Narrow scope, so the system is obviously good at one job and says so. Visible reasoning: it shows the clause, the source, the record it read, so a person can check it in four seconds. It lives where the work already happens, in Slack or Teams. It states what it can’t do and hands those cases to a person, unanswered. And somebody makes the call at the end.

That last one is a design commitment. We don’t build systems that make autonomous decisions about people. AI may inform, surface, rank, summarize or draft. A person makes every call about a person, and the system is built so it can’t do otherwise. A button to disagree with an answer the system already reached is a different build from owning the call, and that button gets pressed less every month.

For the launch, pick the group carrying the most volume. Seniority is the wrong selector here. Let them use it for a few weeks before anything gets announced, publish the limits beside the capabilities, and let that team’s experience be the demo. The signal you want is people using it for something you didn’t design.

Stall four: nobody settled what working would look like

The review meeting arrives. Somebody asks what it moved.

What’s available: a usage chart, three good quotes, a general feeling that people like it. All real, none of it able to survive a CFO’s second question. So the meeting ends with an agreement to keep going and look again next quarter. Do that twice and it has stopped being a pilot.

The gap opened months earlier. Which number this is meant to move, whose report it comes from, what it read before you started, when you’d look, and what result would end the funding. All five take an afternoon between them. Tying an output to a decision, an owner and a number is worked through in employee insights that change a decision, and it applies to a build as much as to a survey.

If you’re already at that review with nothing to show, there’s a version that still works. Pick a number somebody outside HR reports and take the baseline from before the pilot started, since their report has been running the whole time. Say out loud that you’re setting the bar late, then agree the threshold and the date before the next batch of data lands. A bar set once the data is in front of the room gets set where the data already is.

And if nothing anyone outside HR reports could plausibly move, that’s a finding worth the six months. The chain to a number finance already carries is what to fix before the next one.

When the answer is that it shouldn’t have started

Some pilots stall because the idea was the wrong shape, and unblocking doesn’t fix a shape.

Four tells. It only works if three systems get joined at the level of one named person. The work it shortens is a decision about somebody, and the shortening was the point. Security has already declined the one data source that makes it work. Or the sponsor moved on and nobody inherited the number.

Not yet is a result. It comes with a named blocker and a date to score it again, and usually it’s the shape of the idea that changes first. Three moves cover most reshaping: change the grain, change who decides, or change what the system can read. A per-person flight-risk score drops to segment grain, which is a smaller build, clears the review you’re stuck in, and names nobody. A system that tells a manager what to fix from one exit interview becomes quarterly de-identified themes at function grain, read by your HRBPs. That version has some trust to earn, and it can be built this year.

If the answer is to close it, close it properly. One page: what you learned, what would have to be different, and the date it gets scored again. A pilot you ended on purpose with the reason written down costs your next proposal nothing. A pilot still open in month nine costs it a great deal.

What to do differently on the next one

The stallWhat it looks like from insideThe question that prevents itWhose answer it is
Security reviewWeeks of questions, a fortnight per round tripWhich of your questions can we answer in writing first?Your CIO or CISO, with you
Integration and dataWorks on the export, needs a pipeline nobody scheduledWhat’s the production path for every field, and whose name is on it?Your HRIS or data owner, with IT
AdoptionLive, announced, used by the people who built itWhat makes somebody reach for this on a Tuesday, unasked?You, with the team carrying the volume
No agreed numberA review with usage counts and warm quotesWhich number, whose report, and what result ends this?An owner outside HR

Notice how little of that is technical: a slot in one quarter’s IT plan, one written security answer, one launch group chosen for volume, one number with somebody’s name on it. Two habits carry the rest. Put the security reviewer and the data owner in the first conversation, while the idea can still change to suit them. And write the kill condition beside the success condition, because a pilot with a defined ending is one nobody has to be brave to close.

Start with the one that’s already stuck

Take the pilot on your roadmap and work out which of the four it’s in. Then name the person who owns the unblock and give them a date. Most stalls survive on nobody being asked directly.

For a new idea, the 10X HR-AI Framework is the two-minute version of this conversation. Eight questions, and what comes back is an honest tier, one of which says not yet. The map of 122 HR agents across 12 areas shows where the idea sits, and who decides on each.

The strategy month is where a whole list gets sequenced, tested against what your data and your security posture can carry this year, and priced against eLTV. Here’s what that month covers. If the unblock looks like hiring somebody, four kinds of employee experience consulting sets out what you’d own at the end of each, and where we’d tell you to keep the money.

Frequently asked questions

How long should an HR AI pilot run before we call it?

Set the end date when you set the start date, and make it weeks. A proof on test data shows how accurate it is, where it fails, what it costs to run and whether anyone reached for it. None of that needs a quarter. Open-ended pilots stop being experiments around month three.

Our security review has been open for two months. What now?

Ask for the full questionnaire in one go, answer every line in a single document, and ask for a written decision with the conditions attached. Then ask which parts of the review a sandbox version removes. Most of what’s slow is round trips, and a document collapses them.

Should the pilot run on real employee data?

Usually no, at this stage. Synthetic or de-identified data in an approved sandbox tells you almost everything about whether the system works, and skips the heaviest part of the approval. Real records belong in the production build, once the review has cleared it.

We killed one and now nobody wants to try another. How do we restart?

Restart with something small and visible where being wrong is cheap, and tell people what happened to the last one. A function that watched a build get closed for a stated reason will fund a second. One that watched a build go silent hedges on everything you propose next.

  • HR AI Pilot
  • HR AI Strategy
  • CHRO
  • CIO
  • eLTV

Have a problem worth solving?

Thirty minutes. Your business, your bottleneck, and whether AI is the right tool for it.