Contents
- In HR, trust sets the order
- The three lenses, briefly
- Example one: a policy helpdesk for the HR ops inbox
- Example two: screening support for a hiring team
- Example three: flight-risk scores in every manager’s dashboard
- The order, and why the boring one goes first
- The line the framework won’t cross
- What the order is worth
- Start with one idea, then take on the list
- Frequently asked questions
Fifteen ideas, and no defensible order.
Three arrived with vendors who demoed to your team last quarter. Four came from your CEO, who reads the same newsletters as every other CEO. The rest came from your own people, who can tell you exactly which part of their week they want back. A copilot for HR. A policy chatbot. Something that reads exit interviews. Screening. A dashboard that flags who’s about to quit.
Every one of them is plausible. That’s the problem.
Your CEO wants to know which one is first and why that one. Sorting by vendor readiness gives you their roadmap. Sorting by projected return gives you a spreadsheet nobody can defend, because every return on it is your own estimate.
There’s an order that holds up, and it starts somewhere most people don’t expect. Below: three real HR ideas scored end to end, landing in three different places, then the build order and why the least exciting one goes first.
In HR, trust sets the order
Most AI prioritization advice sorts by return. Value on one axis, effort on the other, start with the top-right box. In supply chain or finance that holds up, because a wrong answer costs a correction and a rerun.
HR doesn’t clear that way. A wrong answer here lands on a named person, and named people talk to each other. The cost shows up as a Slack screenshot, a works council meeting, a review site, or a quiet decision by three good people to stop telling you anything true.
Trust works like a shared account for the whole function. Every system you ship draws on it or pays into it. Someone who watched an AI answer a leave question correctly, cite its source, and hand the personal part to a human is a different audience for your next proposal. Someone who watched a system quietly score their team is a different audience too, in the other direction.
This is where employee lifetime value moves without anyone noticing. eLTV is how well people perform and how long they stay, across their whole time with you. It’s the number we’ve worked on since 2017, when Culturro started as a people consultancy. A system people don’t trust changes how they behave long before it changes a dashboard. They route around it. They feed it worse data. Some of them leave, and the exit interview says something else.
So your first build should be the one where being wrong is cheap and being right is visible. That’s almost never the idea with the biggest business case attached.
The three lenses, briefly
| Lens | The question it asks | What it does to eLTV |
|---|---|---|
| Impact | Which number does it move, and how much of the work does it take on? | Raises it |
| Trust to earn | What does a wrong answer cost, who decides, and how sensitive is the data? | Protects it |
| Readiness | Where does the data live, and whose name is on the number? | Keeps it |
Impact and readiness get answered from what you already have: the numbers you report and the names on them. Trust to earn runs the other way. The more of it an idea has to earn before anyone will use it, the later it belongs in the queue. We write that lens in words, Little, Some, Real, A lot, The most, because a score invites averaging, and a strong impact number should never cancel out a trust problem. Inside the lens the worst of the three answers sets the reading, so a cheap mistake never softens sensitive data.
Four tiers come out of the combination: build now, proof first, build with guardrails, not yet.
The framework page has the scorer, and its worked example is an interview scheduling agent, about as easy as HR AI gets. The three below sit at the awkward end of the list.
Example one: a policy helpdesk for the HR ops inbox
Your HR operations team answers the same questions every week. Notice periods. What counts as a dependant. Which parental policy applies to someone hired into the German entity but working from Lisbon. The idea: an agent that answers from your approved handbook, shows the clause it used, and hands anything personal to a human without attempting an answer.
| Lens | Reading |
|---|---|
| Impact | A real number. It moves HR tickets per hundred employees and first-contact resolution, numbers your ops lead already reports, and it takes on most of the volume. |
| Trust to earn | Some. Every answer carries its source clause, so a wrong one is cheap and easy to catch, and no decision about a person sits inside it. What keeps it off the bottom of the scale is the data: it reads location, entity and worker type to pick the right policy version. That’s ordinary employee record, and it’s the answer that sets the lens. Grievance, health, harassment and pay disputes go straight to a person, unanswered. |
| Readiness | Strong, once you scope it to one country and entity. The policies live in one approved space, and your ops lead owns the ticket number and would defend it upstairs. |
Tier 1. Build now.
The proof is small. Load one country’s policies, hand the agent a few months of questions your team has already answered, and let it answer them blind. Then run it live for a pilot group. It works when it cites the right clause, picks the right version for the person asking, and steps back from the personal questions on its own. It’s on the map as a proven agent.
Example two: screening support for a hiring team
Some of your roles draw several hundred applications. Recruiters read the first stack properly and skim the rest. The idea: an agent that reads every application against the must-have criteria you published, summarizes the evidence for and against each one, and orders the pool for recruiter review.
| Lens | Reading |
|---|---|
| Impact | A real number. Recruiter screening hours and time-to-shortlist move first. Further out, so does the quality of who reaches interview, which feeds the front of the eLTV curve. |
| Trust to earn | A lot. Get it wrong and a candidate is treated unfairly at the screen, which is as expensive as a wrong answer gets while the guardrails hold. The recruiter still decides who advances. The agent reads the published criteria and what the candidate wrote, nothing protected. What makes this heavy is the obligations that have to exist before day one: in New York City it counts as an automated employment decision tool, with an annual bias audit and candidate notice, and in the EU it’s high-risk under the AI Act, so your works council will want the conversation first. |
| Readiness | Mixed. The applications and stages sit in your ATS, clean enough. The criteria are the problem: they have to be published before anything can screen against them, and in most teams a good share of them still live in a hiring manager’s head. Your TA lead would own time-to-shortlist. Nobody has volunteered to own the selection rate by demographic group. |
Tier 3. Build with guardrails.
Tier 3 comes with a work package attached, and the work happens before anything is connected. The agent never declines or advances anyone. Every application reaches a human, including the ones at the bottom of the queue. It reads only the published criteria and infers nothing from names, addresses, schools or dates. Notice and audit exist on day one. Skip that work and you’ve shipped a compliance problem with a user interface. The guardrail list sits on the agent page.
Example three: flight-risk scores in every manager’s dashboard
The demo was good. Every employee carries a score, green, amber or red, and each manager gets a weekly list of who’s drifting toward the exit. Your CEO liked it.
| Lens | Reading |
|---|---|
| Impact | On the slide it looks enormous. Under the lens it thins fast. Regretted attrition is the right number to chase. What nobody can say is what a manager does differently on Tuesday because a name turned amber. The work is the conversation, the conversation stays human, and the load this system carries is close to nothing. |
| Trust to earn | The most. It puts a prediction about a named person into a manager’s hands, and it can’t be unseen. A wrong amber changes how someone gets treated for a year. It wants performance, survey and tenure data joined at individual grain, the most sensitive join in your company. In the EU it reads as employee monitoring, and your works council will treat it that way. |
| Readiness | Thin. Below a few thousand employees, the movement history is too small for individual-grain predictions to mean much. And when you ask who’d put their name on the score in the exec meeting, the room goes quiet. |
Tier 4. Not yet.
The trust reading settles this one on its own. Nothing impact or readiness can say will move an idea with the most to earn into a build tier.
Not yet is a result, and it comes with the one thing to fix first. Here that’s the shape of the idea. Cohort attrition forecasting answers what your CEO is actually asking, which is how many people you’ll need to backfill and where. It runs at function-by-level grain, can’t be queried down to a list of names, and hands workforce planning a range with the assumptions attached. That version is on the map, and it still sits late in the queue, because it needs years of clean movement history and a driver analysis your leaders accept.
The order, and why the boring one goes first
| Order | Build | Tier | What it earns you |
|---|---|---|---|
| 1 | Policy helpdesk | Build now | Hours back for HR ops, and a workforce that has watched an AI cite its source and hand off anything personal |
| 2 | Screening support | Build with guardrails | Recruiter hours, a faster shortlist, and a notice-and-audit pattern you can reuse |
| 3 | Cohort attrition forecast, in place of the flight-risk dashboard | Not yet | Backfill planning your planning lead will sign, once the history and driver work are there |
The helpdesk is the least interesting thing on the list and it goes first anyway. Four reasons.
It’s the cheapest place to be wrong. A bad answer about notice periods gets corrected in the same thread it appeared in.
It’s the fastest place to be seen being right. Everyone who asks it something watches the citation appear, and watches it step back when the question turns personal. That’s a demonstration you can’t buy with a policy document.
It builds the muscles the second project needs. By the time TA proposes screening support, your HR team samples logs by habit, your legal team has read one of these designs, and your works council has had a conversation about AI that went fine.
And it tells you what’s broken. The log of questions a policy helpdesk couldn’t answer is the most useful audit of your policy library you’ll ever get, and it arrives free.
In the strategy months we run, the boring build is the one the room argues about hardest, and the one that changes every conversation after it. Skip it and those costs still get paid, in the middle of a harder project with more people watching.
The line the framework won’t cross
One rule doesn’t move. We don’t build systems that make autonomous decisions about people. AI can inform, surface, rank, summarize and draft. A person makes every call about a person, and the system gets built so it can’t do otherwise.
Two things get used interchangeably here and shouldn’t be. “A person decides” is a different design from “a person can override.” Override is a button, and buttons stop getting pressed by week three, when the queue is long and the machine has been right most of the time. Deciding means the person owns the call, sees the whole pool, top to bottom, and the system has no path to act without them.
That’s why the flight-risk dashboard scored the way it did. Nothing in it technically decides anything. In practice it hands a manager a verdict about a named person before the manager has had a thought of their own.
Where the controls sit in your own build is your decision. We recommend, in writing, then build what you agree to.
What the order is worth
Three lenses, one number underneath. Impact raises eLTV, trust protects it, readiness keeps it.
Read the build order through that. The helpdesk gives your team back the hours they spend retyping the handbook, and it shows your workforce what your AI does when a question turns personal. The screening build raises the front of the curve, since better hires start higher and ramp sooner, and the guardrails stop it costing more trust than it earns. The flight-risk dashboard, as demoed, would have raised nothing and spent plenty.
That’s the argument for sequencing by trust. You get to keep spending it.
Start with one idea, then take on the list
Pick the idea with the loudest advocate and score it on the 10X HR-AI Framework. Eight questions, two minutes, an honest tier at the end, including the tier that tells you to stop.
Then see where it sits in your function. The HR agent map covers 122 use cases across 12 HR areas, with who decides and what the guardrails are for each.
Scoring one idea is an afternoon. Ordering fifteen, working out which ones your data can carry this year, and putting a number on what the list does to eLTV is the strategy month. Here’s how that runs.
Frequently asked questions
What if two ideas both come out Tier 1?
Start with the one whose owner is in the room. A Tier 1 with a named person who’ll defend the number upstairs beats one everyone likes and nobody owns. If both have owners, take the one with more volume, because a proof runs on evidence.
Does sequencing by trust mean the high-impact ideas never get built?
No. They get built second or third, with the guardrails designed first and a function that has already seen AI behave well. Tier 3 is a work package with a build at the end of it. The ideas that never make it out of the queue were Tier 4 in a costume: a decision about a person, dressed as a recommendation.
How long before the first one is live?
The proof step for a policy helpdesk runs a couple of weeks on sample data, and you decide on what you watched happen. Build time depends on how many policy versions and entities you carry. Each use case is a fixed fee for an agreed scope, quoted after the proof.
Who should own the list?
You, with your CIO or CISO reading the same document. They’ll ask where the data sits, what gets retained, and what trains on what. Those answers move readiness scores, and hearing them now costs far less than a build stopped in month three.
The CEO wants the flashy one first. What then?
Show them the scorecard for the flashy one beside the scorecard for the boring one, and the order those two produce on their own. A Tier 4 with a named fix and a rescore date reads as diligence to a board. A pilot that stalls in month five reads as something else.