HR AI Strategy

HR Data Readiness: What Each of Your Systems Can Actually Carry

HR Data Readiness: What Each of Your Systems Can Actually Carry
Contents
  1. Existence, ownership and quality are three different questions
  2. The stack, one system at a time
  3. Core HRIS
  4. Applicant tracking
  5. Learning platform
  6. Engagement and survey tools
  7. HR case and ticketing
  8. Payroll and time
  9. The joins are where plans actually die
  10. Where the data lives, and whose name is on the number
  11. Frequently asked questions

Three weeks ago your CEO forwarded something about AI in HR and asked what the plan is. Since then the ideas have piled up. Screening support for the roles you hire in volume. Something that answers policy questions so your operations team stops fielding the same six every Monday. An attrition model, because the CFO keeps asking why the sales org turns over.

Any of those could be the right first move. What decides it sits underneath all three: what your systems can hand to a build this year, and what they’d hand over with a confidence they haven’t earned.

Nobody has one HR system. Sapient Insights Group’s 2025-2026 HR Systems Survey covers 22 HR technology segments, and a two-thousand-person company tends to have live data in six or seven of them. Each was bought for a different job, in a different year, and each holds a different thing well.

Here’s how to read yours.

Existence, ownership and quality are three different questions

Most readiness conversations stop at the first one. Somebody asks whether we have performance data, somebody else says yes, we’ve run reviews for years, and the room moves on.

Existence isn’t ownership. Ownership isn’t quality. All three are separate questions.

Existence is visible on a screen. Ownership is narrower: a named person who maintains the fields you’d depend on and would notice if they went wrong. Owning the system doesn’t count, and neither does owning the vendor relationship. You want whoever would know that half the location fields stopped updating after last year’s office moves.

Quality only becomes answerable once somebody looks, which is why one probe belongs before everything else. Can you pull a representative sample right now, today, without filing a ticket and waiting a week? And has anyone opened it?

Exports get requested, land in a shared drive, and nobody reads past row fifty. Sort each column you’d depend on and look. You’ll find the blanks, the placeholder values from implementation week, and the date somebody typed as 01/01/1900 because the field wouldn’t let them leave it empty. Twenty minutes of that teaches you more than any framework, this one included.

The stack, one system at a time

Four things are worth writing down for each system. What it holds that you can trust. What only looks trustworthy. What an idea reading it could do this year. And the single repair that unlocks the most.

Core HRIS

Trustworthy. Who works here, their job title of record, employment status, start date, legal entity, and today’s reporting line. Payroll depends on all of it, so mistakes get found fast.

Looks trustworthy. History, and anything optional. Effective-dated history is often the first casualty of a migration, so “who did this person report to in March 2024” can be a reconstruction. Termination reason codes are the quiet one: picked by whoever closed the record, from a list nobody has revisited, and “resignation” absorbs everything awkward.

What an idea can do this year. Anything about the present. Headcount and cost views, spans and layers, org structure questions. What it can’t do is anything needing the past to be as accurate as the present. An attrition-driver model built on reason codes learns your dropdown habits and hands them back as insight.

The one repair. Check the history on the five fields your top three ideas depend on. Five fields done properly beats a data quality program that covers everything and finishes next year.

Applicant tracking

Trustworthy. The event trail. Application received, stage changed, interview scheduled, offer sent, offer accepted, dates on all of it. A system built to run a process logs the process well.

Looks trustworthy. The reasons, and the identities. Rejection reason is a dropdown filled at speed by someone with forty more to get through. Source attribution breaks on referrals, reposts and anyone who saw the job in three places. And the same person applies more than once over the years, sometimes under a second email, so your candidate count runs ahead of your candidate population.

What an idea can do this year. Time in stage, drop-off, panel load, scheduling, drafting job ads and debrief summaries. Naming which channel produces your best hires is out of reach. That takes attribution you don’t have, plus an outcome definition living in another system.

The one repair. Resolve candidate identity first. Then decide once, in writing, how a referral that also arrived through a job board gets attributed.

Learning platform

Trustworthy. Assignment and completion. Who was assigned what, who finished, when, and when the certificate expires. Deadlines with consequences attached keep this data honest.

Looks trustworthy. Skills. The course-to-skill mapping in most learning platforms arrived with the vendor’s catalog and describes that catalog, not your job architecture. Self-rated proficiency is a mood. Interest and career-goal fields were filled in during the first month after the platform went in and haven’t moved since.

What an idea can do this year. Assignment tracking, expiry chasing, content curation, drafting learning material. What it can’t do is give you a skills inventory you’d act on, because the skill words in the learning platform and the skill words in your job descriptions are two vocabularies that were never introduced.

The one repair. Map the twenty skills that matter for the roles you can’t afford to lose by hand, once, across both vocabularies. A week of unglamorous work moves several ideas from impossible to straightforward.

Engagement and survey tools

Trustworthy. What people said, and when. One wave’s response set is a clean, timestamped record, and the open text is some of the richest material in your stack.

Looks trustworthy. Everything comparative. Question wording changes between waves, so the trend line has a seam in it the chart doesn’t show. The hierarchy used to roll results up was a file uploaded on fielding day, a snapshot of a company that has since reorganized, so last year’s teams may not exist in a form you can match to today.

What an idea can do this year. Theme synthesis across open text inside a wave, drafting the summary a leadership team reads. Connect the roll-up hierarchy to the responses and an idea can trace engagement into team-level attrition. Skip that step, and the connection isn’t there.

The one repair. Save the org snapshot with every wave from here on, and freeze the wording of the questions you intend to trend. Both cost nothing on fielding day and are unrecoverable afterwards.

HR case and ticketing

Trustworthy. Volume and timing. When a case arrived, when it closed, who touched it, how many times it bounced.

Looks trustworthy. The category. Categories in a case system were designed to route work, so they answer “who handles this”. And the record only holds questions that became tickets. The same question asked in a corridor or a DM is invisible here, so the visible volume understates the demand your team absorbs.

What an idea can do this year. This is often the readiest corner of the stack for a first build. The useful material is the question text, and it’s sitting right there. A policy answering agent grounded in your approved handbook doesn’t care whether the category was right, and neither does intake routing, which reads the question and treats the category as a suggestion. Plenty of the HR operations and shared services agents start here for that reason.

The one repair. Leave the taxonomy alone. Start from the question text and let the categories get fixed later by something that reads better than a dropdown does.

Payroll and time

Trustworthy. The most accurate data you own. It’s reconciled every cycle by people whose job is making it balance, and an error generates a phone call within a day. Pay, hours, cost center, entity, effective dates.

Looks trustworthy. Its accuracy belongs to payroll’s purpose. Cost center gets read as org structure, and it’s a finance construct that survives reorganizations by accident. The manager field is often whoever approves hours, which can be a shift lead or a budget owner with no people-management relationship at all. Compensation columns carry currency, FTE, effective dates and allowances that differ by country, so one total-comp number across entities is usually a fiction assembled in a spreadsheet.

What an idea can do this year. Cost and headcount reporting, anomaly detection on runs, answering payroll questions from documented rules. Pay equity analysis isn’t happening this year. It needs comparable roles, and comparable roles need a job architecture most companies are still arguing about.

The one repair. Agree what counts as a comparable role before anyone compares pay. It’s a people decision, it belongs to you, and it unblocks more analysis than any tooling will.

The joins are where plans actually die

Almost every HR AI idea worth doing touches two systems. An attrition model needs the HRIS, the survey tool and something about pay. A manager-facing assistant needs to know who reports to whom, which sounds like the easiest question in the building.

Inside one system the data is at least internally consistent. Wrong sometimes, and wrong in a stable way you can describe. Between two systems you get disagreements nobody owns, since each is correct by its own definition and neither team has had to reconcile them.

Three disagreements account for most of it.

Who this person is. Your systems key people differently: an employee number in the HRIS, a candidate ID in the applicant tracker, a work email in the learning platform, a payroll number in payroll, an SSO identity in your identity provider. Match on email and you’ll cover most of the population. The group you miss is never random. It’s the rehires, the contractors who converted, the population that arrived through an acquisition with its own numbering, and the frontline staff who never had a company email. Those tend to be exactly the people your question was about.

Who their manager is. This one earns its own table, because six systems will give you six answers and each is defensible on its own terms.

SystemWhat it calls a manager
Core HRISThe supervisor on the position hierarchy, as of the last effective date
Applicant trackingThe hiring manager named on the requisition
Learning platformWhoever assignments roll up to, often set at import
Engagement and surveyThe roll-up used on fielding day, frozen inside that wave
Case and ticketingThe approver in the workflow, sometimes a shared inbox
Payroll and timeWhoever signs off hours, or owns the cost center

Ask how many managers you have and the honest answer is a range. Any idea that changes what managers see, or measures them, has to pick one definition and defend it, and that’s a decision with your name on it. A good share of people analytics work turns out to be this decision, made once and made properly, followed by the analysis everyone thought they were asking for.

Which entity they sit in. Legal entity, business unit, cost center, country, region. A reorganization re-cuts one or two of those and leaves the rest, so a slice that reconciles this quarter won’t reconcile against last year’s. In a multi-country company, “Europe” means four different populations depending on which system drew the line.

A fourth hides inside all three, and it’s time. Systems disagree about when. Effective date, entry date, pay period, fielding date, snapshot date. A join that reconciles today can be quietly wrong the moment somebody asks about last March, and it won’t announce itself. It returns a number that’s slightly off, which is worse than one that’s obviously broken.

We ran into a clean version of this on an offshore energy build. The documents an agent would need all existed, and people knew where they were. Ownership sat across several regional IT teams, with no single access path reaching all of them. Nothing about the idea was wrong. The AI scoping waited behind a consolidation phase, because a system that reaches two thirds of the documents answers confidently and incompletely, which is the worst behavior available to it.

That’s the shape of it. Ideas rarely die because one system holds bad data. They die in the gap between two systems that were each doing their own job correctly.

Two things to do with this. For your top three ideas, write down which systems each one touches and what the join key is. If a join takes more than a sentence to describe, you’ve found your first project, and it’s smaller and duller and more valuable than the idea it was blocking. And a first build inside a single system is worth real weight: a working thing, a named owner and a number, while the join work runs on its own timeline.

Where the data lives, and whose name is on the number

The Readiness lens in our 10X HR-AI Framework asks two things about any idea: where the data lives, and whose name is on the number. Everything above is the long answer to the first. The second is shorter and often harder, since somebody has to say they’ll defend a number in the exec meeting on the month it moves the wrong way.

If the question on your desk is the board’s version, where HR sits on AI and what to say about it is the post for that. Readiness looks inward at your systems and answers whether you can start. Maturity moves in stages, looks outward at the board, and answers where you are. And if the list of ideas has no order to it yet, three of them scored end to end works through that.

So take the idea at the top of your list and run it through the scorer. The readiness questions will land differently now you know which systems it touches. Getting that answer for the whole list, with the order to build in and the repairs sequenced underneath, is what the strategy month produces.

Frequently asked questions

Do we need a data warehouse before we can build anything?

For a first build, usually no. A warehouse is one answer to the join problem, and a good one over a few years. It’s also why plenty of HR AI plans go quiet for eighteen months. Pick a first build inside one system, get it working with an owner and a number attached, and let the warehouse conversation run on its own schedule.

Our HRIS vendor says their embedded AI already reads all of this. Does that change the answer?

It changes who does the plumbing. What the fields contain stays the same, since embedded features read the same records with the same blanks and the same reason codes. Two questions are worth putting to any vendor: which fields does it read, and what does it do when one of them is empty. There’s a legal version of the second question, and it’s narrower than it first looks. The EU AI Act treats some employment uses as high risk, including systems used for recruitment and for decisions on promotion and termination. For those, Article 10(3) requires training, validation and testing data sets to be “relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose.” That’s a duty on the data a model is built and tested on, and it sits with whoever provides the system. Plenty of HR AI never lands in that category, and the two questions are worth asking anyway.

Who should actually do this read?

Two people and about a morning per system. Someone from HR operations who knows what the fields mean, and someone from IT who knows how data moves between systems. HR alone will miss the pipes. IT alone will call a field clean because it’s populated. Sit them in front of the same export, let them disagree, and write down what they disagree about. That list is most of your answer.

  • HR Data Readiness
  • HR AI Strategy
  • HRIS
  • People Analytics
  • CHRO

Have a problem worth solving?

Thirty minutes. Your business, your bottleneck, and whether AI is the right tool for it.