Skip to content
AI Is Already Interpreting Your Employees' Feedback. We Built the First Way to Measure Whether It's Any Good.

AI Is Already Interpreting Your Employees' Feedback. We Built the First Way to Measure Whether It's Any Good.

Key Takeaways: PYX Labs has launched PYX-Voice, the industry’s first benchmark for evaluating how AI models handle employee feedback. While frontier models are fluent, the study found they struggle with nuanced, emotional judgment—scoring between 54% and 76% against expert criteria. The results highlight that while AI is useful for theme identification, it falls short on the nuanced, high-stakes calls, where expert-grounded judgment is what makes AI trustworthy.

AI has already moved into the middle of people decisions. In a 2025 survey of more than 1,300 U.S. managers, six in ten said they use AI to help make calls about their direct reports: raises, promotions, layoffs, and terminations. HR teams are analyzing engagement survey results with general-purpose AI tools to find themes. Managers are using consumer chatbots to navigate difficult conversations. Leaders are drafting action plans from listening pulse data without guardrails or I-O psychology grounding.

And until now, no one could answer a basic question: how good is any of this, really?

For objective work like math and coding, the AI industry has public benchmarks that make a model's strengths and weaknesses clear. For the work that touches employees directly, understanding how people feel, interpreting what they mean, recommending what to do about it, there was no such yardstick. That work is subjective and high-stakes, which makes "good" both harder to define and more consequential to get wrong.

This isn't an abstract concern, either. Ethan Burris, who studies employee voice at the University of Texas at Austin and advises PYX Labs, puts the stakes plainly:

"Employees don't always speak up, and when they do, whether it leads to anything depends on how well the organization hears them. That makes employee feedback some of the highest-stakes, most easily misread data there is — tied to people's livelihoods and to who holds power at work. We need to ask not whether a model sounds capable — we need measures that assess whether it meets the bar that expert practitioners actually hold and what executives need to make better decisions."

The risk is compounding. A model can sound authoritative while misreading what employees actually mean — or, in rare cases, fabricate a statistic that was never in the data — and that output can flow straight into a decision about someone's pay, promotion, or role. Without a way to measure how well a model actually handles employee feedback, organizations have no mechanism to catch those failures before they reach a real person.

To fill that gap, Perceptyx sponsored PYX Labs to create the industry's first benchmark evaluating AI performance on employee listening work.

How we built the first benchmark for AI in employee listening

Perceptyx sponsored the launch of PYX Labs, an independent research initiative with a focused mission: to understand, and help improve, how well AI understands people at work. Its first release is PYX-Voice — the industry's first benchmark evaluating how leading frontier AI models handle real employee-listening work.

  • Models Tested: 7 frontier models from OpenAI, Google, Anthropic, and xAI.

  • Evaluation Scope: 84 tasks based on real-world employee experience analysis.

  • Grading Criteria: 208 specific grading criter defined by I-O psychologists and organizational behavior experts.

That criteria is critical. Anyone can hand a model a spreadsheet of survey data. The hard, valuable part is defining what separates an excellent interpretation from a mediocre one.

 As Melissa Valentine, Professor of Management Science at Stanford and Senior Fellow at the Stanford Institute for Human-Centered AI, who advises PYX Labs, explains: "What makes PYX Labs' approach distinctive is the attention paid to defining what 'good' actually looks like when AI interprets the human experience at work. Most benchmarks measure whether an AI can complete a task. This work asks a harder and more important question: whether AI is applying the right values and expertise when evaluating that task."

Is the recommendation grounded in the data? Built on behavioral science rather than management fads? Does it handle sensitive comments responsibly, and read like something a leader could actually act on? Encoding that judgment into criteria a machine can be measured against is what makes this a trustworthy benchmark.

What benchmark testing reveals about frontier AI models

The short version: today's frontier models are more fluent than they are reliable, and no single model is good at everything.

  • AI Strengths: Predictable themes and consistent feedback (e.g., Performance Enablement).

  • AI Weaknesses: Nuanced, emotional, and context-dependent feedback requiring human judgment.

  • Overall Model Range: Scores ranged from 54% to 76%, with instances of data fabrication and drift.

Model performance also varied by model and by task: the top overall score was 76%, with the field ranging down to 54%, and leadership shifting depending on the specific task. And in rare but real cases, models fabricated a statistic or drifted from the constraints of the underlying data. That matters because AI outputs can flow straight into decisions before anyone independently checks them. 

None of this says 'don't use AI.' It says something more useful. Know where these tools are reliable, and where they still need expert judgment, better prompting, or human review before they touch a real decision.

Why HR leaders need a standard for measuring AI accuracy

For the HR leaders and teams we work with, the takeaway is practical. You can't manage what you can't see, and PYX-Voice gives the industry its first clear view of where AI is trustworthy on people work and where it isn't. That's the difference between deploying these tools responsibly and hoping for the best.

For the industry, it matters more. Perceptyx has spent two decades helping the world's largest organizations understand and improve the employee experience. PYX Labs is the next evolution of that mission. We're working to make the AI models themselves better at understanding people. Doing that credibly takes deep vertical expertise, the technical ability to build and run these evaluations, and years of real employee-experience data across industries, roles, and geographies.

That combination is why we can define the standard, and why frontier labs and those building AI agents have reason to work with us to reach it. The benchmark shows where models fall short; expert post-training and evaluation are how they improve. Either way, the outcome is the same one we've pursued for twenty years: a better, more trustworthy experience for the people doing the work.

As Melissa Valentine, Professor of Management Science at Stanford and Senior Fellow at the Stanford Institute for Human-Centered AI, who advises PYX Labs, put it: "Most benchmarks measure whether an AI can complete a task. This work asks a harder and more important question: whether AI is applying the right values and expertise when evaluating that task. The workplace is one of the most consequential domains for AI to get right, and work like this is what the field needs to move from capability to trustworthiness."

What's next for PYX Labs and benchmarking workplace AI

PYX-Voice is the first in a planned series of benchmarks, with employee listening as the starting point because it's core to what we do. We're already looking toward coaching, development, and the other ways employees and leaders rely on AI to understand and act on human behavior at work.

The full benchmark results, methodology, and model-by-model scores are available at pyxlabs.ai.

What HR leaders should take away from the first AI benchmark

If there's one thing to carry out of this, it's that judgment can't be outsourced to a machine. Not on work this human, and not yet. Frontier AI is genuinely useful for employee-experience work, but it's uneven: the model that's strong on one topic can be unreliable on the next.

Knowing those limits, for AI in general and for the specific model in front of you, is now part of doing this work well.

It's also a reason to be deliberate about the tools you choose. General-purpose models weren't built for the subjective, context-dependent, emotionally loaded signal in employee feedback; systems tuned for it, with expert judgment and guardrails built in, are. That's the standard we built PYX-Voice to define. Now there's a defined standard for evaluating model performance on employee listening tasks—one grounded in I-O psychology criteria rather than general capability claims. We'd rather measure it and help make it better.

Explore the full PYX-Voice benchmark at pyxlabs.ai and see how the models you rely on performed.

Frequently asked questions

How is AI being used in HR?

In a 2025 survey of more than 1,300 U.S. managers, six in ten said they already use AI for decisions about direct reports: raises, promotions, layoffs, and terminations. HR teams paste engagement survey results into large language models to identify themes. Managers use AI tools to draft talking points for difficult conversations. Leaders use it to build action plans from employee listening pulses. The use cases are real and growing, but the quality of AI output varies widely depending on the model and the task.

Is HR being replaced by AI?

No. AI handles certain HR tasks well, especially where employee feedback is consistent and maps to well-defined categories. But it performs unevenly on work that requires real judgment: reading nuanced, emotional, or context-dependent feedback and turning it into a clear recommendation. In rare but documented cases, AI models have also fabricated statistics or drifted from the actual data. HR professionals who apply expert review, set clear guardrails, and understand where a specific model is reliable are the ones who get the most from these tools, without handing over decisions the work doesn't support.

What are the risks of using AI for HR decisions?

Three risks stand out. First, reliability gaps: in the first industry benchmark of AI on employee-listening work, top model scores ranged from 54% to 76%, and performance shifted significantly by task. Second, fabrication: models most often overstated what the data supported, and in one case a model invented a statistic that wasn't there, a serious problem when outputs feed directly into decisions about raises, promotions, or terminations. Third, model mismatch: general-purpose AI was not built for the subjective, high-stakes signal in employee feedback. Using a model not trained or evaluated on this type of work increases the chance of misreading what employees actually mean.


PYX Labs is a research lab sponsored by Perceptyx, focused on defining evaluation standards and expert post-training for how AI systems understand and support people at work. It is advised by Melissa Valentine, PhD (Stanford University / Stanford HAI) and Ethan Burris, PhD (UT Austin).

Subscribe to our blog

Opt-in for our weekly recap and never miss a post.

Getting started is easy

Advance from data to insights to focused action