Skip to content
Perceptyx Built the Most Accurate Employee Feedback Sentiment Model

Perceptyx Built the Most Accurate Employee Feedback Sentiment Model

Key Takeaways: Perceptyx has released a new sentiment model for employee feedback, and out of the box it beats every frontier and general-purpose AI we tested. For customers, it replaces the former production model and scores comments with 89.4% accuracy — the same agreement rate our own I/O psychologists reach with each other. This more accurate read means a cleaner signal feeding everything downstream, from themes and intent to the narratives our other agents build. The only approach that matched it was a frontier model with a painstaking, professionally tuned prompt, at far higher cost and slower speed. Bottom line: the most accurate option for reading sentiment from employee comments is now a model built for the job, not the biggest LLM you can find.

We’ve released a new sentiment model for employee feedback, one built to read a comment more accurately and give every analysis downstream a cleaner signal to work from. For customers, it replaces the previous production model, and with no further tuning it beats every frontier and general-purpose AI we tested. It scores comments with 89.4% accuracy, the same agreement rate our own I/O psychologists reach with each other, which is as good as this task allows for a model or a trained I/O psychologist. It’s live now across Comments Report, Comment Copilot, Third-party Dashboards, and the Narrative Analysis Agent, and it scores every new listening event automatically.

It’s worth noting that this isn’t our first sentiment model, either. Our earlier models performed well, but we know it takes iteration to get this right — reading workplace language well takes the kind of employee-listening data and I/O expertise most AI is never trained on. The only thing that came close this time was a frontier model running a carefully and professionally engineered prompt, and it still cost far more per comment and ran far slower across the tens of thousands of comments that real enterprise-grade listening programs produce. So for anyone who’s been getting by on a general AI tool, or who looked at employee-comment sentiment a while ago and decided it wasn’t accurate enough to trust, here’s the update: the most accurate way to read your people’s comments is no longer the biggest, most expensive LLM. It’s the model we built for exactly this.

  • At 89.4% accuracy, our fine-tuned encoder model beat every approach we tested, including small language models, zero-shot methods, off-the-shelf sentiment models, and out-of-the-box frontier models. With a carefully engineered prompt, a frontier model drew level with ours, while costing far more per comment and taking far longer to run.
  • A domain-specific benchmark is what made this possible. Until we could score an experiment against an objective, domain-specific benchmark, we had no way to confirm which of the approaches was the right direction.
  • Human feedback can be genuinely ambiguous. Even human annotators disagreed about sentiment on about 7% of our benchmark data. Some comments can be interpreted more than one way, which sets a realistic ceiling on the task for models and humans alike.

Why Does Sentiment Matter in Employee Listening Data?

When analyzing employee feedback, sentiment is a quick signal about where to look first. Paired with quantitative results, it helps explain the "why" behind a score. For example, your favorability scores may show that career growth is a challenge; employee comments reveal what is behind that result, such as limited visibility into advancement opportunities, unclear career paths, or a lack of development support. Comment data often contains mixed or contradictory sentiment — a single response can praise a manager while criticizing a process — and manually sorting thousands of responses is inconsistent and slow. Manually reviewing large datasets is time-consuming and difficult to scale, and differences in how reviewers interpret responses can lead to inconsistent findings.

AI offers scalability, but most sentiment models are trained on public content like product and movie reviews. That training data is a poor proxy for employee listening data. Generic tools often rely on keyword-level signals, treating words like "fine" or "okay" as neutral when the surrounding context makes them clearly negative. Survey comments frequently contain mixed sentiment within a single response, such as praise for a manager alongside criticism of a policy. They are also written in response to a specific question, so the same short answer can carry opposite meaning depending on what was asked. Rule-based approaches that scan for positive or negative keywords miss these patterns entirely. An accurate sentiment model needs labeled workplace data and an architecture that reads each comment as a whole.

How Did Perceptyx Build a Domain-Specific Sentiment Model?

How Did We Build a Reliable Benchmark?

We built the benchmark before the model, using 1,000 anonymized survey responses labeled independently by four I/O psychologists against a shared rubric.

We used a dataset of anonymized survey responses, sourced from customers who consented to provide their data for benchmarking purposes. The data was aggregated and de-identified across many sources, then randomly sampled, so no single customer's information is identifiable or reconstructible.

Before any labeling started, the team wrote a rubric: what makes a comment negative rather than neutral, how to handle comments that are positive about a manager and critical of a process. Four of our I/O psychologists then used this rubric to label 1,000 responses as positive, neutral, or negative, working independently so that no one's judgment influenced anyone else's.

Even expert humans disagreed: on roughly 70 of the 1,000 responses, the four annotators split evenly, two against two. We reviewed those cases individually and found complex comments that were difficult to label.

That disagreement rate set a ceiling on the task. In a large enough dataset, some percentage of comments will always defy classification within a narrow taxonomy. No model can achieve a perfect score, and neither can a human reviewer working alone.

Which Model Performed Best?

A fine-tuned encoder model achieved 89.4% accuracy. We fine-tuned it on a separate set of survey responses, synthetically labeled with the same rubric our I/O psychologists use.

Encoders are built to classify text as a whole, not generate new text. That architecture suits sentiment analysis: a comment like 'my manager is great, but the workload is unsustainable' shifts halfway through, and an encoder reads the full response before scoring any part of it. The result is a single confidence score across all three labels — 85% negative, 12% neutral, 3% positive — that is consistent, auditable, and fast to produce at scale.

Sentiment analysis benefits from that architecture because sentiment is a property of a whole comment, not something that accumulates left to right. For example, "my manager is great, but the workload is unsustainable" flips halfway through.

An encoder reads the entire response at once, so the second clause is available when it interprets the first. A generative model reads forward and has to carry the meaning along with it as it goes.

An encoder's output also suits this task. It returns a confidence score across all three labels at once: 85% negative, 12% neutral, 3% positive. That makes it simple to train, use, and evaluate.

Our own benchmark results support this directly: the fine-tuned encoder, trained specifically on employee listening data, matched or outperformed every general-purpose model we tested.

How Did the Fine-Tuned Model Compare to Other Approaches?

Approach

Accuracy

Key Trade-offs

Fine-tuned encoder (Perceptyx)

89.4%

Fastest inference, lowest cost per comment, explainable confidence scores

Frontier model (tuned prompt)

89.3%

Equivalent accuracy, higher cost and latency at scale

Frontier model (out-of-the-box)

85.7%

No prompt engineering needed, lower accuracy

Small language model

88.5%

Competitive accuracy, less efficient than an encoder

Zero-shot NLI

84.8%

No labeled training data required, lowest accuracy tested

Frontier models: 85.7% accuracy, increasing to 89.3% with a tuned prompt. The best performance, from Claude Opus 4.8, scored 85.7%. Tuning the prompt increased the score to 89.3%, tied with our fine-tuned model. Smaller frontier model variants scored 83.8%. Across all frontier options, costs per comment are meaningfully higher, latency compounds across tens of thousands of responses, and it is harder to audit why a specific comment received the label it did — a real concern when HR leaders need to explain their findings to executives or works councils.

Small language models: 88.5% accuracy. Competitive performance, and cheaper than a frontier model, but without the efficiency of an encoder or the accuracy of the fine-tuned version.

Zero-shot natural language inference: 84.8% accuracy. NLI models judge whether a hypothesis follows from a premise, which can be reframed as a sentiment question without any training data. The results were comparatively strong for an approach that needs no labeled examples, but fell short.

Why a Purpose-Built Model Is the Right Choice?

General-purpose models aren’t trained and evaluated on survey questions. They’re trained on text that is available, and score employee comments the same way they’d score product reviews or social media posts. We built this one around employee survey comments: the benchmark it was measured against, the rubric behind the labels, and the way it interprets a comment alongside the question that prompted it. Three things follow from that.

The benchmark uses the right data. Our I-O psychologists labeled real employee comments against a rubric written for this task, so the test reflects real-world conditions. A domain-specific benchmark is what distinguishes meaningful improvements from noise.

The model scores comments and questions together. "Not really" means something different in response to "do you feel supported by your manager?" than in response to "is your workload manageable?" Scoring both together removes that ambiguity.

The model is consistent. Because we version and control the model, labels also remain comparable across surveys and years — making trend analysis reliable.

Where Is the New Sentiment Model Available Today?

The new model is automatically applied to all future events launched with Perceptyx. You'll see the results alongside our other models for theme and intent wherever you already work with comment data. The model also scores comment data collected outside the Perceptyx platform through Third-Party Dashboards.

Sentiment analysis turns thousands of individual comments into a pattern leaders can act on. That only works if the labels underneath it are accurate. With a model scored against the same rubric our I/O psychologists use, organizations can move from reading a sample of comments to acting on all of them.

Frequently asked questions

What is AI sentiment analysis?

AI sentiment analysis is the automated process of reading a piece of text and classifying the emotional tone as positive, negative, or neutral. In employee listening, the model reads open-ended survey comments and returns a label, so analysts can see where feedback clusters before reading every response manually. Most general-purpose models are trained on product reviews or social media posts. Models fine-tuned on workplace feedback classify employee comments more accurately because the language, context, and question-response format differ significantly from consumer data.

Can you use ChatGPT or other general AI tools for employee sentiment analysis?

Yes, but with trade-offs on accuracy, cost, and consistency. In Perceptyx's benchmark testing, the best-performing frontier model scored 85.7% accuracy out of the box. With a carefully engineered prompt, that rose to 89.3%, matching the fine-tuned Perceptyx model. The practical problem is cost and speed: frontier models charge more per comment and take longer to process, so the gap compounds across the tens of thousands of comments a typical listening program generates. A fine-tuned, task-specific model running at the same accuracy level costs meaningfully less and returns results faster. There is also a consistency concern. General AI tools update frequently, which means the same comment can score differently across time periods, making trend analysis unreliable.

Is AI sentiment analysis accurate enough to act on?

At 89.4% accuracy, the Perceptyx sentiment model matches the agreement rate of trained I/O psychologists labeling the same comments. That figure also reflects a realistic ceiling on the task: even expert human annotators disagreed on roughly 7% of the benchmark responses because some comments carry genuinely mixed signals. For leaders using sentiment as a directional signal, such as identifying which survey topics skew negative before reading comments in full, that accuracy level is sufficient. For higher-stakes decisions, leaders should review the underlying comments and broader context alongside sentiment scores.

What does AI sentiment analysis look like on a real employee comment?

Consider the comment: "My manager is great, but the workload is unsustainable." A general-purpose model reading left to right may carry the positive tone from the first clause and label the whole comment positive. A fine-tuned encoder model reads the entire comment at once, so it registers the negative second clause alongside the positive first clause. The model then returns a confidence score across all three labels (for example, 72% negative, 20% neutral, and 8% positive) and assigns the dominant class. That output tells an analyst to flag the comment for workload-related themes rather than manager effectiveness, which is where the actual concern sits.

What are the main limitations of AI sentiment analysis for employee feedback?

Three limitations apply to any sentiment model, including purpose-built ones. First, some comments are genuinely ambiguous. When Perceptyx's I/O psychologists labeled 1,000 survey responses, they split evenly on roughly 70 of them. No model can score those cases correctly every time because humans do not agree on the right answer. Second, general-purpose models trained on product reviews or social posts perform worse on workplace language, where phrasing like "not really" or "it's fine" carries different weight depending on the survey question asked. Third, models that update or change versions produce inconsistent labels over time, which breaks trend analysis. A versioned, domain-specific model addresses the second and third limitations, though the first remains a property of the task itself.

Want a fuller picture of how listening programs are evolving industry-wide? Read The State of Employee Listening 2026 for benchmark data from 750+ HR leaders.

Ready to see the new sentiment model on your own comment data? Talk to a Perceptyx expert.

Subscribe to our blog

Opt-in for our weekly recap and never miss a post.

Getting started is easy

Advance from data to insights to focused action