Voice of the Employee: A Real Say While the Change Is Still Moving
Voice of the employee is an old idea with a new way to gather it: ordinary spoken conversation, from everyone rather than a sample, fast enough to still change a decision. The evidence for it is real but uneven, and this is an account of where it holds and where it does not.
HECTOR BENITEZ VENTURA AND NOAH ALEXANDER · LATENT VARIABLES
The published record splits on one variable, and it is not the technology. Where a machine stands between a person and something they need, people object. Where it asks what they think and leaves the choice open, most take it up. Everything else here follows from which side of that line a program is built on.
Change has become a continuing condition
The next period of organizational change will be shaped less by a single program than by several overlapping shifts. Organizations are integrating acquisitions, deploying AI and other major technologies, redesigning work, reorganizing teams and testing new approaches to wellbeing and retention, usually while normal operations continue. The World Economic Forum's 2025 employer survey[1], covering more than 1,000 employers representing over 14 million workers, gives a sense of the scale employers themselves expect.
Each number answers a different question. Open one to see which.
World Economic Forum, Future of Jobs Report 2025, a survey of more than 1,000 employers representing over 14 million workers. These are employer forecasts and not measured outcomes. The three figures answer different questions and are not parts of a single total.
This pace creates room to improve how work gets done, and it raises the value of hearing from people while a change is still in motion. Research across industries links employees' own experience of frequent or extensive change to strain: a two-wave study of 260 employees[2] tied perceived change to insomnia through affective resistance, and frontline nursing research on 746 nurses[3] documented change fatigue after large-scale reorganization. Neither finding argues against necessary change. Both argue for understanding how a change is landing while there is still time to adjust it, which is a more useful moment than a retrospective.
Where people go unheard
Most organizations have good methods for the beginning and the end of a change. At the beginning, leaders set out a strategy, hold town halls and distribute toolkits. At the end, they measure adoption, performance, engagement or retention. The stretch in between is where people have the most to say and the fewest ways to say it, during the weeks and months when they are interpreting the change, finding workarounds, meeting friction and deciding whether the new way of working makes sense.
Annual surveys provide comparability. Pulse surveys provide speed. Focus groups and interviews provide depth. Manager check-ins provide context. Each of these remains valuable. The gap opens when an organization needs depth from many people, quickly and consistently, and a rating scale or a small text box cannot hold the answer.
An employee voice agent is an AI-led spoken interview built for that gap. It asks a consistent core set of questions, listens to an answer in ordinary language, follows up where that helps, and turns the resulting conversations into aggregate patterns for people to interpret. It does not make decisions, score individuals, infer emotion, or feed anything into hiring, promotion or performance review. It is a way to collect better information before leaders decide what to do.
What voice changes
The case for voice does not rest on speaking being better than typing. It rests on a well-designed conversation combining three qualities that have been difficult to achieve together: consistency, depth and reach. Every participant can receive the same structured core, while follow-up questions invite examples and draw out the reasoning behind an answer. Conversations can be offered across shifts, sites and time zones without recruiting a large interviewer team.
The most useful distinction in the published record is whether AI removes access or creates it. People object, reasonably, to a system placed between them and a person they need to reach: 64 percent of 5,728 surveyed customers[12] would prefer that companies did not use AI for customer service. Their response looks different when AI adds an opportunity to speak, preserves choice and makes clear that an accountable leader will use what is learned.
Open a bar for who was asked, and what they were asked.
Left: people trying to get something they need, with AI standing between them and a person (Gartner, 5,728 customers, December 2023, customer service). Right: people invited to say what they think, free to choose the AI or a person (Jabarian and Henkel, choice arm of a roughly 70,000-applicant field experiment, 2025, recruitment research). Two different questions with separate denominators. They are not two parts of one total.
What the evidence shows so far
The research base is early and uneven. It spans peer-reviewed experiments, observational deployments, public-agency data and clearly labeled working papers. Taken together it does not show that voice agents are ready for every employee conversation. It does mark out a credible frontier: conversational AI can gather structured information at scale, add useful detail, and in some settings reduce the pressure people feel to manage how they come across, provided the task, the safeguards and the handoffs are well designed.
Open a study for what it does not show.
Evidence snapshot. These studies answer different questions and should not be treated as a single pooled estimate. See sources 8 to 10 and 13.
The most directly relevant peer-reviewed experiment compared a voice-based virtual interviewer with a chatbot[8]. Across 80 UK participants and 2,265 responses, the voice condition produced more informative and detailed answers and more time-efficient engagement. Satisfaction did not improve significantly, and some participants disliked turn-taking delays or found the avatar uncanny. A separate real-world telephone study[9] reached 75 people in the United States and 2,739 in Peru. Quality on structured items approached human-led standards, while the AI interviewer was weaker at qualitative probing.
The largest test of scale so far comes from recruitment research, and it remains a working paper. About 70,000 people were randomly assigned[10] to be interviewed by a person, by a voice AI agent, or to choose between the two. Among those free to choose, 78 percent chose the AI, and 80 percent of the women did. The AI conversations covered more of the interview guideline in less time, and the structured format produced less variation in outcome by gender than 131 human interviewers produced. People made every decision in the study. The setting is recruitment, which is not the same as a conversation about how a change is landing, and the finding that carries over is the choice itself.
One likely mechanism is reduced evaluation pressure. In a 2014 health-screening study[11], people who believed a virtual interviewer was computer-controlled reported less fear of disclosure and less impression management. That effect is not universal. It suggests that the arrangement around the technology, meaning what participants understand it to be, who will see the information and what could follow, matters as much as the voice itself.
Useful technology is bounded technology
Evidence from adjacent settings shows how much scope decides the outcome. A study of 5,179 customer-support agents[13] found that a generative AI assistant increased issues resolved per hour by about 14 percent, with the largest gains among less-experienced workers. A large ambient-scribe deployment[16] reported substantial savings in documentation time and higher physician work satisfaction. In both cases AI improved the information and administrative layer around human work while people kept responsibility for the judgment.
The United States Internal Revenue Service offers the clearest illustration, and the figures come from independent oversight. The agency's narrowly scoped Collection voice bots[14] handled more than 4.8 million calls, with 40 percent contained without escalation. The Conversational 1040 bot on the main taxpayer line was given a far broader job. In fiscal year 2025 about 23.4 million callers were directed to it[15] and roughly 14.5 million of them, about 62 percent, transferred to a live assistor. The National Taxpayer Advocate reports the reason plainly: the agency cannot carry information from the bot interaction over to the assistor, so people repeat themselves. Across the wider set of automated lines measured on containment, the rate for the year was 2.98 percent.
Open a bar for the scope behind the number.
Containment is the share of calls resolved inside the voice system. Both figures use that measure, and they cover different bot sets and different years. Internal Revenue Service and National Taxpayer Advocate, sources 14 and 15.
The underlying technology did not change between those deployments. The scope, the handoff and the experience did. Voice systems deserve to be judged on usefulness, resolution and the quality of the handoff to a person, and not on the number of conversations they begin.
Where this fits across the change lifecycle
Voice agents earn their place when leaders have a real decision in front of them, when employees hold distributed knowledge about how the work is changing, and when the organization needs more depth than a survey can carry. The use case is a voice tied to a decision, and not generic engagement measurement.
Six contexts. Choose one.
QUESTION TO UNDERSTAND
Where are identity, role, process or customer handoffs creating friction?
WHERE VOICE ADDS VALUE
Day 1, then 30, 60 and 90-day integration pulses, with targeted follow-up by function or site.
Six contexts where the decision is usually still open when the conversation happens. The right timing follows the decision, not the calendar quarter.
Timing should follow the decision cycle. A voice with no defined path to a decision produces data and not influence.
From measuring change to shaping it
Traditional measurement asks whether a change worked. Voice of the employee asks a more useful question earlier: what should be adjusted now so that it works? That requires a closed loop and not a one-time interview campaign.
Open a step to see what it involves.
The value is created by the return path from experience to action, and not by the conversation alone. A program that stops at step three has collected material without giving anyone influence.
A strong program starts with the decision and not the questionnaire. Leaders define what they are prepared to change, which populations hold the relevant experience, and what evidence would alter the plan. The conversation then becomes one input into a broader judgment that may also draw on operational data, manager observations, survey trends and direct dialogue.
The return step is the one most often skipped and the one that matters most. People need to see what the organization heard, what it will act on, what it will not change and why. Research on change consistently points to involvement and to informational fairness: formal involvement in a change had stable positive effects on change-supportive behavior[4] at 18 and 42 months, and perceived informational justice helped sustain commitment to change[5], which predicted later behavioral support and turnover intention. Voice becomes meaningful when it travels into a decision and the path back is visible.
What employees gain
People can answer in their own words, give examples, and describe the context behind a rating or a concern, which a fixed scale cannot hold. Participants can choose when to respond, ask for a question to be repeated, pause, skip, stop, or use a different channel. A consistent core means people across roles, shifts, sites and levels receive the same opportunity to be heard, including those who rarely sit at a desk.
The most substantial gain is a clearer path from daily experience into decisions about how work, technology and support are designed. There is some evidence that the ability to contribute matters on its own: a two-wave study of 152 certified nursing assistants[7] found that voice behavior predicted later job satisfaction after controlling for baseline satisfaction. The responsible claim is narrow. Being able to contribute can matter when leaders listen, respect the input and act on it.
What organizations gain
Leaders can hear where uncertainty, workarounds, workload or role confusion are forming, before any of it hardens into resistance. Aggregate patterns can be compared across sites, functions, tenure or stages of a rollout, and recurring conversations show how experience shifts over time. A repeatable evidence trail shows what the organization was told, what it did, and what happened next, which is useful to the next program and the next leader. An analysis of 17,890 European firms[6] found verbal, written and indirect employee-voice mechanisms positively associated with innovation, though firm-level observational data does not establish causality. A dashboard records the outcome. The conversation records the cause, and what might improve it.
What participants are told first
Voice creates immediacy, which can put one person at ease and make another wary. The opportunity becomes responsible only when six commitments are clear before the first substantive question, and visible to the person being asked.
Open a commitment for what it means in practice.
Each of these is a design and governance decision rather than a matter of tone, and each should be visible to the participant before anything of substance is asked. See sources 17 to 19.
Trust is not a matter of vocal style. It is a set of design and governance decisions that participants can see for themselves. Accessibility asks for more than offering speech. W3C guidance notes that speech input can be essential[17] for people who cannot use a keyboard or mouse, while a voice-only system can create barriers[18] for people with hearing, speech, language, memory, processing or anxiety-related needs. The inclusive design is multimodal, with clear pacing, easy repetition, generous response time, language support and an alternate channel where practical.
The same logic applies to bias. Question wording, speech recognition, follow-up behavior, transcript quality, theme generation and reporting can each produce uneven results across groups. Responsible programs test quality across the groups that matter, keep people accountable for interpretation, and offer a way to correct or challenge an error. These practices align with the governance logic in the NIST AI Risk Management Framework[19].
A well-built system does check whether a conversation produced usable material, so that a thin or off-topic exchange does not distort the aggregate. That check operates on the response and never on the person. The system does not infer emotion, sincerity or truthfulness, does not score anyone, and does not report on an individual. Where an exchange carries little information, the appropriate action is a better follow-up question or a note that the item needs rewording.
Employee voice agents should not be used for covert recording or hidden monitoring, emotion or truthfulness inference, individual performance scoring, automated employment decisions, or retaliation. A channel built to carry someone's voice should not quietly become a monitoring system, and saying so in advance is part of what makes it work.
How to tell whether a program is working
Participation on its own is a weak measure of success. A responsible pilot looks at the quality of the conversation, the usefulness of the insight, the experience of the participant, and what the organization did in response.
Eight measures. Choose one.
Who took part, who did not, and whether access differed by role, shift, site, language or ability.
Eight checks, each on a different way a program can fail. A pilot that reports only participation can look successful while failing every other row.
A useful benchmark is already inside the organization. Any employer running an annual engagement survey knows its own response rate, and that figure is the fair comparison for a new channel. Survey fatigue is widely reported and poorly measured, and most published figures on it trace back to vendor material with no study behind them, so none are cited here.
Where this is heading
The future of voice of the employee is not an AI imitation of the annual survey, and it is not a synthetic interviewer trying to pass as a person. The more consequential opportunity is a standing voice of the employee: a channel people can speak into during a change, with real influence over it, so plans can be adjusted before friction compounds and the organization keeps a record of what its people told it.
Organizations will still need surveys, managers, town halls, focus groups and skilled human interviews. Voice agents add an option between high-scale measurement and high-depth conversation. Their value is greatest where they widen access, preserve agency and connect to a decision that is still open. The organizations that lead here will not be the ones with the most human-sounding synthetic voice. They will be the ones that build the strongest loop from employee experience to organizational action.
REFERENCES
- 1.The Future of Jobs Report 2025. World Economic Forum survey of more than 1,000 employers representing over 14 million workers: 86% expect AI and information-processing technologies to transform their business by 2030, 39% of current skill sets will change or become outdated, and 59 in 100 workers will need training. Employer expectations are forecasts and not guarantees. www.weforum.org/publications/the-future-of-jobs-report-2025
- 2.Subjective Perceptions of Organizational Change and Employee Resistance to Change. Peer-reviewed British Journal of Management study, 260 employees surveyed twice four months apart. Frequent and extensive perceived change was linked to insomnia through affective resistance. Observational relationships do not establish causation.
- 3.Change fatigue: the frontline nursing experience of large-scale organisational change and the influence of teamwork. Peer-reviewed Journal of Nursing Management study of 746 frontline nurses after large-scale change. High change fatigue was reported in both cohorts. The cross-sectional design limits causal claims.
- 4.Change-Supportive Employee Behavior: Antecedents and the Moderating Role of Time. Peer-reviewed Journal of Management study, two-wave panel of 72 hospital employees during a strategic reorientation. Formal involvement in the change had stable positive effects on change-supportive behavior at 18 and 42 months.
- 5.Maintaining Employees' Commitment to Organizational Change: The Role of Leaders' Informational Justice and Transformational Leadership. Peer-reviewed longitudinal study. Perceived informational justice and transformational leadership helped sustain commitment to change, which predicted later behavioral support and turnover intention.
- 6.Direct and Indirect Employee Voice and Firm Innovation in Small and Medium Firms. Peer-reviewed British Journal of Management analysis of 17,890 European firms. Verbal, written and indirect employee-voice mechanisms were positively associated with innovation. Firm-level observational data does not establish causality.
- 7.The relationships between certified nursing assistants' voice behaviour and job satisfaction, work engagement and turnover intentions. Peer-reviewed Journal of Advanced Nursing study, two-wave sample of 152 certified nursing assistants. Voice behavior predicted later job satisfaction after controlling for baseline satisfaction. The design does not show that a new channel for employee voice will cause satisfaction to improve.
- 8.Krajcovic, M., Demcak, P. and Kuric, E. (2026). Talking Surveys: How Photorealistic Embodied Conversational Agents Shape Response Quality, Engagement, and Satisfaction. Peer-reviewed Behavior Research Methods study, randomized voice-agent against chatbot design, 80 UK participants and 2,265 responses. Voice produced more informative and detailed responses and higher, more time-efficient engagement, with no significant satisfaction gain. Some participants reported turn-taking delays or uncanny reactions. Preprint at arxiv.org/abs/2508.02376. doi.org/10.3758/s13428-026-03091-0
- 9.Telephone Surveys Meet Conversational AI: Evaluating an LLM-Based Telephone Survey System at Scale. 2025 preprint accepted at the 80th AAPOR Conference. US pilot of 75 and Peru deployment of 2,739. Structured-item quality approached human-led standards, while qualitative probing was weaker than skilled human interviewing.
- 10.Jabarian, B. and Henkel, L. (2025). Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews. Working paper, about 70,000 applicants across 48 positions and 43 client firms. Human recruiters made every hiring decision. Reported offer, start and retention gains are promising and remain provisional pending peer review, and should not be generalized to all employee conversations. papers.ssrn.com/sol3/papers.cfm?abstract_id=5395709
- 11.Lucas, G. M., Gratch, J., King, A. and Morency, L. P. (2014). It's only a computer: Virtual humans increase willingness to disclose. Peer-reviewed Computers in Human Behavior, 37, 94-100. 239 participants in a health-screening setting. Computer framing reduced evaluation fears and impression management. The result is context-specific and does not establish a universal disclosure advantage. doi.org/10.1016/j.chb.2014.04.043
- 12.Gartner. Survey Finds 64% of Customers Would Prefer That Companies Didn't Use AI For Customer Service. Survey of 5,728 customers conducted December 2023. Useful for understanding resistance where AI is seen to restrict access to a person. It is not directly comparable with research on invited interviews. www.gartner.com/en/newsroom/press-releases/2024-07-09-gartner-survey-finds-64-percent-of-customers-would-prefer-that-companies-didnt-use-ai-for-customer-service
- 13.Brynjolfsson, E., Li, D. and Raymond, L. Generative AI at Work. NBER study of 5,179 customer-support agents. AI assistance increased issues resolved per hour by about 14 percent, with larger gains for less-experienced workers. This evaluates AI augmenting a person, and no autonomous voice agent was involved.
- 14.Using voice and chat bots to improve the collection taxpayer experience. Internal Revenue Service operational account of its narrowly scoped Collection voice bots. More than 4.8 million calls were handled and 40 percent were contained without escalation. This is reported by the deploying agency and the page is historical. www.irs.gov/about-irs/using-voice-and-chat-bots-to-improve-the-collection-taxpayer-experience
- 15.National Taxpayer Advocate 2025 Annual Report to Congress, Publication 2104. Independent oversight. The telephone-service section reports that about 23.4 million callers to the toll-free 1040 line were directed to the Conversational 1040 voicebot in fiscal year 2025 and about 14.5 million, roughly 62 percent, transferred to a live assistor, and that the agency cannot pass information from the bot interaction to the assistor. For the Small Business and Self-Employed voicebots measured on containment, the fiscal year 2025 rate through the week ending 13 September was 2.98 percent. www.taxpayeradvocate.irs.gov/wp-content/uploads/2026/01/ARC_Publication-2104_2025_Web.pdf
- 16.AI scribes save 15,000 hours and restore the human side of medicine. American Medical Association summary of a Permanente Medical Group and NEJM Catalyst evaluation. This large observational deployment reported savings in documentation time and improved physician work satisfaction. It was not a randomized trial, and its subject is ambient documentation and not employee interviewing.
- 17.Speech Recognition. W3C Web Accessibility Initiative overview of who depends on speech input and its broader usability benefits. www.w3.org/WAI/perspective-videos/voice
- 18.Voice Systems and Conversational Interfaces: Cognitive Accessibility Research Module. W3C research guidance on cognitive, language, speech, hearing and anxiety-related barriers in voice systems.
- 19.AI Risk Management Framework. The National Institute of Standards and Technology voluntary framework for managing AI risks to individuals, organizations and society. AI RMF 1.0 is under revision, and its governance logic remains a useful baseline. www.nist.gov/itl/ai-risk-management-framework