Customer effort score (CES): The metric that predicts loyalty

Customer effort score reveals loyalty risk that satisfaction dashboards often miss. Your dashboards may look stable even as experience budgets still grow every cycle. Churn and repeat-contact rates remain stubborn because customers remember long holds and repeated explanations. CES turns the ease of each resolution path into a loyalty signal that CX leaders can track and improve across channels.
Most effort programs assume that a human handles every call. Automated resolutions now create a growing measurement blind spot. If a customer resolved an issue this morning without speaking to anyone, would your program register the time, repetition, and friction involved? Coverage across human and automated resolution paths becomes essential as service automation grows.
What is customer effort score (CES)?
Customer effort score (CES) is a post-interaction metric that measures how easy it was for a customer to resolve an issue with a company. Teams collect it with a single question; typically, "On a scale of 1 to 7, how strongly do you agree: [Company] made it easy for me to resolve my issue", and calculate the score as the share of respondents who agree the interaction was easy.
Dixon, Freeman, and Toman introduced CES in Harvard Business Review and showed it better predicts loyalty than customer satisfaction measures or the Net Promoter Score, drawing on more than 75,000 customer interactions. Their finding: exceeding expectations barely moves loyalty, while ease of resolution drives repurchase and spending, and high-effort experiences produce negative word of mouth. A 2025 peer-reviewed study of 359 CES responses confirmed that lower effort improved satisfaction, which then strengthened repurchase and recommendation intentions.
That predictive power comes with limits every CX leader should understand before the score reaches a board slide.
Where CES breaks down as a standalone metric
The metric has known blind spots, and each one narrows the claim a leader can make from the number. Design around these before you publish the score.
No expectation context: The standard effort question doesn't account for cases where more effort isn't necessarily worse, or how effort compares with what the customer expected.
No published benchmark: No universal industry CES standard exists, so a score is most meaningful against your own history and trend rather than as a direct comparison with others.
A 2010 research base: The foundational dataset dates to 2010, and a leader presenting the metric at board level should be able to explain the scope and limits of that research.
These constraints matter more as automation reshapes the contact center. A metric that already depends on trend and context loses meaning when it only surveys interactions humans handle. Applying CES to resolutions that no human touched is the next design problem.
Measuring effort when AI agents answer the call
Quality frameworks for human agents do not describe failures inside an automated resolution. Script adherence and coaching notes say nothing about an AI agent that answered confidently and wrongly or looped a caller through the same clarifying question three times. Four measurement rules make the score defensible across automated and human service paths.
Separate resolution paths: Survey AI-resolved interactions separately from human-resolved ones. Separate scores reveal whether the AI agent lowered effort or shifted it downstream.
Trigger surveys after resolution: Fire the survey when the operation resolves the issue. A customer who abandoned automation and called back the next day experienced substantial effort, regardless of what containment dashboards report.
Assign handoff effort to the leg that generated it: When a caller repeats their issue after a transfer, the handoff design produced that effort, not the receiving agent.
Audit survey suppression: Programs that survey only successful automated resolutions flatter the deployment. Record excluded interactions and who authorized each exclusion.
Effort visibility is only useful when it changes how the service team designs interactions. On the phone, high effort concentrates in the depth of the menu tree before a caller reaches anything useful, account details recited a second time after a transfer, the issue explained again to whoever received the escalation, and the hold music in between. Fixing those patterns is where voice AI earns its return.
Improving loyalty with voice AI agents
The voice channel concentrates customer effort at five predictable points: wait time, menu navigation, authentication, repeated explanations, and transfers. A well-designed AI voice agent can compress each one, and customers don't have to change anything to benefit.
Immediate answer, shorter waits
An AI agent picks up on the first ring and holds capacity through peak volume, so callers stop queuing to describe a problem they already know. In the orderbird customer story, orderbird cut customer wait time by 60%, from 98 to 39 seconds. Faster pickup removes the earliest, most measurable form of effort a caller encounters, and it makes the rest of the interaction feel shorter regardless of resolution time.
Natural language over menu trees
Callers state their need in plain language and skip the IVR entirely. Accurate intent recognition routes the call to the right resolution path without a menu tree or a probing series of "press 1 for" prompts. The customer describes the problem once, in their own words, and the AI agent either resolves it or hands it forward. Removing the menu removes the interaction most callers name when asked what makes support feel hard.
Authentication that persists
Identity verification runs once at the start of the conversation and holds through the rest, including any transfer to a human agent. The caller doesn't recite an account number a second time after a handoff, and the human receiving the call opens with verified context already loaded. Persistent authentication is a small design choice with an outsized effect on effort, because repeating security questions is one of the moments customers cite most often as frustrating.
Context that travels across handoffs
Escalation logic carries the verified identity and diagnosed intent forward, so a transfer becomes a continuation of the same conversation rather than a fresh start. Swiss Life became 60% faster at addressing customer concerns, with its AI agent routing callers at 96% accuracy. Accurate routing prevents the second queue and the second explanation, two of the highest-effort moments in any phone interaction.
Each of these design choices shows up in CES as lower effort scores on the interactions AI agents touch, and in operational data as fewer repeat contacts and shorter handle times. Pair the customer signal with the operational signal to know whether the interaction actually felt easier, not just faster.
Use customer effort score to expose loyalty risk
Loyalty risk lives in interaction effort, not the satisfaction number your dashboard reports. A survey program that covers only calls handled by human agents measures a shrinking share of that risk, and it hides the automated interactions where effort now most often accumulates. CES becomes a strategic instrument the moment it covers every resolution path: human, automated, and every handoff between them.
Parloa's AI Agent Management Platform follows the Build, Optimize, and Observe lifecycle, testing AI agents against real conversation complexity before launch and monitoring resolution, drop-off, and frustration signals afterward across 140+ languages. That observability closes the measurement gap CES was designed to expose and gives CX leaders a defensible view of effort across every interaction their contact center handles.
Book a demo to see how AI agents lower customer effort in your phone channel. Customers do not remember being surveyed; they remember whether getting help was easy.
Get in touch with our teamFAQs about customer effort score
What is a good customer effort score?
No truly universal, authoritative industry standard applies across all organizations, and vendors publish very different benchmark ranges depending on scale, industry, and calculation method. The more useful comparison is against your own prior quarters, because CES was designed as a trend instrument rather than a leaderboard. Segment the score by resolution path, channel, and issue type so a stable overall number doesn't hide a deteriorating automated queue or a specific intent that consistently forces callers to work harder. Manage the direction of travel, and treat any single quarter as one data point in a longer series.
How do teams calculate CES?
The most common convention is to sum positive responses, divide by total responses and multiply by 100, producing a percentage that reads as the share of customers who found the interaction easy. Some teams instead report a mean score on the 1-to-7 or 1-to-5 scale, and others convert responses into a net figure by subtracting negative responses from positive ones. Each method produces a different number from the same raw data, so calculation conventions vary across survey vendors. Choose one method, document it, and keep it consistent so trends remain comparable across quarters, teams, and resolution paths.
Is CES better than NPS or CSAT?
Each answers a different question, and a mature CX program uses all three rather than picking one. CES is an interaction-level ease signal that predicts repeat business, CSAT reflects in-the-moment satisfaction with a specific transaction, and NPS reflects brand advocacy and long-term relationship strength. CSAT tells you whether the customer felt taken care of, NPS tells you whether they would recommend you, and CES tells you whether they'll come back. A CX program needs the effort signal because ease of resolution tracks more directly with repeat business than either satisfaction or advocacy scores.
Should AI agent interactions get their own CES survey?
Yes. Blending AI-resolved and human-resolved interactions into a single score obscures which resolution path actually created effort and hides the failure modes unique to automation, such as clarification loops or confident wrong answers. Score interactions AI agents resolve separately from those human agents resolve, fire the survey on resolution rather than at the end of a session, and assign effort across escalation handoffs to the leg that generated it. Separate scoring lets you see whether automation lowered effort or simply shifted it downstream to the human agent who took the transfer.