CSAT survey questions: Question bank for post-call survey design
A monthly CSAT report that lands on a Tuesday. The contact center handled several hundred thousand calls last month, and an AI agent resolved many. Every caller who stayed on the line heard the same three questions. Survey governance should make the resulting score traceable to the sample, each question, and the handling path.
The score ticked up a point from the month before, and escalation complaints rose. By Friday, you have to decide whether the AI agent improved or the response pool changed. A governed survey makes that headcount decision defensible.
What is CSAT and how to calculate it?
CSAT, or Customer Satisfaction Score, is a post-interaction metric that measures how satisfied a customer was with a specific interaction. Contact centers typically collect it on a 1–5 scale immediately after a call ends, so the response ties to a single handling path rather than a general impression of the brand.
The reported figure is a top-two-box percentage. Count the responses that scored 4 or 5, divide by the total number of responses, then multiply by 100. If 800 of 1,000 respondents rated a queue 4 or 5, the CSAT for that queue is 80%; the remaining 200 responses, whether they scored 1, 2, or 3, do not enter the numerator.
An 85% or higher CSAT benchmark offers one point of comparison, though benchmarks vary by context. Improving CSAT starts by asking why a given queue sits below it. Deciding who responds and what they respond to requires governance, which is where survey program design begins.
Why post-call survey scores mislead at enterprise scale
A CSAT number from a self-selected sample can be an unreliable performance metric, especially at high call volumes. In a contact center running millions of calls a month, the respondents are not a random draw. They are whoever chose to stay on the line, which on an Interactive Voice Response (IVR) survey means opting in at the moment the caller most wants to hang up.
As fewer callers respond, the sample moves further from the population it describes. At enterprise scale, a persistent bias can move headcount plans and change which queues get investment. Without sample audits, those decisions rest on unchecked assumptions.
Surveying every caller fatigues frequent callers and inflates a self-selected sample that already skews toward extreme experiences. Programs that hold up under audit govern who receives the survey and enforce rules that keep month-over-month comparisons meaningful, so a shift in the score reflects service rather than sampling:
Contact-frequency quarantine: no caller receives a second survey within a set period, preventing heavy callers from dominating the sample and learning to ignore it.
Stratified sampling by intent and agent type: the program offers the survey in fixed proportions per intent and per handling path, so a queue of easy password resets cannot inflate the score for claims disputes.
Suppression logic: the program excludes callers with an open complaint or an active collections case. It also excludes recent survey recipients, so a single unresolved case does not produce three angry responses.
Governance fixes who answers. The wording and length of what they are asked are a separate design problem, and instruments can be under-specified in both areas.
10 post-call CSAT survey questions to ask
Each question below earns its place by tracing to one outcome or behavior you can act on, and a post-call survey has to stay short enough to hold the respondent's interest. Pick the items whose trace-back target matches the decision you need to make.
A voice-first survey lets the caller speak the answer, with keypad entry as a fallback for noisy environments, accessibility needs, or survey platforms outside the handling agent. The wording below assumes callers cannot see the scale anchors, so each item reads its options aloud and accepts either a spoken response or a matching keypress.
1. Was your issue resolved on this call?
A binary answer hides calls where the customer still needs a follow-up. Use the response options Yes / Partly / No, where "Partly" captures those calls. The agent reads the three options and accepts either the spoken word or a matching keypress. The resolution question ties to first call resolution rate, and it belongs ahead of the satisfaction item.
2. How satisfied were you with this call overall?
An overall score alone cannot explain what changed. The overall satisfaction question is the core CSAT item and feeds the top-two-box figure. Use a scale of 1 (very dissatisfied) to 5 (very satisfied), captured as a spoken number or a keypress. A rating of 3 offers no guidance about what to change. The surrounding items give the rating value by attaching a cause to it.
3. How easy was it to get your issue handled?
Satisfaction can hide avoidable effort. A caller can be satisfied with the outcome and still have repeated a policy number three times to reach it. This question measures effort on a scale of 1 (very difficult) to 5 (very easy), captured by voice or keypad.
4. Is this the first time you have contacted us about this issue?
A resolved answer can hide a previous attempt. A Yes / No response traces to repeat contact, the metric that first call resolution (FCR) self-reports miss because the caller who says "resolved" today may be the same caller from last week. A "No" here, joined to the caller's contact history, tells you whether the previous contact failed on resolution or on follow-through. Phone surveys accept the spoken answer or a keypress; email versions can add "How many times?"
5. Did you try another channel before calling?
Phone CSAT can absorb another channel's failure. This question traces to channel switching, using the response options Yes / No. A high "Yes" rate for a given intent means the app or web form failed before the phone call, and the phone CSAT you measure partly reflects cleaning up another channel's failure. Callers can say the answer or press a key. Email and in-app versions can ask which channel the caller tried first.
6. How would you describe how you felt at the end of the call?
A single satisfaction rating can hide whether disappointment became anger. Use this Customer Emotion scale: Perfect / Excellent / Good / Frustrating / Totally Unacceptable. "Frustrating" and "Totally Unacceptable" separate a mild disappointment from a caller who will churn or complain publicly. The emotion question traces to the customer's emotion. The agent reads the five labels aloud and accepts the spoken label; a numeric keypress remains available for callers in noisy conditions.
7. How knowledgeable and courteous was the human agent who helped you?
A combined rating can blur knowledge and courtesy. This rating maps to quality assurance (QA) rubric dimensions, where it earns its keep, since the QA form already scores the human agent's knowledge and courtesy from the recording. Use a 1-5 scale, captured by voice or keypad. Splitting the stem into two items when the channel allows gives cleaner data; a voice survey keeps the combined stem. Read the rating as a signal about the rubric dimension.
8. How natural did the conversation feel?
Human-handled calls do not have the speech-system failure modes this question measures, so asking it on those calls wastes a slot. Reserve this question for AI-handled calls and use a 1-5 scale. The naturalness question traces to speech and turn-taking quality: whether the agent interrupted or paused too long before answering. It also detects when the agent misheard a policy number spoken from a moving car, which is exactly the environment where keypad fallback earns its place alongside voice capture.
The naturalness score often moves first after a change to voice or speech settings.
9. When transferred to a human agent, how well was your request handed over?
A handoff score does not apply when no transfer occurred. Skip this item in that case. On AI-handled calls with a transfer, use a scale of 1 to 5, captured by voice or keypad.
The handoff question traces to escalation quality: whether the caller had to repeat what they had already told the AI agent, and whether the human agent picked up with the context already on screen. Survey teams need to wire the skip logic to the transfer flag on the call record, or every non-transferred caller will answer a question that doesn't apply.
10. What is one thing we could have done better?
A broad invitation can produce feedback that is difficult to categorize. Keep this open-ended item last. Record the response as voice on the phone and as free text elsewhere, using the same spoken-answer capture that carries the earlier items.
The question traces to verbatim analysis. The wording asks for one thing, which raises completion and produces a specific complaint you can categorize by intent.
Select items based on the trace-back target. A queue where the CRM already measures resolution does not need question 1 on the survey; a queue where callers arrive from a failing app needs question 5 more than question 3. The survey is a short instrument focused on the questions operational data cannot answer on its own.
Adapting the question bank for AI-handled calls
AI agents and human agents produce satisfaction through different mechanisms, so the survey record for an AI-handled call needs different items and metadata. Where a human agent's low score usually traces to manner or knowledge, an AI agent's traces to retrieval accuracy or a handoff that dropped context, and neither yields to coaching.
Stamp the AI agent version on every response
An AI agent's behavior can change on a single prompt update, so a three-point drop in question 2 cannot be traced to a Thursday release if every response carries only a month. Attach the prompt and configuration identifier in force at the time of the call, then compare scores before and after each change.
Run the survey on a neutral platform when neutrality matters
An AI agent surveying callers about its own performance raises questions no executive wants to answer, so many enterprises hand the post-call survey to a separate platform such as Genesys, NiCE, or Qualtrics. The voice-or-keypad capture pattern carries across both the handling agent and the detached stack.
Map each item to a behavior you can change
For AI-handled calls, there is no individual to coach, so the useful unit is the mapping from item to behavior. Question 7 maps to a QA rubric dimension, question 8 to response quality, and question 9 to escalation logic. That mapping is what makes AI observability actionable on the resulting scores.
Stamp the metadata that joins scores to conversations
Without metadata stamped at call end, you can't trace a low rating to its prompt version, handling path, or the exact conversation behind it. Attach an agent type flag, a containment or escalation outcome, the AI agent version, and a call recording or transcript ID to every survey response.
Close the loop with joined evidence
Württembergische Versicherung results show 3.8 out of 5 CSAT on its AI agent, alongside a 33% reduction in call wait times within 4 weeks. Joined to a version and a transcript, that satisfaction score becomes the input for the next change to the agent.
Turn CSAT survey questions into a governed measurement loop
A governed CSAT program turns the support channel into a source of decisions, not reassurance. When each score joins to a resolution flag, an agent type, and an AI agent version, sales and support leaders can see whether a queue is losing customers to friction, whether an AI update helped, or whether the response pool shifted underneath the number. That traceability is what lets a Tuesday report support a Friday decision without confusing a changing sample with better service.
Parloa's AI Agent Management Platform covers the Build, Optimize, and Observe lifecycle in 140+ languages, feeding satisfaction signals from AI-handled calls back into agent development and monitoring. Its compliance coverage includes ISO 27001:2022, ISO 17422:2020, SOC 2 Type I & II, PCI DSS, HIPAA, the General Data Protection Regulation (GDPR), and DORA.
Book a demo to measure and improve CSAT on AI-handled calls.
Get in touch with our teamFAQs about CSAT survey questions
Delivery and reporting choices determine whether post-call scores remain comparable across queues, channels, and handling paths.
How many questions should a post-call phone survey include?
Long phone surveys lose callers before they finish, so cap a voice survey at two or three items chosen by the trace-back target that matters for that queue. Reserve longer instruments for email or in-app follow-ups, where the respondent can see the whole survey instead of holding a phone to their ear.
Should CSAT for AI agents be reported separately from human agents?
A blended average hides the mechanism behind any movement, and a rise in one call type can mask a fall in the other. Report AI-handled and human-handled calls as separate figures because prompt and model changes affect AI-handled scores, whereas staffing and coaching affect human-handled scores.
Do I need consent to send a post-call text-message survey?
Auto-dialed text messages create consent obligations because the Telephone Consumer Protection Act (TCPA) requires prior express written consent. The Federal Communications Commission vacated its one-to-one consent rule in January 2025, so document opt-in for every number you text. That record supports the consent requirement, while GDPR requires explicit consent for call recordings used alongside the survey.
Which is better after a call: CSAT, customer effort score (CES), or Net Promoter Score (NPS)?
Using one metric for every post-call decision obscures whether the result reflects the interaction, its friction, or the wider relationship. Use CSAT for the interaction, CES for friction, and NPS for the relationship; an 11-point NPS scale remains a poor fit for a spoken or single-keypress answer on a phone.
Should I use a 1–5 or 1–7 scale on an IVR survey?
Longer scales are harder to hold in mind when callers cannot see the anchors, whether they speak the number or press it. Use a five-point scale on phone surveys to match the reporting convention in existing dashboards, and reserve longer scales for email or in-app surveys where respondents can see each option.
:format(webp))