Skip to content
Cati.botResearch at the speed of sound.
Cati.bot
Research notes

Response rates

Telephone survey response rates and the limits of adjustment

Pew Research Center reported a telephone survey response rate of 36% in 1997 and 6% in 2018. The decline is substantial, but the figures belong to Pew’s US telephone polling programme and its fieldwork protocols. They are neither a universal history of telephone research nor a current benchmark for every country, population and study design.

They also leave the most consequential question unanswered: how much error did nonresponse introduce into a particular estimate? A response rate describes the extent of participation under a stated definition. Assessing the consequences requires information about the people who did and did not participate, and about their relationship to the subject being measured.

Count cases before counting calls

A sampled case may receive several call attempts. Treating those attempts as separate opportunities in a response-rate denominator confuses participation with fieldwork productivity. In a hypothetical study with 1,000 known eligible cases, 100 complete interviews and no partial interviews, the case response rate is 10%. If the team makes 4,000 calls to those cases, it obtains 2.5 completes per 100 dials. Both quantities are useful, but they answer different questions.

Real samples are less tidy because eligibility is often unknown for some cases. AAPOR’s Standard Definitions, tenth edition, distinguishes several response rates. RR1 counts complete interviews in the numerator and includes all cases of unknown eligibility in the denominator. RR3 instead includes an estimated eligible share of those unknown cases. RR4 uses that eligibility treatment while also counting qualifying partial interviews in the numerator.

A methods report should identify the chosen definition, supply the final case dispositions and explain any eligibility estimate. Contact and cooperation rates help locate losses at different stages of recruitment. Call attempts, interviewer hours and cost per complete belong alongside them. Without those details, an apparently favourable response rate may reflect a different denominator, a more familiar sample or a more intensive protocol.

What call screening tells us

There is direct evidence that unknown callers face a difficult reception. In a Pew survey conducted in July 2020, 19% of US adults said they generally answered cellphone calls from unknown numbers. Most said they would leave such calls unanswered, although many would check a voicemail. This was self-reported behaviour in a particular period, rather than an experiment measuring which interventions increased survey participation.

It is reasonable to investigate call screening when planning recruitment. It is harder to assign a precise share of a decades-long response-rate decline to caller identification, unwanted calls, changes in telephone ownership or attitudes towards research. A historical trend alone cannot separate those mechanisms. Nor can it establish that a decline is irreversible.

For a new study, the practical task is to learn where contact is being lost. Examine attempts by day and time, available delivery outcomes, callbacks and the point at which people leave the introduction. Distinguish failure to reach anyone from a decision to decline the study. If an advance notice is feasible, test it against the existing protocol using comparable sampled cases. An improvement in answer rates is useful; whether it changes the composition of the achieved sample is a separate empirical question.

The same response rate can conceal different errors

For a simple illustration, set sampling variation aside. Suppose 10% of a population responds. An attribute is present among 30% of respondents and 40% of nonrespondents. The population proportion is 39%: one tenth multiplied by 30%, plus nine tenths multiplied by 40%. The unadjusted respondent estimate of 30% understates it by nine percentage points.

If the attribute were present among 30% of both groups, the same 10% response rate would produce no nonresponse error for that attribute. Low participation leaves considerable scope for error, but its size depends on differences relevant to the outcome. A sample can therefore estimate one quantity reasonably well and another poorly.

The methodological literature has long examined this distinction. Groves and Peytcheva’s 2008 meta-analysis assembled evidence from 59 studies of nonresponse bias, examining how its relationship with nonresponse rates varies with the survey and the statistic being estimated. This is why a single participation percentage is an inadequate substitute for investigating the estimates a study is intended to produce.

Pew’s 2017 benchmarking study provides a concrete example. Across 13 demographic, lifestyle and health measures compared with federal survey benchmarks, the average absolute difference was 2.7 percentage points in 2016, compared with 2.8 points in 2012. Yet its wider analysis found substantial overrepresentation of civic engagement, including volunteering. Agreement on some outcomes coexisted with sizeable discrepancies on others.

Benchmarking also has limits. Differences in question wording, reference periods, population coverage and interview mode can produce discrepancies even when nonresponse is similar. A benchmark survey has errors of its own. The comparison is evidence about specified measures under specified conditions; it cannot certify an unrelated outcome for which no benchmark exists.

What weighting asks us to assume

Calibration can make a sample reproduce known population totals for characteristics such as age, education or region. Matching those totals is a check that the adjustment worked as specified. It does not, by itself, show that the adjusted outcome estimates are unbiased.

Consider a customer study in which response is especially likely among people with an unresolved complaint. Weighting by age and region will not necessarily remove that selection if it persists within each age–region group. A reliable administrative measure of complaint history might support a better adjustment, provided it is available for the relevant population and measured consistently. The useful auxiliary information is information that helps explain participation and the outcome, rather than simply offering more demographic categories.

Inspect how much influence the adjustment gives individual cases and small groups. A few respondents carrying very large weights can leave an estimate unstable. Trimming weights may improve precision while reintroducing some imbalance, so document the rule and examine sensitivity to reasonable alternatives. Use variance estimation that reflects the sample design and weighting. A conventional sampling margin of error does not quantify residual nonresponse bias.

Make the fieldwork decision outcome-specific

A useful fieldwork review brings three records together: the disposition history, the available characteristics of respondents and nonrespondents, and the study’s key estimates. Where frame data permit, examine participation across relevant groups. Compare achieved composition before and after adjustment, and use external benchmarks that match the concepts and population closely enough to be informative.

Additional callbacks may be valuable when they bring in people whose experiences are otherwise missing. They may add less information when they mainly recruit more people resembling existing respondents. Evaluate the marginal interviews as well as their number. Late respondents can help diagnose fieldwork differences, but assuming they represent everyone who never responds requires justification.

Mixed-mode follow-up deserves the same scrutiny. It may reach people missed by the initial approach, while also changing how questions are understood or answered. Preserve mode and recruitment information so those effects can be investigated. If the proposed change is AI voice interviewing, test its effects on participation and measurement alongside its cost. Cheaper administration creates room to reconsider a design; it does not establish that the resulting sample is representative.

A defensible account of a low-response survey therefore explains which outcomes were assessed, what evidence supports the adjustment, and where uncertainty remains. Reporting the rate accurately is the beginning of that account.

Designing telephone fieldwork?

Share your target population, sample source and fieldwork constraints with Voxworks to discuss a possible Cati.bot pilot.

Get in touch