Sample Size Calculator
How many to ask or measure — for a survey, a mean or a comparison of two groups.
Sample size
This is the last answer worked out. Fix the input above to update it.
Steps
Other margins
About the Sample Size Calculator
Work out how many people to survey — or how many values to measure — before you start, with every step of the formula shown. For a survey or prevalence study the calculator uses Cochran’s formula with the finite population correction, and can add a design effect for cluster sampling and allow for people who do not respond. It also finds the sample size to estimate a mean to a chosen precision (with z, or with t when the standard deviation will be estimated from the sample).
For studies that compare two groups — two proportions or two means — it finds the size of each group needed to detect a chosen difference with a chosen power, including unequal groups, paired designs and the exact t-test answer next to the textbook approximation. And for a sample you already have, it gives the margin of error, with the Wilson interval for proportions.
How to use it
- Choose what you are planning: a survey (a proportion), a mean, a comparison of two proportions or two means, or the margin of error of a sample you have.
- Pick the confidence level (95% is usual), and for comparisons the power (80% or 90%) and one- or two-sided testing.
- Fill in the rest: the margin of error and expected proportion (keep 50% if unsure), or the standard deviation and margin, or the two proportions, or the difference and standard deviation (or Cohen’s d).
- Add the population size if it is small, and optionally a design effect and the response rate you expect.
- Read the sample size, the steps and — for surveys — the table of sizes for other margins and confidence levels. Copy the result for your protocol or report.
Examples
Margin 5 points, p = 50%, 95% confidence
n₀ = 1.96² × 0.5 × 0.5 ÷ 0.05² = 384.15 → 385 Population 1,000: 384.15 ÷ (1 + 383.15 ÷ 1,000) = 277.74 → 278 · population 10,000: 370
As above, design effect 2, 80% expected response
384.15 × 2 ÷ 0.8 = 960.37 → invite 961
σ = 15, within ±3 at 95% confidence
(1.96 × 15 ÷ 3)² = 96.04 → 97 — or 99 with t, because σ will be estimated from the sample
50% vs 65%, α = 0.05 two-sided, 80% power
170 in each group, 340 in total
Effect size d = 0.5 (medium), α = 0.05 two-sided, 80% power
64 in each group (exact t-test; the normal approximation gives 63)
Common uses
- Planning a questionnaire, a customer survey or a poll, and checking whether a target margin of error is affordable.
- Sample size sections of theses, dissertations and research protocols, with the formula and the numbers that went into it.
- Prevalence studies in public health, with a design effect for cluster sampling.
- Reporting the margin of error of a survey that is already done.
The formulas
- Proportion (Cochran): n₀ = z² × p(1 − p) ÷ e², with z for the confidence level (1.96 for 95%), p the expected proportion and e the margin of error.
- Finite population correction: n = n₀ ÷ (1 + (n₀ − 1) ÷ N). Then × design effect, and ÷ the expected response rate.
- Mean: n₀ = (z × σ ÷ E)²; with t instead of z the formula is repeated with t for n − 1 degrees of freedom until n stops changing, as the NIST/SEMATECH e-Handbook of Statistical Methods describes in §7.2.2.2.
- Two proportions: n₁ = [z_α √(p̄(1 − p̄)(1 + 1/k)) + z_β √(p₁(1 − p₁) + p₂(1 − p₂)/k)]² ÷ (p₁ − p₂)², n₂ = k × n₁, with p̄ = (p₁ + k p₂) ÷ (1 + k) (Fleiss, Levin & Paik, Statistical Methods for Rates and Proportions, 3rd ed.); the optional continuity correction is that of Fleiss, Tytun & Ury.
- Two means: n₁ = (z_α + z_β)² × (1 + 1/k) ÷ d², d = Δ ÷ σ — and the exact answer from the noncentral t distribution, which R’s
power.t.testalso gives.
Cochran’s formula and the correction are those of W. G. Cochran, Sampling Techniques (3rd ed., Wiley); the effect sizes d and h and their conventions are from J. Cohen, Statistical Power Analysis for the Behavioral Sciences (2nd ed.).
Choosing the inputs
The margin of error is the ± you can accept around a result: ±5 percentage points means a measured 40% could really be anything from 35% to 45% (at the chosen confidence). Halving the margin quadruples the sample.
The expected proportion matters because p(1 − p) is largest at 50%: if you have no estimate, 50% is the safe choice. If earlier studies found about 20%, using 20% needs fewer people (246 instead of 385 for ±5% at 95%). With a relative margin (say ±20% of the proportion), rare outcomes need much larger samples.
The population size only matters when the sample is a sizeable share of it: 278 for 1,000 people, 370 for 10,000, 377 for 20,000 and 383 for 100,000. A design effect above 1 allows for cluster sampling (whole schools or villages instead of individuals), which gives less information per person; values of 1.5 to 2 are common, but take yours from similar studies.
Power and effect size
When comparing groups, α (1 − confidence) is the chance of a false alarm — finding a difference that is not there — and the power is the chance of finding a difference of the size you entered if it is real. 80% power is the usual minimum, 90% common for important studies. Cohen’s conventions call d = 0.2 a small, 0.5 a medium and 0.8 a large difference between means (and the same values of h for proportions); a smaller effect needs many more participants — 64 per group for d = 0.5 but about 394 for d = 0.2.
What a sample size cannot fix
These formulas assume a random sample: everyone in the population had a known chance of being picked. A large sample from a self-selected online poll or a convenient group can still be badly biased, and the margin of error says nothing about that — nor about non-response, leading questions or measurement errors.
Limitations
- Simple random sampling is assumed; for cluster or stratified designs use a design effect taken from similar studies, or a statistician’s calculation.
- The proportion formulas use the normal approximation; with expected proportions close to 0% or 100% and small samples, exact methods give more reliable answers.
- The comparisons assume independent groups (or simple pairs) and, for means, equal standard deviations; they do not cover survival, repeated-measures, cluster-randomised or sequential designs.
- The result is the number of completed responses or measurements you need, rounded up; dropouts beyond the response rate you enter are not included.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
How many people do I need to survey?
For a margin of error of ±5 percentage points at 95% confidence, 385 people — whatever the size of a large population. Fewer are needed for a small population (278 for a population of 1,000) and more for a tighter margin (1,068 for ±3 points).
Why is 50% used for the expected proportion?
Because p × (1 − p) — the part of the formula that depends on the answer — is largest at p = 50%, so 50% gives the largest, safest sample size. If you know the answer is around 10% or 90%, using that value gives a smaller sample with the same margin of error.
Does the size of the population matter?
Only when the sample is a noticeable share of it. The finite population correction n = n₀ ÷ (1 + (n₀ − 1) ÷ N) shrinks 385 to 278 for 1,000 people, but barely changes it for 100,000 (383). Beyond that, the population size makes almost no difference.
What is the margin of error of my survey?
For a proportion p̂ from n people it is z × √(p̂(1 − p̂) ÷ n): with 1,000 people and a result of 50% at 95% confidence, 1.96 × √(0.25 ÷ 1,000) ≈ ±3.1 percentage points. Choose Margin of error to work it out, with the Wilson interval, which is more accurate for small samples and extreme results.
What power should I choose?
80% is the conventional minimum: a real difference of the size you entered is missed one time in five. 90% is often required for clinical trials and other studies where missing a real effect is costly; it needs about a third more participants.
Why are there two answers for comparing two means?
The textbook formula uses the normal distribution, as if the standard deviation were known. A real study estimates it from the data and uses a t-test, which needs slightly more participants — usually one or two per group. The calculator shows both and uses the exact t-test value, computed from the noncentral t distribution.