← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Standard deviation you expect in the variable, taken from a previous study or a pilot. If you only know the range, a rough approximation is range/4.

Half-width of the confidence interval you want, in the units of the variable.

Number of subjects in the population you will sample from, if it is finite and known. Leave it empty (or 0) otherwise.

Proportion of subjects you expect to lose (dropout, loss to follow-up, unusable samples). Leave it empty if you expect none.

Inverse mode: if you already have a given number of subjects, enter it here to see the precision you would reach. Leave it empty (or 0) if you are not using it.

Example loaded

Illustrative example: estimating the mean platelet count (×10³/µL) in patients with dengue, assuming a standard deviation of 60, within ±10 and anticipating 10% losses (fictitious data).

Illustrative data, not real.

Results

Normal quantile (z)

1.960

95%

Size with the normal formula (n_z)

139

exact value: 138.29

Size with Student's t (n_t)

141

half-width achieved: 9.99

Required size (n)

139

exact value: 138.29

Recruitment with losses

154

exact value: 153.66

Precision reachable with the available n

not defined

enter an available sample size to compute it

Interpretation

With 139 subjects, and if the true standard deviation is 60, the mean would be estimated with a 95% confidence interval of about ±10 units.

Using Student's t quantile instead of the normal one, the requirement rises to 141 subjects: 2 more than with the normal formula, because estimating the standard deviation as well costs precision.

With no bounded source population, the size stays at 139 subjects. If you were sampling from a finite population not much larger than that figure, entering it would reduce the number needed.

With 10% expected losses, you need to recruit 154 subjects so that 139 finish.

If you already have a given number of subjects, enter it under "Sample size already available (n, optional)" to see how precisely you could estimate the mean.

  • The figures depend on the assumptions (expected standard deviation, precision) taken from previous studies or pilots; document where they come from.
  • The formula assumes an approximately symmetric variable. If yours is markedly skewed, consider transforming it (logarithm) or describing it with the median rather than the mean.
Reachable precision by sample sizeReachable precision con 139 Subjects (n): ±9.975; target precision ±10.Half-width of the CIReachable precision0102050100150200250Subjects (n)target precision ±10computed n: 139
Reachable precision by sample size

Explanation

To estimate a mean (platelets, age, length of stay) the sample size is computed so that the confidence interval is narrow enough. The absolute precision d is the half-width of that interval, in the units of the variable: with d = 10 the mean will be estimated within ±10 units.

The formula needs the expected standard deviation σ, taken from previous studies, from a pilot or, at worst, from an approximation such as the range divided by four. The size grows with the square of σ and falls with the square of d: doubling the precision demanded quadruples the sample, and getting σ wrong by 20% changes the size by almost 50%.

The t variant recognises that, when the standard deviation is also estimated, the right quantile is not the normal one but Student's t with n−1 degrees of freedom. Since that quantile depends on n itself, there is no closed form: the calculator searches for the smallest whole size that meets the condition, starting from the normal result and stepping up one at a time. The difference is negligible in large samples and becomes decisive below 30 subjects, where the normal formula falls short.

Three adjustments close the calculation. If you sample from a finite, known population, the finite population correction reduces the size. Expected losses L raise recruitment to n/(1−L). And the inverse mode answers the opposite question: with the subjects I already have, how precisely can I estimate the mean? Every figure is rounded up once, at the end. If the variable is markedly skewed, transform it or describe it with the median instead.

Equations

nz=(z1−α/2 σd)2n_z=\left(\dfrac{z_{1-\alpha/2}\,\sigma}{d}\right)^{2}
σ\sigma
expected standard deviation of the variable
dd
absolute precision: half-width of the confidence interval
z1−α/2z_{1-\alpha/2}
normal quantile for the confidence level (1.96 at 95%)
Cochran's (1977) normal formula for estimating a mean with absolute precision.
nt=min⁡{ n∈Z, n≥2 : n≥(tn−1,  1−α/2 σd)2}n_t=\min\left\{\,n\in\mathbb{Z},\ n\ge 2\ :\ n\ge\left(\dfrac{t_{n-1,\;1-\alpha/2}\,\sigma}{d}\right)^{2}\right\}
tν,  1−α/2t_{\nu,\;1-\alpha/2}
quantile of Student's t with ν degrees of freedom
The t variant (Student 1908): the smallest whole size that meets the condition, searched one by one from ⌈n_z⌉. Not solved by fixed point, which can land below the minimum.
n=nz1+nz−1Nn=\dfrac{n_z}{1+\dfrac{n_z-1}{N}}
NN
size of the finite population being sampled
Finite population correction; applied to the unrounded n_z. With no N declared, n = n_z.
naj=⌈n1−L⌉n_{aj}=\left\lceil\dfrac{n}{1-L}\right\rceil
LL
expected proportion of losses (0 to 0.5)
Adjustment for expected losses (Lwanga and Lemeshow 1991): recruitment rises so that n subjects finish. The ceiling is applied once, at the end; the unrounded value is shown in the cell's detail.
d(n)=z1−α/2 σnz(n),nz(n)=n (N−1)N−nd(n)=\dfrac{z_{1-\alpha/2}\,\sigma}{\sqrt{n_z(n)}},\qquad n_z(n)=\dfrac{n\,(N-1)}{N-n}
nz(n)n_z(n)
uncorrected size equivalent to n in a population of N
Inverse mode: the same n(d) relationship solved for d. With no N declared, n_z(n) = n.

R code

# Sample size to estimate one mean (absolute precision) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

sigma <- 60; d <- 10; nivel <- 0.95
poblacion <- 0   # size of the finite population; 0 = population not bounded
perdidas <- 0.1     # expected losses to follow-up, 0 to 0.5
n_dado <- 0         # inverse mode: sample size already available; 0 = not used

z <- qnorm(1 - (1 - nivel) / 2)
n_z <- (z * sigma / d)^2                                        # Cochran 1977
# t variant (Student 1908): the SMALLEST integer n >= 2 with n >= (t_{n-1} * sigma / d)^2.
# Searched one by one from max(2, ceiling(n_z)), a lower bound because t > z; the
# condition is monotone, so the first n that meets it is the minimum.
# NOT a fixed point: iterating n <- (t_{ceiling(n)-1} * sigma / d)^2 falls into a
# period-2 cycle and can publish a size BELOW the minimum (sigma = 1, d = 0.36,
# nivel = 0.80 gives 14, but 14 subjects reach only 0.3608, not 0.36).
n_t <- NA_real_
n_i <- max(2, ceiling(n_z))
for (i in 1:64) {
  cumple <- n_i >= (qt(1 - (1 - nivel) / 2, max(1, n_i - 1)) * sigma / d)^2
  if (cumple || n_i + 1 == n_i) { n_t <- n_i; break }   # above 2^53, +1 leaves the double unchanged
  n_i <- n_i + 1
}
n <- if (poblacion >= 2) n_z / (1 + (n_z - 1) / poblacion) else n_z   # finite population correction
n_ajustado <- n / (1 - perdidas)                                # Lwanga & Lemeshow 1991
# Inverse mode: the very same n(d) solved for d, with the same correction
n_efectivo <- if (poblacion >= 2) n_dado * (poblacion - 1) / (poblacion - n_dado) else n_dado
d_dado <- if (n_dado >= 2) z * sigma / sqrt(n_efectivo) else NA_real_

res <- list(z = z, n_z = n_z, n_t = n_t, n = n, n_ajustado = n_ajustado, d_dado = d_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# n_t is already an integer; the interface applies the ceiling once to the rest:
# ceiling(n_z), ceiling(n), ceiling(n_ajustado).
# Equivalent in RStudio (not run in the browser):
# presize::prec_mean(mean = 0, sd = sigma, conf.width = 2 * d, conf.level = nivel)
# epiR::epi.sssimpleestc(N = poblacion, xbar = 0, sigma = sigma, epsilon = d, error = "absolute", conf.level = nivel)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Sample size was computed to estimate a mean with absolute precision, assuming a standard deviation of 60 and an interval half-width of 10, with a confidence level of 95%, using Cochran's normal formula [1], without finite population correction, with the iterative variant based on Student's t [2] as a check (141 subjects); 10% was added for expected losses [3], giving a recruitment target of 154 subjects. Calculations used the "Sample size for one mean" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-media), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Cochran WG. Sampling Techniques. 3rd ed. New York: John Wiley & Sons; 1977. Original source
  2. 02 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
  3. 03 Lwanga SK, Lemeshow S. Sample Size Determination in Health Studies: A Practical Manual. Geneva: World Health Organization; 1991. Complementary
  4. 04 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
  5. 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading