Tools · Open Biostatistics
Sample size to estimate one mean, with t variant and finite population correction
Enter the standard deviation you expect and how precisely you want to estimate the mean; get the sample size from the normal formula and from the t variant, the effect of a finite population, the recruitment needed if you expect losses, and the precision you would reach with a sample size you already have.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-media
This link does not include the pasted data: they are too long for a URL.
Results
Normal quantile (z)
1.960
95%
Size with the normal formula (n_z)
139
exact value: 138.29
Size with Student's t (n_t)
141
half-width achieved: 9.99
Required size (n)
139
exact value: 138.29
Recruitment with losses
154
exact value: 153.66
Precision reachable with the available n
not defined
enter an available sample size to compute it
Interpretation
With 139 subjects, and if the true standard deviation is 60, the mean would be estimated with a 95% confidence interval of about ±10 units.
Using Student's t quantile instead of the normal one, the requirement rises to 141 subjects: 2 more than with the normal formula, because estimating the standard deviation as well costs precision.
With no bounded source population, the size stays at 139 subjects. If you were sampling from a finite population not much larger than that figure, entering it would reduce the number needed.
With 10% expected losses, you need to recruit 154 subjects so that 139 finish.
If you already have a given number of subjects, enter it under "Sample size already available (n, optional)" to see how precisely you could estimate the mean.
- The figures depend on the assumptions (expected standard deviation, precision) taken from previous studies or pilots; document where they come from.
- The formula assumes an approximately symmetric variable. If yours is markedly skewed, consider transforming it (logarithm) or describing it with the median rather than the mean.
Explanation
To estimate a mean (platelets, age, length of stay) the sample size is computed so that the confidence interval is narrow enough. The absolute precision d is the half-width of that interval, in the units of the variable: with d = 10 the mean will be estimated within ±10 units.
The formula needs the expected standard deviation σ, taken from previous studies, from a pilot or, at worst, from an approximation such as the range divided by four. The size grows with the square of σ and falls with the square of d: doubling the precision demanded quadruples the sample, and getting σ wrong by 20% changes the size by almost 50%.
The t variant recognises that, when the standard deviation is also estimated, the right quantile is not the normal one but Student's t with n−1 degrees of freedom. Since that quantile depends on n itself, there is no closed form: the calculator searches for the smallest whole size that meets the condition, starting from the normal result and stepping up one at a time. The difference is negligible in large samples and becomes decisive below 30 subjects, where the normal formula falls short.
Three adjustments close the calculation. If you sample from a finite, known population, the finite population correction reduces the size. Expected losses L raise recruitment to n/(1−L). And the inverse mode answers the opposite question: with the subjects I already have, how precisely can I estimate the mean? Every figure is rounded up once, at the end. If the variable is markedly skewed, transform it or describe it with the median instead.
Equations
- expected standard deviation of the variable
- absolute precision: half-width of the confidence interval
- normal quantile for the confidence level (1.96 at 95%)
- quantile of Student's t with ν degrees of freedom
- size of the finite population being sampled
- expected proportion of losses (0 to 0.5)
- uncorrected size equivalent to n in a population of N
R code
# Sample size to estimate one mean (absolute precision) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
sigma <- 60; d <- 10; nivel <- 0.95
poblacion <- 0 # size of the finite population; 0 = population not bounded
perdidas <- 0.1 # expected losses to follow-up, 0 to 0.5
n_dado <- 0 # inverse mode: sample size already available; 0 = not used
z <- qnorm(1 - (1 - nivel) / 2)
n_z <- (z * sigma / d)^2 # Cochran 1977
# t variant (Student 1908): the SMALLEST integer n >= 2 with n >= (t_{n-1} * sigma / d)^2.
# Searched one by one from max(2, ceiling(n_z)), a lower bound because t > z; the
# condition is monotone, so the first n that meets it is the minimum.
# NOT a fixed point: iterating n <- (t_{ceiling(n)-1} * sigma / d)^2 falls into a
# period-2 cycle and can publish a size BELOW the minimum (sigma = 1, d = 0.36,
# nivel = 0.80 gives 14, but 14 subjects reach only 0.3608, not 0.36).
n_t <- NA_real_
n_i <- max(2, ceiling(n_z))
for (i in 1:64) {
cumple <- n_i >= (qt(1 - (1 - nivel) / 2, max(1, n_i - 1)) * sigma / d)^2
if (cumple || n_i + 1 == n_i) { n_t <- n_i; break } # above 2^53, +1 leaves the double unchanged
n_i <- n_i + 1
}
n <- if (poblacion >= 2) n_z / (1 + (n_z - 1) / poblacion) else n_z # finite population correction
n_ajustado <- n / (1 - perdidas) # Lwanga & Lemeshow 1991
# Inverse mode: the very same n(d) solved for d, with the same correction
n_efectivo <- if (poblacion >= 2) n_dado * (poblacion - 1) / (poblacion - n_dado) else n_dado
d_dado <- if (n_dado >= 2) z * sigma / sqrt(n_efectivo) else NA_real_
res <- list(z = z, n_z = n_z, n_t = n_t, n = n, n_ajustado = n_ajustado, d_dado = d_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# n_t is already an integer; the interface applies the ceiling once to the rest:
# ceiling(n_z), ceiling(n), ceiling(n_ajustado).
# Equivalent in RStudio (not run in the browser):
# presize::prec_mean(mean = 0, sd = sigma, conf.width = 2 * d, conf.level = nivel)
# epiR::epi.sssimpleestc(N = poblacion, xbar = 0, sigma = sigma, epsilon = d, error = "absolute", conf.level = nivel)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Sample size was computed to estimate a mean with absolute precision, assuming a standard deviation of 60 and an interval half-width of 10, with a confidence level of 95%, using Cochran's normal formula [1], without finite population correction, with the iterative variant based on Student's t [2] as a check (141 subjects); 10% was added for expected losses [3], giving a recruitment target of 154 subjects. Calculations used the "Sample size for one mean" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-media), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Cochran WG. Sampling Techniques. 3rd ed. New York: John Wiley & Sons; 1977. Original source
- 02 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
- 03 Lwanga SK, Lemeshow S. Sample Size Determination in Health Studies: A Practical Manual. Geneva: World Health Organization; 1991. Complementary
- 04 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
- 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading