← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Arithmetic mean of the variable, as reported in the paper or in the software output.

Sample standard deviation (the one with n − 1 in the denominator). If the paper reports the standard error, multiply it by the square root of n.

Number of observations with valid data (integer of 2 or more).

Example loaded

Illustrative example: platelet count in 50 febrile patients, mean 95 ×10³/µL and standard deviation 40 ×10³/µL (fictitious data).

Illustrative data, not real.

Results

Mean

95.00

83.63 to 106.37

95% CI · Student's t, n − 1 df

Standard error of the mean (SEM)

5.66

SD divided by the square root of n

Standard deviation (SD)

40.00

33.41 to 49.85

95% CI · chi-squared distribution

Student's t multiplier

2.010

Degrees of freedom: 49

Interpretation

The mean was 95.00 (SD 40.00) in 50 observations; the 95% confidence interval (83.63 to 106.37) gives the range of population means compatible with the data.

That interval does not describe individual spread: for that, use mean ± 2 SD, which here covers roughly 15.00 to 175.00.

The population standard deviation is compatible with values from 33.41 to 49.85; with small samples that interval is far wider than usually assumed.

With n = 50 the multiplier is t = 2.010 rather than the 1.96 of the normal distribution, and the standard error of the mean is 5.66.

With n = 50 (30 or more) the central limit theorem makes the interval for the mean robust even if the data are not normal; the interval for the standard deviation, by contrast, does depend on that assumption.

The mean with its interval against the spread of the dataMean: 95.00 (83.63 to 106.37); Mean ± 2 SD: 15.00 to 175.00.MeanMean ± 1 SDMean ± 2 SD50.0100.0150.0
The mean with its interval against the spread of the data

Explanation

A confidence interval for a mean answers one specific question: which population means are compatible with what I observed? It says nothing about where individual patients fall. Confusing the two is the most common mistake when reading a paper: a narrow interval alongside a large standard deviation means the mean is well estimated, not that the values resemble each other.

The multiplier is not 1.96 but the quantile of Student's t distribution with n − 1 degrees of freedom, because the standard deviation was estimated from the same data. With small samples that multiplier grows noticeably (with n = 5 it is 2.78) and the interval widens. Convergence towards 1.96 is slow: with n = 30 the multiplier is still 2.045, with n = 100 it is 1.984, and about 240 observations are needed for the difference to fall below 0.01.

The standard deviation is also an estimate and also carries uncertainty. Its interval comes from the chi-squared distribution of (n − 1)s²/σ², is markedly asymmetric and, with small n, surprisingly wide: a good way to see why standard deviations from small studies should not be compared as if they were exact figures.

Both intervals assume the data come from an approximately normal distribution. The one for the mean is robust thanks to the central limit theorem as soon as the sample grows; the one for the standard deviation is not, and with strongly skewed distributions it is better to report the median and the interquartile range instead of the mean and the standard deviation.

Equations

xˉ±tn−1, 1−α/2 sn\bar x \pm t_{n-1,\,1-\alpha/2}\,\frac{s}{\sqrt{n}}
xˉ\bar x
observed mean
ss
sample standard deviation (n − 1 denominator)
nn
number of observations
tn−1, 1−α/2t_{n-1,\,1-\alpha/2}
two-sided quantile of Student's t distribution with n − 1 degrees of freedom
Confidence interval for the mean (Student 1908). It is the same interval t.test() returns from the raw data.
(n−1) s2σ2∼χn−12[ sn−1χn−1, 1−α/22,  sn−1χn−1, α/22 ]\begin{gathered} \frac{(n-1)\,s^2}{\sigma^2} \sim \chi^2_{n-1} \\ \left[\ s\sqrt{\frac{n-1}{\chi^2_{n-1,\,1-\alpha/2}}},\ \ s\sqrt{\frac{n-1}{\chi^2_{n-1,\,\alpha/2}}}\ \right] \end{gathered}
χn−1, q2\chi^2_{n-1,\,q}
quantile q of the chi-squared distribution with n − 1 degrees of freedom
σ\sigma
population standard deviation, the quantity being estimated
Equal-tailed interval for σ, derived from the fact that (n − 1)s²/σ² follows a chi-squared distribution with n − 1 degrees of freedom. It is asymmetric: the upper limit sits farther from s than the lower one.
SE⁡(xˉ)=sn\se(\bar x) = \frac{s}{\sqrt{n}}
SE⁡(xˉ)\se(\bar x)
standard error of the mean: how much the mean varies from sample to sample
The standard error describes the precision of the mean; the standard deviation describes the spread of the data. Only the former shrinks as n grows.

R code

# Confidence interval for a mean (t) and for the SD (chi-squared) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

media <- 95; de <- 40; n <- 50; nivel <- 0.95

# Two-sided Student t quantile with n - 1 degrees of freedom, and the standard error of the mean
tcrit <- qt(1 - (1 - nivel)/2, n - 1)
eem <- de/sqrt(n)
media_ic <- c(media, media - tcrit*eem, media + tcrit*eem)

# CI for the population SD from the chi-squared distribution of (n - 1)s^2/sigma^2 (equal tails)
de_ic <- c(de, de*sqrt((n - 1)/qchisq(1 - (1 - nivel)/2, n - 1)), de*sqrt((n - 1)/qchisq((1 - nivel)/2, n - 1)))

res <- list(media = media_ic, eem = eem, de = de_ic, t_crit = tcrit)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio with raw data: t.test(x, conf.level = nivel)$conf.int; DescTools::MeanCI(x)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The mean (95.00; SD 40.00; n = 50) is reported with its 95% confidence interval based on Student's t distribution with n − 1 degrees of freedom [1,2]: 83.63 to 106.37. The interval for the standard deviation was obtained from the chi-squared distribution of (n − 1)s²/σ², with equal tails: 33.41 to 49.85. Calculations used the "CI for a mean" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/ic-media), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
  2. 02 Gardner MJ, Altman DG. Confidence intervals rather than P values: estimation rather than hypothesis testing. British Medical Journal (Clinical Research Edition). 1986;292(6522):746–750. doi:10.1136/bmj.292.6522.746 PMID: 3082422 Didactic reading
  3. 03 Altman DG, Machin D, Bryant TN, Gardner MJ. Statistics with Confidence: Confidence Intervals and Statistical Guidelines. 2nd ed. London: BMJ Books; 2000. Didactic reading
  4. 04 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading