Tools · Open Biostatistics
Mean and standard deviation from the median, the range or the interquartile range
Choose which statistics the study reported, enter them together with the number of participants and get the mean and standard deviation estimated by the three published methods, with the skewness warning, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/media-desde-mediana
This link does not include the pasted data: they are too long for a URL.
Results
Estimated mean (Luo 2018)
6.924
recommended
Estimated mean (Wan 2014)
7.625
Estimated mean (Hozo 2005)
6
Estimated SD (Wan 2014)
4.071
recommended
Estimated SD (Hozo 2005)
4.75
Range skewness ratio
3.75
1 = symmetric summary
Interquartile skewness ratio
1.5
1 = symmetric summary
Interpretation
From a median of 6, an interquartile range of 4 to 9 and a range of 2 to 21 in 45 participants, the estimated mean is 6.924 (Luo 2018) and the standard deviation 4.071 (Wan 2014).
The three methods estimate means of 6.924 (Luo), 7.625 (Wan) and 6 (Hozo), and the standard deviations are 4.071 (Wan) and 4.75 (Hozo). The wider they spread, the more uncertainty the conversion itself adds.
The skewness ratios of the summary (3.75 for the range and 1.5 for the interquartile range) stray from 1: the distribution looks skewed and the formulas, which assume normality, may bias the mean and above all the standard deviation.
Use these figures only for meta-analysis or approximate comparisons: they are estimates derived from a published summary, not the original statistics of the study, and that should be stated in the methods.
- The reported summary is clearly skewed: the formulas assume normality and the result may be biased (Cochrane Handbook §6.5.2.5 and §6.5.2.6).
Explanation
Many studies describe their continuous variables with the median and the interquartile range, which is the right choice when the distribution is skewed. The problem shows up in a meta-analysis: most synthesis methods need a mean and a standard deviation. These formulas convert one summary into the other without access to the raw data.
All three proposals assume the data come from an approximately normal distribution and use the sample size to calibrate the conversion. Hozo et al. (2005) came first and works in ranges of n, with abrupt jumps at n = 15, 25 and 70. Wan et al. (2014) fixed those jumps by dividing the observed range by the range expected from a normal sample of the same size. Luo et al. (2018) refined the mean with optimal weights that depend on the scenario: only in S1 does the estimate tend to the median itself as n grows; in S2 and S3 the weights converge to 0.70 and the estimate stays close to the midpoint of the interquartile range.
That is why the calculator shows all three at once. If they agree, the conversion is stable and any of them will do; if they differ widely, that difference is real uncertainty worth exploring in a sensitivity analysis. Common practice in evidence synthesis is to use Luo's mean and Wan's standard deviation, and to keep Hozo as a historical comparison; the Cochrane Handbook (§6.5.2.5) supports the conversion from the interquartile range and (§6.5.2.6) advises against doing it from the range, so scenario S1, with only a minimum and a maximum, should be read with more caution than S2 and S3.
The skewness ratio of the summary itself is the quality check: if (max − median)/(median − min) or (Q₃ − median)/(median − Q₁) stray far from 1, the distribution does not look symmetric and the conversion may bias the mean and, above all, the standard deviation. In that case it is worth reporting the result as approximate or looking for the original data.
Equations
- reported minimum
- reported median
- reported maximum
- reported first and third quartiles
- number of participants in the study
- expected range of n standard normal values
- expected interquartile range of n standard normal values
- quantile of the standard normal distribution
R code
# Mean and SD from median, range or IQR (Luo 2018, Wan 2014, Hozo 2005) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
escenario <- "s3" # "s1" = min, median, max; "s2" = Q1, median, Q3; "s3" = the five numbers
n <- 45
a <- 2; q1 <- 4; m <- 6; q3 <- 9; b <- 21 # NA where the scenario does not use them
# Mean - Luo et al. 2018, with the optimal weights of each scenario
if (escenario == "s1") {
w1 <- 4/(4 + n^0.75)
media_luo <- w1*(a + b)/2 + (1 - w1)*m
} else if (escenario == "s2") {
w2 <- 0.70 + 0.39/n
media_luo <- w2*(q1 + q3)/2 + (1 - w2)*m
} else {
w3 <- 2.2/(2.2 + n^0.75)
w4 <- 0.70 - 0.72/n^0.55
media_luo <- w3*(a + b)/2 + w4*(q1 + q3)/2 + (1 - w3 - w4)*m
}
# Mean - Wan et al. 2014
media_wan <- switch(escenario,
s1 = (a + 2*m + b)/4,
s2 = (q1 + m + q3)/3,
s3 = (a + 2*q1 + 2*m + 2*q3 + b)/8)
# SD - Wan et al. 2014: expected range xi(n) and expected IQR eta(n) of n standard normal values
xi <- 2*qnorm((n - 0.375)/(n + 0.25))
eta <- 2*qnorm((0.75*n - 0.125)/(n + 0.25))
de_wan <- switch(escenario,
s1 = (b - a)/xi,
s2 = (q3 - q1)/eta,
s3 = ((b - a)/xi + (q3 - q1)/eta)/2)
# Hozo et al. 2005, by ranges of n; it needs the minimum and the maximum, so it does not apply to S2
if (escenario == "s2") {
media_hozo <- NA_real_
de_hozo <- NA_real_
} else {
media_hozo <- if (n <= 25) (a + 2*m + b)/4 else m
de_hozo <- if (n <= 15) sqrt(((a - 2*m + b)^2/4 + (b - a)^2)/12) else if (n <= 70) (b - a)/4 else (b - a)/6
}
# Skewness of the reported summary itself: 1 means a symmetric summary
asim_rango <- if (escenario == "s2") NA_real_ else (b - m)/(m - a)
asim_iqr <- if (escenario == "s1") NA_real_ else (q3 - m)/(m - q1)
res <- list(media_luo = media_luo, media_wan = media_wan, media_hozo = media_hozo,
de_wan = de_wan, de_hozo = de_hozo, asim_rango = asim_rango, asim_iqr = asim_iqr)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# No CRAN package on the browser list implements these; estmeansd::qe.mean.sd is an alternative in RStudio
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The unreported mean and standard deviation were estimated from the median, the interquartile range and the range in 45 participants with the methods of Luo et al. [3] and Wan et al. [2] (with Hozo et al. [1] for comparison); the Cochrane Handbook [4] covers the conversion from the interquartile range (§6.5.2.5) and advises against estimating the SD from the range (§6.5.2.6). The estimated mean was 6.924 and the standard deviation 4.071. Calculations used the "Mean and SD from the median" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/media-desde-mediana), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Hozo SP, Djulbegovic B, Hozo I. Estimating the mean and variance from the median, range, and the size of a sample. BMC Medical Research Methodology. 2005;5:13. doi:10.1186/1471-2288-5-13 PMID: 15840177 Original source
- 02 Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Medical Research Methodology. 2014;14:135. doi:10.1186/1471-2288-14-135 PMID: 25524443 Original source
- 03 Luo D, Wan X, Liu J, Tong T. Optimally estimating the sample mean from the sample size, median, mid-range, and/or mid-quartile range. Statistical Methods in Medical Research. 2018;27(6):1785–1805. doi:10.1177/0962280216669183 PMID: 27683581 Original source
- 04 Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA. Cochrane Handbook for Systematic Reviews of Interventions. 2nd ed. Chichester: John Wiley & Sons; 2019. doi:10.1002/9781119536604 Didactic reading
- 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading