← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Choose which statistics the paper reports: only the range, only the interquartile range, or the five-number summary. Fields the scenario does not use are ignored in the calculation, even if you leave them filled in.

Number of participants with valid data for that variable (integer of 2 or more; quartiles need at least 4).

Lowest reported value.

Reported 25th percentile.

Reported 50th percentile.

Reported 75th percentile.

Highest reported value.

Example loaded

Illustrative example: days of fever in 45 patients, median 6 (interquartile range 4 to 9) and range 2 to 21 (fictitious data).

Illustrative data, not real.

Results

Estimated mean (Luo 2018)

6.924

recommended

Estimated mean (Wan 2014)

7.625

Estimated mean (Hozo 2005)

6

Estimated SD (Wan 2014)

4.071

recommended

Estimated SD (Hozo 2005)

4.75

Range skewness ratio

3.75

1 = symmetric summary

Interquartile skewness ratio

1.5

1 = symmetric summary

Interpretation

From a median of 6, an interquartile range of 4 to 9 and a range of 2 to 21 in 45 participants, the estimated mean is 6.924 (Luo 2018) and the standard deviation 4.071 (Wan 2014).

The three methods estimate means of 6.924 (Luo), 7.625 (Wan) and 6 (Hozo), and the standard deviations are 4.071 (Wan) and 4.75 (Hozo). The wider they spread, the more uncertainty the conversion itself adds.

The skewness ratios of the summary (3.75 for the range and 1.5 for the interquartile range) stray from 1: the distribution looks skewed and the formulas, which assume normality, may bias the mean and above all the standard deviation.

Use these figures only for meta-analysis or approximate comparisons: they are estimates derived from a published summary, not the original statistics of the study, and that should be stated in the methods.

  • The reported summary is clearly skewed: the formulas assume normality and the result may be biased (Cochrane Handbook §6.5.2.5 and §6.5.2.6).
Reported summary, estimated means and implied normal distributionMedian 6; Estimated mean (Luo 2018): 6.924; Estimated SD (Wan 2014): 4.071.Implied normal distributionMean (Luo)Mean (Wan)Mean (Hozo)5.010.015.020.0Value of the variable
Reported summary, estimated means and implied normal distribution

Explanation

Many studies describe their continuous variables with the median and the interquartile range, which is the right choice when the distribution is skewed. The problem shows up in a meta-analysis: most synthesis methods need a mean and a standard deviation. These formulas convert one summary into the other without access to the raw data.

All three proposals assume the data come from an approximately normal distribution and use the sample size to calibrate the conversion. Hozo et al. (2005) came first and works in ranges of n, with abrupt jumps at n = 15, 25 and 70. Wan et al. (2014) fixed those jumps by dividing the observed range by the range expected from a normal sample of the same size. Luo et al. (2018) refined the mean with optimal weights that depend on the scenario: only in S1 does the estimate tend to the median itself as n grows; in S2 and S3 the weights converge to 0.70 and the estimate stays close to the midpoint of the interquartile range.

That is why the calculator shows all three at once. If they agree, the conversion is stable and any of them will do; if they differ widely, that difference is real uncertainty worth exploring in a sensitivity analysis. Common practice in evidence synthesis is to use Luo's mean and Wan's standard deviation, and to keep Hozo as a historical comparison; the Cochrane Handbook (§6.5.2.5) supports the conversion from the interquartile range and (§6.5.2.6) advises against doing it from the range, so scenario S1, with only a minimum and a maximum, should be read with more caution than S2 and S3.

The skewness ratio of the summary itself is the quality check: if (max − median)/(median − min) or (Q₃ − median)/(median − Q₁) stray far from 1, the distribution does not look symmetric and the conversion may bias the mean and, above all, the standard deviation. In that case it is worth reporting the result as approximate or looking for the original data.

Equations

S1:  xˉ≈w1 a+b2+(1−w1) m,w1=44+n0.75S2:  xˉ≈w2 q1+q32+(1−w2) m,w2=0.70+0.39nS3:  xˉ≈w3 a+b2+w4 q1+q32+(1−w3−w4) mw3=2.22.2+n0.75,w4=0.70−0.72n0.55\begin{aligned} \text{S1:}\ \ \bar x &\approx w_1\,\frac{a+b}{2} + (1-w_1)\,m, \qquad w_1 = \frac{4}{4+n^{0.75}} \\ \text{S2:}\ \ \bar x &\approx w_2\,\frac{q_1+q_3}{2} + (1-w_2)\,m, \qquad w_2 = 0.70 + \frac{0.39}{n} \\ \text{S3:}\ \ \bar x &\approx w_3\,\frac{a+b}{2} + w_4\,\frac{q_1+q_3}{2} + (1-w_3-w_4)\,m \\ &\qquad w_3 = \frac{2.2}{2.2+n^{0.75}}, \qquad w_4 = 0.70 - \frac{0.72}{n^{0.55}} \end{aligned}
aa
reported minimum
mm
reported median
bb
reported maximum
q1, q3q_1,\ q_3
reported first and third quartiles
nn
number of participants in the study
Mean estimated by Luo et al. (2018). The weights were obtained by minimising the mean squared error under normality. As n grows, w₁ and w₃ tend to 0 and the S1 mean approaches the median; w₂ and w₄, by contrast, tend to 0.70, so S2 and S3 converge to 0.70·(q₁+q₃)/2 + 0.30·m.
xˉ≈{a+2m+b4S1q1+m+q33S2a+2q1+2m+2q3+b8S3\bar x \approx \begin{cases} \dfrac{a+2m+b}{4} & \text{S1} \\ \dfrac{q_1+m+q_3}{3} & \text{S2} \\ \dfrac{a+2q_1+2m+2q_3+b}{8} & \text{S3} \end{cases}
Mean estimated by Wan et al. (2014): weighted averages of the reported statistics, with no dependence on n.
s≈{b−aξ(n)S1q3−q1η(n)S212[b−aξ(n)+q3−q1η(n)]S3ξ(n)=2 Φ−1 ⁣(n−0.375n+0.25),η(n)=2 Φ−1 ⁣(0.75n−0.125n+0.25)\begin{gathered} s \approx \begin{cases} \dfrac{b-a}{\xi(n)} & \text{S1} \\ \dfrac{q_3-q_1}{\eta(n)} & \text{S2} \\ \dfrac{1}{2}\left[\dfrac{b-a}{\xi(n)} + \dfrac{q_3-q_1}{\eta(n)}\right] & \text{S3} \end{cases} \\ \xi(n) = 2\,\Phi^{-1}\!\left(\frac{n-0.375}{n+0.25}\right), \qquad \eta(n) = 2\,\Phi^{-1}\!\left(\frac{0.75n-0.125}{n+0.25}\right) \end{gathered}
ξ(n)\xi(n)
expected range of n standard normal values
η(n)\eta(n)
expected interquartile range of n standard normal values
Φ−1\Phi^{-1}
quantile of the standard normal distribution
Standard deviation estimated by Wan et al. (2014): the observed range (or interquartile range) divided by the one expected from a normal sample of the same size.
xˉ≈{a+2m+b4n≤25mn>25s≈{112[(a−2m+b)24+(b−a)2]n≤15b−a415<n≤70b−a6n>70\begin{gathered} \bar x \approx \begin{cases} \dfrac{a+2m+b}{4} & n \le 25 \\ m & n > 25 \end{cases} \\ s \approx \begin{cases} \sqrt{\dfrac{1}{12}\left[\dfrac{(a-2m+b)^2}{4} + (b-a)^2\right]} & n \le 15 \\ \dfrac{b-a}{4} & 15 < n \le 70 \\ \dfrac{b-a}{6} & n > 70 \end{cases} \end{gathered}
Hozo et al. (2005), the original proposal: it works in ranges of n and only with the minimum and the maximum. The jumps between ranges are discontinuous, which is exactly what Wan et al. (2014) corrected.

R code

# Mean and SD from median, range or IQR (Luo 2018, Wan 2014, Hozo 2005) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

escenario <- "s3"   # "s1" = min, median, max; "s2" = Q1, median, Q3; "s3" = the five numbers
n <- 45
a <- 2; q1 <- 4; m <- 6; q3 <- 9; b <- 21   # NA where the scenario does not use them

# Mean - Luo et al. 2018, with the optimal weights of each scenario
if (escenario == "s1") {
  w1 <- 4/(4 + n^0.75)
  media_luo <- w1*(a + b)/2 + (1 - w1)*m
} else if (escenario == "s2") {
  w2 <- 0.70 + 0.39/n
  media_luo <- w2*(q1 + q3)/2 + (1 - w2)*m
} else {
  w3 <- 2.2/(2.2 + n^0.75)
  w4 <- 0.70 - 0.72/n^0.55
  media_luo <- w3*(a + b)/2 + w4*(q1 + q3)/2 + (1 - w3 - w4)*m
}

# Mean - Wan et al. 2014
media_wan <- switch(escenario,
                    s1 = (a + 2*m + b)/4,
                    s2 = (q1 + m + q3)/3,
                    s3 = (a + 2*q1 + 2*m + 2*q3 + b)/8)

# SD - Wan et al. 2014: expected range xi(n) and expected IQR eta(n) of n standard normal values
xi  <- 2*qnorm((n - 0.375)/(n + 0.25))
eta <- 2*qnorm((0.75*n - 0.125)/(n + 0.25))
de_wan <- switch(escenario,
                 s1 = (b - a)/xi,
                 s2 = (q3 - q1)/eta,
                 s3 = ((b - a)/xi + (q3 - q1)/eta)/2)

# Hozo et al. 2005, by ranges of n; it needs the minimum and the maximum, so it does not apply to S2
if (escenario == "s2") {
  media_hozo <- NA_real_
  de_hozo <- NA_real_
} else {
  media_hozo <- if (n <= 25) (a + 2*m + b)/4 else m
  de_hozo <- if (n <= 15) sqrt(((a - 2*m + b)^2/4 + (b - a)^2)/12) else if (n <= 70) (b - a)/4 else (b - a)/6
}

# Skewness of the reported summary itself: 1 means a symmetric summary
asim_rango <- if (escenario == "s2") NA_real_ else (b - m)/(m - a)
asim_iqr   <- if (escenario == "s1") NA_real_ else (q3 - m)/(m - q1)

res <- list(media_luo = media_luo, media_wan = media_wan, media_hozo = media_hozo,
            de_wan = de_wan, de_hozo = de_hozo, asim_rango = asim_rango, asim_iqr = asim_iqr)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# No CRAN package on the browser list implements these; estmeansd::qe.mean.sd is an alternative in RStudio

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The unreported mean and standard deviation were estimated from the median, the interquartile range and the range in 45 participants with the methods of Luo et al. [3] and Wan et al. [2] (with Hozo et al. [1] for comparison); the Cochrane Handbook [4] covers the conversion from the interquartile range (§6.5.2.5) and advises against estimating the SD from the range (§6.5.2.6). The estimated mean was 6.924 and the standard deviation 4.071. Calculations used the "Mean and SD from the median" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/media-desde-mediana), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Hozo SP, Djulbegovic B, Hozo I. Estimating the mean and variance from the median, range, and the size of a sample. BMC Medical Research Methodology. 2005;5:13. doi:10.1186/1471-2288-5-13 PMID: 15840177 Original source
  2. 02 Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Medical Research Methodology. 2014;14:135. doi:10.1186/1471-2288-14-135 PMID: 25524443 Original source
  3. 03 Luo D, Wan X, Liu J, Tong T. Optimally estimating the sample mean from the sample size, median, mid-range, and/or mid-quartile range. Statistical Methods in Medical Research. 2018;27(6):1785–1805. doi:10.1177/0962280216669183 PMID: 27683581 Original source
  4. 04 Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA. Cochrane Handbook for Systematic Reviews of Interventions. 2nd ed. Chichester: John Wiley & Sons; 2019. doi:10.1002/9781119536604 Didactic reading
  5. 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading