← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

40 values read

Copy the column in your spreadsheet (Excel, Google Sheets, Numbers) and paste it here: one value per row. Values separated by tabs, commas or semicolons also work, and the decimal comma is recognised. A header row and any non-numeric text are dropped with a notice.

Example loaded

Illustrative example: platelet count (×10³/µL) in 40 febrile patients, simulated in R with set.seed(1190) and round(rnorm(40, 150, 45)) (fictitious data, not real).

Illustrative data, not real.

Results

Values (n)

40

Mean

142.6

128.6 to 156.6

95% CI · t distribution

Standard deviation (SD)

43.87

Standard error of the mean (SEM)

6.936

Median

147.5

type 7 quantile

First quartile (Q₁)

122.8

type 7 quantile

Third quartile (Q₃)

166

type 7 quantile

Interquartile range (IQR)

43.25

type 7 quantile

Minimum

47

Maximum

226

Range

179

Coefficient of variation (CV)

30.8%

Skewness (G1)

-0.45

Joanes and Gill, type 2

Kurtosis (G2)

0.03

Joanes and Gill, type 2

Geometric mean

134.3

Shapiro-Wilk (W)

0.962

AS R94 (Royston 1995)

Shapiro-Wilk (p)

0.196

AS R94 (Royston 1995)

Outliers (Tukey)

3

1.5 × IQR fences

Interpretation

40 values were described: mean 142.6 (SD 43.87; 95% CI for the mean 128.6 to 156.6), median 147.5 (IQR 122.8 to 166), range 47 to 226. The standard error of the mean is 6.936. The coefficient of variation is 30.8%: the standard deviation as a fraction of the mean.

Skewness G1 = -0.45: the distribution is roughly symmetric; kurtosis G2 = 0.03, where 0 corresponds to the normal, positive values to heavier tails and negative values to a flatter shape.

Shapiro-Wilk: W = 0.962, p = 0.196: there is no evidence against normality in these 40 values. The test does not prove that the variable is normal, only that these data are not enough to rule it out; with large n, small p values do not imply relevant departures.

For the manuscript: report mean (SD) when the distribution is roughly symmetric and free of outliers, that is 142.6 (43.87); report median (IQR) when it is not, that is 147.5 (122.8 to 166). Geometric mean, useful when the variable is multiplicative (titres, viral loads, dilutions): 134.3.

3 values fall outside Tukey's fences (below 57.88 or above 230.9, that is, more than 1.5 interquartile ranges away from the box). An outlier is not an error in itself: check the data entry and the units before excluding it, and if you do exclude it, say so.

  • 3 values fall outside Tukey's fences (more than 1.5 interquartile ranges below Q₁ or above Q₃). Check for data entry or unit errors before excluding them.
Histogram with the fitted normal curve and box plot of the same columnValues (n): 40; Mean: 142.6; Standard deviation (SD): 43.87; Median: 147.5 (122.8 to 166); Range: 47 to 226; Outliers (Tukey): 3Fitted normal curveFrequency051050100150200250Value of the variableMean
Histogram with the fitted normal curve and box plot of the same column

Explanation

Description comes before comparison. The practical question is which pair of numbers summarises the variable best: the mean with its standard deviation, which only reads well when the distribution is roughly symmetric, or the median with the interquartile range, which survives long tails and extreme values. This calculator gives both and adds the confidence interval of the mean, which answers "where is the population mean?" and must not be confused with the standard deviation, which describes how individuals are spread out.

Skewness G1 and kurtosis G2 put a number on the shape. G1 is 0 for a symmetric distribution, positive when there is a right tail (a few high values, as in length of stay or viral load) and negative when the tail is on the left. G2 is measured as an excess over the normal: 0 is the normal, above it the tails are heavier and the peak sharper, below it the distribution is flatter. Both use the G1 and G2 estimators of Joanes and Gill (1998), the ones reported by SPSS and SAS; as a practical cut-off, |G1| below 0.5 reads as roughly symmetric.

The Shapiro-Wilk test compares the ordered data with what a normal distribution would produce: W close to 1 means a good fit, and the p value is the probability of seeing a W that low if the variable were normal. It has two limits worth keeping in mind. With small samples it almost never rejects, even when the distribution is far from normal; with large samples it rejects because of minimal departures that change no clinical decision. That is why the result is read alongside the histogram and not instead of it, and why the test is not computed with fewer than 3 or more than 5,000 values.

Tukey's fences flag as an outlier every value more than 1.5 interquartile ranges below the first quartile or above the third. It is an exploratory rule, not an exclusion criterion: an outlier may be a data entry error, a wrong unit, or the most interesting patient in the study. The histogram uses Sturges' rule (1926) to choose the number of classes and the box plot draws those values as separate points, beyond the whiskers.

Equations

xˉ=1n∑i=1nxi,s=∑i=1n(xi−xˉ)2n−1,SE⁡(xˉ)=sn\bar x=\frac{1}{n}\sum_{i=1}^{n}x_i,\qquad s=\sqrt{\frac{\sum_{i=1}^{n}(x_i-\bar x)^2}{n-1}},\qquad \se(\bar x)=\frac{s}{\sqrt{n}}
nn
number of values in the column
xix_i
i-th value
ss
sample standard deviation (n − 1 denominator)
The standard deviation describes individuals; the standard error of the mean describes how precisely the mean was estimated.
CI1−α(xˉ)=xˉ±t1−α/2,  n−1⋅sn\CI_{1-\alpha}(\bar x)=\bar x\pm t_{1-\alpha/2,\;n-1}\cdot\frac{s}{\sqrt{n}}
t1−α/2,  n−1t_{1-\alpha/2,\;n-1}
Student's t quantile with n − 1 degrees of freedom
Confidence interval of the mean (Student 1908); the one returned by t.test() in R.
Q(p)=x(⌊h⌋)+(h−⌊h⌋)(x(⌈h⌉)−x(⌊h⌋)),h=1+(n−1) pQ(p)=x_{(\lfloor h\rfloor)}+\left(h-\lfloor h\rfloor\right)\left(x_{(\lceil h\rceil)}-x_{(\lfloor h\rfloor)}\right),\qquad h=1+(n-1)\,p
x(i)x_{(i)}
i-th value of the sorted column
pp
0.25 (Q₁), 0.5 (median) or 0.75 (Q₃)
Type 7 quantile of Hyndman and Fan (1996), R's default.
G1=m3m23/2⋅n(n−1)n−2,G2=n−1(n−2)(n−3)[(n+1)(m4m22−3)+6],mk=1n∑i=1n(xi−xˉ)kG_1=\frac{m_3}{m_2^{3/2}}\cdot\frac{\sqrt{n(n-1)}}{n-2},\qquad G_2=\frac{n-1}{(n-2)(n-3)}\left[(n+1)\left(\frac{m_4}{m_2^{2}}-3\right)+6\right],\qquad m_k=\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar x)^k
mkm_k
k-th sample central moment
The G1 and G2 estimators of Joanes and Gill (1998), "type 2": the SPSS and SAS pair. G1 needs n ≥ 3 and G2 needs n ≥ 4.
W=(∑i=1nai x(i))2∑i=1n(xi−xˉ)2W=\frac{\left(\sum_{i=1}^{n}a_i\,x_{(i)}\right)^{2}}{\sum_{i=1}^{n}(x_i-\bar x)^{2}}
aia_i
normalised coefficients derived from the expected normal order statistics
The statistic of Shapiro and Wilk (1965) with the coefficients and p value of Royston's AS R94 algorithm (1995), the one shapiro.test runs in R.
IQR=Q3−Q1,[ Q1−1.5⋅IQR,    Q3+1.5⋅IQR ]\mathrm{IQR}=Q_3-Q_1,\qquad \left[\,Q_1-1.5\cdot \mathrm{IQR},\;\; Q_3+1.5\cdot \mathrm{IQR}\,\right]
IQR\mathrm{IQR}
interquartile range, Q₃ − Q₁
Tukey's fences (1977): every value outside that interval is flagged as an outlier.

R code

# Descriptive statistics for one pasted column - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

# The pasted column, one numeric value per row:
x <- c(47, 165, 97, 138, 202, 156, 166, 141, 96, 132, 177, 51, 150, 117, 210,
  178, 208, 138, 171, 50, 162, 139, 134, 126, 160, 147, 196, 75, 166, 204,
  226, 148, 99, 123, 160, 147, 153, 148, 79, 122)
nivel <- 0.95
n <- length(x)

media <- mean(x); de <- sd(x); eem <- de/sqrt(n)   # sd() uses the n - 1 denominator

# t interval for the mean, written out because t.test() stops with constant data;
# otherwise identical to t.test(x, conf.level = nivel)$conf.int
tcrit <- qt(1 - (1 - nivel)/2, n - 1)
media_ic <- c(media, media - tcrit*eem, media + tcrit*eem)

q <- quantile(x, c(0.25, 0.5, 0.75), type = 7, names = FALSE)   # type 7 = R's default (Hyndman & Fan 1996)
iqr <- q[3] - q[1]
cv <- de/media

# Sample skewness G1 and excess kurtosis G2 (Joanes & Gill 1998, "type 2": the SPSS/SAS pair)
m <- function(k) mean((x - media)^k)
g1 <- if (n >= 3 && m(2) > 0) m(3)/m(2)^1.5 * sqrt(n*(n - 1))/(n - 2) else NA_real_
g2 <- if (n >= 4 && m(2) > 0) ((n + 1)*(m(4)/m(2)^2 - 3) + 6)*(n - 1)/((n - 2)*(n - 3)) else NA_real_

media_geom <- if (all(x > 0)) exp(mean(log(x))) else NA_real_   # undefined with zero or negative values

# Shapiro-Wilk (Royston 1995, AS R94): needs 3 <= n <= 5000 and a non-zero range
sw <- if (n >= 3 && n <= 5000 && diff(range(x)) > 0) shapiro.test(x) else NULL
sw_w <- if (is.null(sw)) NA_real_ else unname(sw$statistic)
sw_p <- if (is.null(sw)) NA_real_ else sw$p.value

n_atipicos <- sum(x < q[1] - 1.5*iqr | x > q[3] + 1.5*iqr)   # Tukey's fences

res <- list(n = n, media = media_ic, de = de, eem = eem,
            mediana = q[2], q1 = q[1], q3 = q[3], iqr = iqr,
            min = min(x), max = max(x), rango = diff(range(x)), cv = cv,
            g1 = g1, g2 = g2, media_geom = media_geom,
            sw_w = sw_w, sw_p = sw_p, n_atipicos = n_atipicos)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run in the browser):
# DescTools::Skew(x, method = 2); DescTools::Kurt(x, method = 2); e1071::skewness(x, type = 2)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Continuous variables were summarised as mean (standard deviation) or median (interquartile range) according to their distribution, assessed with the Shapiro-Wilk test [2,3] and visual inspection of the histogram and box plot [1]; quartiles were computed with type 7 of Hyndman and Fan [5], skewness and kurtosis with the G1 and G2 estimators [6] and the 95% confidence interval of the mean with Student's t distribution [8]. 40 values were described. Calculations used the "Descriptive statistics for one variable" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/descriptivos), verified against R (mean, sd, quantile with type = 7 and shapiro.test).

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Tukey JW. Exploratory Data Analysis. Reading, MA: Addison-Wesley; 1977. Original source
  2. 02 Shapiro SS, Wilk MB. An analysis of variance test for normality (complete samples). Biometrika. 1965;52(3-4):591–611. doi:10.1093/biomet/52.3-4.591 Original source
  3. 03 Royston P. Remark AS R94: a remark on algorithm AS 181: the W-test for normality. Journal of the Royal Statistical Society. Series C (Applied Statistics). 1995;44(4):547–551. doi:10.2307/2986146 Original source
  4. 04 Royston P. Approximating the Shapiro-Wilk W-test for non-normality. Statistics and Computing. 1992;2(3):117–119. doi:10.1007/BF01891203 Complementary
  5. 05 Hyndman RJ, Fan Y. Sample quantiles in statistical packages. The American Statistician. 1996;50(4):361–365. doi:10.1080/00031305.1996.10473566 Original source
  6. 06 Joanes DN, Gill CA. Comparing measures of sample skewness and kurtosis. Journal of the Royal Statistical Society. Series D (The Statistician). 1998;47(1):183–189. doi:10.1111/1467-9884.00122 Original source
  7. 07 Sturges HA. The choice of a class interval. Journal of the American Statistical Association. 1926;21(153):65–66. doi:10.1080/01621459.1926.10502161 Original source
  8. 08 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Complementary
  9. 09 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading