Tools · Open Biostatistics
Descriptive statistics for one variable: mean, SD, confidence interval, quartiles, skewness, kurtosis and the Shapiro-Wilk test
Paste a column from your spreadsheet and get the full summary of that variable: mean with confidence interval, standard deviation, median and quartiles, skewness and kurtosis, the Shapiro-Wilk normality test and outliers, together with the histogram and box plot you should look at before deciding what to report.
https://udgca1190.com.mx/en/herramientas/bioestadistica/descriptivos
This link does not include the pasted data: they are too long for a URL.
Results
Values (n)
40
Mean
142.6
128.6 to 156.6
95% CI · t distribution
Standard deviation (SD)
43.87
Standard error of the mean (SEM)
6.936
Median
147.5
type 7 quantile
First quartile (Q₁)
122.8
type 7 quantile
Third quartile (Q₃)
166
type 7 quantile
Interquartile range (IQR)
43.25
type 7 quantile
Minimum
47
Maximum
226
Range
179
Coefficient of variation (CV)
30.8%
Skewness (G1)
-0.45
Joanes and Gill, type 2
Kurtosis (G2)
0.03
Joanes and Gill, type 2
Geometric mean
134.3
Shapiro-Wilk (W)
0.962
AS R94 (Royston 1995)
Shapiro-Wilk (p)
0.196
AS R94 (Royston 1995)
Outliers (Tukey)
3
1.5 × IQR fences
Interpretation
40 values were described: mean 142.6 (SD 43.87; 95% CI for the mean 128.6 to 156.6), median 147.5 (IQR 122.8 to 166), range 47 to 226. The standard error of the mean is 6.936. The coefficient of variation is 30.8%: the standard deviation as a fraction of the mean.
Skewness G1 = -0.45: the distribution is roughly symmetric; kurtosis G2 = 0.03, where 0 corresponds to the normal, positive values to heavier tails and negative values to a flatter shape.
Shapiro-Wilk: W = 0.962, p = 0.196: there is no evidence against normality in these 40 values. The test does not prove that the variable is normal, only that these data are not enough to rule it out; with large n, small p values do not imply relevant departures.
For the manuscript: report mean (SD) when the distribution is roughly symmetric and free of outliers, that is 142.6 (43.87); report median (IQR) when it is not, that is 147.5 (122.8 to 166). Geometric mean, useful when the variable is multiplicative (titres, viral loads, dilutions): 134.3.
3 values fall outside Tukey's fences (below 57.88 or above 230.9, that is, more than 1.5 interquartile ranges away from the box). An outlier is not an error in itself: check the data entry and the units before excluding it, and if you do exclude it, say so.
- 3 values fall outside Tukey's fences (more than 1.5 interquartile ranges below Q₁ or above Q₃). Check for data entry or unit errors before excluding them.
Explanation
Description comes before comparison. The practical question is which pair of numbers summarises the variable best: the mean with its standard deviation, which only reads well when the distribution is roughly symmetric, or the median with the interquartile range, which survives long tails and extreme values. This calculator gives both and adds the confidence interval of the mean, which answers "where is the population mean?" and must not be confused with the standard deviation, which describes how individuals are spread out.
Skewness G1 and kurtosis G2 put a number on the shape. G1 is 0 for a symmetric distribution, positive when there is a right tail (a few high values, as in length of stay or viral load) and negative when the tail is on the left. G2 is measured as an excess over the normal: 0 is the normal, above it the tails are heavier and the peak sharper, below it the distribution is flatter. Both use the G1 and G2 estimators of Joanes and Gill (1998), the ones reported by SPSS and SAS; as a practical cut-off, |G1| below 0.5 reads as roughly symmetric.
The Shapiro-Wilk test compares the ordered data with what a normal distribution would produce: W close to 1 means a good fit, and the p value is the probability of seeing a W that low if the variable were normal. It has two limits worth keeping in mind. With small samples it almost never rejects, even when the distribution is far from normal; with large samples it rejects because of minimal departures that change no clinical decision. That is why the result is read alongside the histogram and not instead of it, and why the test is not computed with fewer than 3 or more than 5,000 values.
Tukey's fences flag as an outlier every value more than 1.5 interquartile ranges below the first quartile or above the third. It is an exploratory rule, not an exclusion criterion: an outlier may be a data entry error, a wrong unit, or the most interesting patient in the study. The histogram uses Sturges' rule (1926) to choose the number of classes and the box plot draws those values as separate points, beyond the whiskers.
Equations
- number of values in the column
- i-th value
- sample standard deviation (n − 1 denominator)
- Student's t quantile with n − 1 degrees of freedom
- i-th value of the sorted column
- 0.25 (Q₁), 0.5 (median) or 0.75 (Q₃)
- k-th sample central moment
- normalised coefficients derived from the expected normal order statistics
- interquartile range, Q₃ − Q₁
R code
# Descriptive statistics for one pasted column - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
# The pasted column, one numeric value per row:
x <- c(47, 165, 97, 138, 202, 156, 166, 141, 96, 132, 177, 51, 150, 117, 210,
178, 208, 138, 171, 50, 162, 139, 134, 126, 160, 147, 196, 75, 166, 204,
226, 148, 99, 123, 160, 147, 153, 148, 79, 122)
nivel <- 0.95
n <- length(x)
media <- mean(x); de <- sd(x); eem <- de/sqrt(n) # sd() uses the n - 1 denominator
# t interval for the mean, written out because t.test() stops with constant data;
# otherwise identical to t.test(x, conf.level = nivel)$conf.int
tcrit <- qt(1 - (1 - nivel)/2, n - 1)
media_ic <- c(media, media - tcrit*eem, media + tcrit*eem)
q <- quantile(x, c(0.25, 0.5, 0.75), type = 7, names = FALSE) # type 7 = R's default (Hyndman & Fan 1996)
iqr <- q[3] - q[1]
cv <- de/media
# Sample skewness G1 and excess kurtosis G2 (Joanes & Gill 1998, "type 2": the SPSS/SAS pair)
m <- function(k) mean((x - media)^k)
g1 <- if (n >= 3 && m(2) > 0) m(3)/m(2)^1.5 * sqrt(n*(n - 1))/(n - 2) else NA_real_
g2 <- if (n >= 4 && m(2) > 0) ((n + 1)*(m(4)/m(2)^2 - 3) + 6)*(n - 1)/((n - 2)*(n - 3)) else NA_real_
media_geom <- if (all(x > 0)) exp(mean(log(x))) else NA_real_ # undefined with zero or negative values
# Shapiro-Wilk (Royston 1995, AS R94): needs 3 <= n <= 5000 and a non-zero range
sw <- if (n >= 3 && n <= 5000 && diff(range(x)) > 0) shapiro.test(x) else NULL
sw_w <- if (is.null(sw)) NA_real_ else unname(sw$statistic)
sw_p <- if (is.null(sw)) NA_real_ else sw$p.value
n_atipicos <- sum(x < q[1] - 1.5*iqr | x > q[3] + 1.5*iqr) # Tukey's fences
res <- list(n = n, media = media_ic, de = de, eem = eem,
mediana = q[2], q1 = q[1], q3 = q[3], iqr = iqr,
min = min(x), max = max(x), rango = diff(range(x)), cv = cv,
g1 = g1, g2 = g2, media_geom = media_geom,
sw_w = sw_w, sw_p = sw_p, n_atipicos = n_atipicos)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# DescTools::Skew(x, method = 2); DescTools::Kurt(x, method = 2); e1071::skewness(x, type = 2)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Continuous variables were summarised as mean (standard deviation) or median (interquartile range) according to their distribution, assessed with the Shapiro-Wilk test [2,3] and visual inspection of the histogram and box plot [1]; quartiles were computed with type 7 of Hyndman and Fan [5], skewness and kurtosis with the G1 and G2 estimators [6] and the 95% confidence interval of the mean with Student's t distribution [8]. 40 values were described. Calculations used the "Descriptive statistics for one variable" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/descriptivos), verified against R (mean, sd, quantile with type = 7 and shapiro.test).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Tukey JW. Exploratory Data Analysis. Reading, MA: Addison-Wesley; 1977. Original source
- 02 Shapiro SS, Wilk MB. An analysis of variance test for normality (complete samples). Biometrika. 1965;52(3-4):591–611. doi:10.1093/biomet/52.3-4.591 Original source
- 03 Royston P. Remark AS R94: a remark on algorithm AS 181: the W-test for normality. Journal of the Royal Statistical Society. Series C (Applied Statistics). 1995;44(4):547–551. doi:10.2307/2986146 Original source
- 04 Royston P. Approximating the Shapiro-Wilk W-test for non-normality. Statistics and Computing. 1992;2(3):117–119. doi:10.1007/BF01891203 Complementary
- 05 Hyndman RJ, Fan Y. Sample quantiles in statistical packages. The American Statistician. 1996;50(4):361–365. doi:10.1080/00031305.1996.10473566 Original source
- 06 Joanes DN, Gill CA. Comparing measures of sample skewness and kurtosis. Journal of the Royal Statistical Society. Series D (The Statistician). 1998;47(1):183–189. doi:10.1111/1467-9884.00122 Original source
- 07 Sturges HA. The choice of a class interval. Journal of the American Statistical Association. 1926;21(153):65–66. doi:10.1080/01621459.1926.10502161 Original source
- 08 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Complementary
- 09 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading