Tools · Open Biostatistics
Sample size to detect a correlation (Fisher's z transformation)
Enter the correlation you want to be able to detect, the significance level and the desired power, and get how many pairs of observations you need, the power you would reach with the ones you already have, the interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-correlacion
This link does not include the pasted data: they are too long for a URL.
Results
Quantile of α (z)
1.960
Quantile of power (z)
0.842
Fisher's z of r (C)
0.3095
Pairs needed
85
exact value 84.93
Pairs according to pwr.r.test
85
comparison (bias correction and t quantile)
Pairs with expected losses
95
exact value 94.36
Power with the pairs available
no n available
Interpretation
To detect a correlation of at least 0.30 with a power of 80% and α = 0.050 (two-sided), 85 pairs of observations are required.
Fisher's z transformation gives C = 0.3095 and an exact size of 84.93 pairs. R's pwr package, which corrects the bias of C and uses the t quantile, asks for 85: both routes answer the same question and the difference is one of rounding, not of method.
Allowing for 10% losses or exclusions, 95 pairs must be recruited in order to end with 85 analysable ones.
If you already have a sample, enter its number of pairs and the calculator will tell you what power it reaches for this correlation.
- The figures depend entirely on the correlation that is assumed: take it from previous studies or a pilot and document its source in the protocol. The formula also assumes independent pairs from a bivariate normal distribution.
Explanation
The correlation coefficient r is not normally distributed: its distribution piles up against 1 as the correlation grows. Fisher (1915, 1921) solved the problem with a change of scale, C = ½·ln((1 + r)/(1 − r)), which is nearly normal and whose standard error depends only on sample size: 1/√(n − 3).
On that scale the sample size calculation is the usual one: the distance between the null value (C = 0) and the value worth detecting must cover z for α plus z for power, in standard error units. That gives n = ((z + z)/C)² + 3, the formula found in clinical research textbooks.
What matters is the magnitude of the correlation, not its sign: detecting r = −0.30 costs exactly the same as detecting r = 0.30. And the cost grows very fast when the correlation is small, because C ≈ r for low values and n goes with 1/C²: moving from r = 0.30 to r = 0.10 multiplies the size by nine.
R's `pwr` package solves the same problem with two refinements: it corrects the bias of C and uses the t quantile instead of the normal one. It lands one or two subjects lower, and it appears here as a comparison so that the result is reproducible along either route. The calculator reports when the two figures differ after rounding.
Equations
- Pearson correlation worth detecting
- Fisher's z transformation of r, nearly normal
- number of pairs of observations
- normal quantile of α, in whichever tail the test uses
- 2 if the test is two-sided, 1 if it is one-sided
- normal quantile of power
- cumulative distribution function of the standard normal
- expected losses, as a proportion
- ceiling: rounding up is applied only once, at the end
R code
# Sample size to detect a correlation (Fisher's z transformation) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
library(pwr)
r <- 0.3 # correlation worth detecting (not 0, |r| <= 0.98)
alfa <- 0.05 # significance level
lateralidad <- "bilateral" # "bilateral" or "unilateral"
poder <- 0.8 # target power
perdidas <- 0.1 # expected losses (0-0.5)
n_dado <- 0 # pairs already available (inverse mode); 0 = not given
tside <- if (lateralidad == "bilateral") 2 else 1
z_alfa <- qnorm(alfa / tside, lower.tail = FALSE)
z_beta <- qnorm(poder)
# Fisher (1915, 1921): atanh(r) is nearly normal with standard error 1/sqrt(n - 3)
c_fisher <- atanh(r) # 0.5 * log((1 + r) / (1 - r))
n_clasico <- ((z_alfa + z_beta) / c_fisher)^2 + 3
n_ajustado <- n_clasico / (1 - perdidas) # Lwanga & Lemeshow (1991)
# pwr adds the bias correction r/(2*(n - 1)) and uses the t quantile instead of
# the normal one, so it lands one or two subjects below the classic formula.
# With a very large r and a modest power its uniroot finds no sign change (the
# four pairs of the lower bound already exceed the power) and it stops: that
# combination is reported as NA instead of aborting the whole script.
alternativa <- if (tside == 2) "two.sided" else if (r > 0) "greater" else "less"
n_pwr <- tryCatch(pwr.r.test(r = r, sig.level = alfa, power = poder, alternative = alternativa)$n,
error = function(e) NA_real_)
# Inverse mode: power reached with the pairs already available.
poder_dado <- if (n_dado >= 4) pnorm(abs(c_fisher) * sqrt(n_dado - 3) - z_alfa) else NA_real_
# Values travel WITHOUT rounding: the interface takes the ceiling once, at the end.
res <- list(z_alfa = z_alfa, z_beta = z_beta, c_fisher = c_fisher, n_clasico = n_clasico,
n_pwr = n_pwr, n_ajustado = n_ajustado, poder_dado = poder_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio: presize::prec_cor(r = r, conf.width = 0.2, conf.level = 1 - alfa, method = "fisher")
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The sample size was computed to detect a Pearson correlation of 0.30 with a power of 80% and α = 0.050 (two-sided), using Fisher's z transformation [1,2]: 85 pairs of observations. An additional 10% was added for expected losses or exclusions, bringing planned recruitment to 95 pairs. The result was checked against the pwr.r.test function of the pwr package [4], which gives 85 pairs. Calculations used the "Sample size for a correlation" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-correlacion), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Fisher RA. Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population. Biometrika. 1915;10(4):507–521. doi:10.1093/biomet/10.4.507 Original source
- 02 Fisher RA. On the “probable error” of a coefficient of correlation deduced from a small sample. Metron. 1921;1:3–32. Original source
- 03 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
- 04 Champely S. pwr: Basic Functions for Power Analysis. R package version 1.3-0. CRAN; 2020. Complementary
- 05 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
- 06 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading