← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

The smallest correlation you would consider worth detecting, from −0.98 to 0.98 and different from 0. The sign does not change the sample size.

Probability of concluding that there is a correlation when there is none; usually 0.05.

Two-sided unless the direction of the correlation is fixed beforehand by the design, not by the data.

Probability of detecting the correlation if it really exists; usually 0.80 or 0.90.

Proportion of pairs you expect to lose or exclude (0.1 = 10%); leave it empty if you expect none.

If you already have a sample, enter its number of pairs (4 or more) and you will see the power it reaches; leave it empty for the forward calculation.

Example loaded

Illustrative example: a study is planned to detect a correlation of 0.30 between platelet count and days of fever, with α = 0.05 two-sided, 80% power and 10% expected losses (fictitious data).

Illustrative data, not real.

Results

Quantile of α (z)

1.960

Quantile of power (z)

0.842

Fisher's z of r (C)

0.3095

Pairs needed

85

exact value 84.93

Pairs according to pwr.r.test

85

comparison (bias correction and t quantile)

Pairs with expected losses

95

exact value 94.36

Power with the pairs available

no n available

Interpretation

To detect a correlation of at least 0.30 with a power of 80% and α = 0.050 (two-sided), 85 pairs of observations are required.

Fisher's z transformation gives C = 0.3095 and an exact size of 84.93 pairs. R's pwr package, which corrects the bias of C and uses the t quantile, asks for 85: both routes answer the same question and the difference is one of rounding, not of method.

Allowing for 10% losses or exclusions, 95 pairs must be recruited in order to end with 85 analysable ones.

If you already have a sample, enter its number of pairs and the calculator will tell you what power it reaches for this correlation.

  • The figures depend entirely on the correlation that is assumed: take it from previous studies or a pilot and document its source in the protocol. The formula also assumes independent pairs from a bivariate normal distribution.
Power against the number of pairs, with the computed size and the target power markedComputed size 85 Pairs of observations (n); Target power 80%.PowerPower from the classic formula0%25%50%75%100%50100150Pairs of observations (n)Target powerComputed size
Power against the number of pairs, with the computed size and the target power marked

Explanation

The correlation coefficient r is not normally distributed: its distribution piles up against 1 as the correlation grows. Fisher (1915, 1921) solved the problem with a change of scale, C = ½·ln((1 + r)/(1 − r)), which is nearly normal and whose standard error depends only on sample size: 1/√(n − 3).

On that scale the sample size calculation is the usual one: the distance between the null value (C = 0) and the value worth detecting must cover z for α plus z for power, in standard error units. That gives n = ((z + z)/C)² + 3, the formula found in clinical research textbooks.

What matters is the magnitude of the correlation, not its sign: detecting r = −0.30 costs exactly the same as detecting r = 0.30. And the cost grows very fast when the correlation is small, because C ≈ r for low values and n goes with 1/C²: moving from r = 0.30 to r = 0.10 multiplies the size by nine.

R's `pwr` package solves the same problem with two refinements: it corrects the bias of C and uses the t quantile instead of the normal one. It lands one or two subjects lower, and it appears here as a comparison so that the result is reproducible along either route. The calculator reports when the two figures differ after rounding.

Equations

C=artanh⁡(r)=12ln⁡1+r1−r,SE⁡(C)=1n−3C=\operatorname{artanh}(r)=\tfrac{1}{2}\ln\frac{1+r}{1-r},\qquad \se(C)=\frac{1}{\sqrt{n-3}}
rr
Pearson correlation worth detecting
CC
Fisher's z transformation of r, nearly normal
nn
number of pairs of observations
Fisher's z transformation (1915, 1921); it assumes independent pairs from a bivariate normal distribution.
n=(z1−α/k+z1−βC)2+3n=\left(\frac{z_{1-\alpha/k}+z_{1-\beta}}{C}\right)^{2}+3
z1−α/kz_{1-\alpha/k}
normal quantile of α, in whichever tail the test uses
kk
2 if the test is two-sided, 1 if it is one-sided
z1−βz_{1-\beta}
normal quantile of power
Classic formula: the 3 gives back the degrees of freedom spent by the transformation. With a two-sided hypothesis α is split between both tails (k = 2) and with a one-sided one it is concentrated in a single tail (k = 1), which is what makes the study cheaper.
1−β=Φ(∣C∣n−3−z1−α/k),naj=⌈n1−L⌉1-\beta=\Phi\left(|C|\sqrt{n-3}-z_{1-\alpha/k}\right),\qquad n_{aj}=\left\lceil\frac{n}{1-L}\right\rceil
Φ\Phi
cumulative distribution function of the standard normal
LL
expected losses, as a proportion
⌈  ⌉\lceil\;\rceil
ceiling: rounding up is applied only once, at the end
Inverse mode (power with the n available) and adjustment for losses from Lwanga and Lemeshow (1991).

R code

# Sample size to detect a correlation (Fisher's z transformation) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
library(pwr)

r <- 0.3                      # correlation worth detecting (not 0, |r| <= 0.98)
alfa <- 0.05                # significance level
lateralidad <- "bilateral"  # "bilateral" or "unilateral"
poder <- 0.8              # target power
perdidas <- 0.1        # expected losses (0-0.5)
n_dado <- 0            # pairs already available (inverse mode); 0 = not given

tside <- if (lateralidad == "bilateral") 2 else 1
z_alfa <- qnorm(alfa / tside, lower.tail = FALSE)
z_beta <- qnorm(poder)

# Fisher (1915, 1921): atanh(r) is nearly normal with standard error 1/sqrt(n - 3)
c_fisher <- atanh(r)                               # 0.5 * log((1 + r) / (1 - r))
n_clasico <- ((z_alfa + z_beta) / c_fisher)^2 + 3
n_ajustado <- n_clasico / (1 - perdidas)           # Lwanga & Lemeshow (1991)

# pwr adds the bias correction r/(2*(n - 1)) and uses the t quantile instead of
# the normal one, so it lands one or two subjects below the classic formula.
# With a very large r and a modest power its uniroot finds no sign change (the
# four pairs of the lower bound already exceed the power) and it stops: that
# combination is reported as NA instead of aborting the whole script.
alternativa <- if (tside == 2) "two.sided" else if (r > 0) "greater" else "less"
n_pwr <- tryCatch(pwr.r.test(r = r, sig.level = alfa, power = poder, alternative = alternativa)$n,
                  error = function(e) NA_real_)

# Inverse mode: power reached with the pairs already available.
poder_dado <- if (n_dado >= 4) pnorm(abs(c_fisher) * sqrt(n_dado - 3) - z_alfa) else NA_real_

# Values travel WITHOUT rounding: the interface takes the ceiling once, at the end.
res <- list(z_alfa = z_alfa, z_beta = z_beta, c_fisher = c_fisher, n_clasico = n_clasico,
            n_pwr = n_pwr, n_ajustado = n_ajustado, poder_dado = poder_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio: presize::prec_cor(r = r, conf.width = 0.2, conf.level = 1 - alfa, method = "fisher")

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The sample size was computed to detect a Pearson correlation of 0.30 with a power of 80% and α = 0.050 (two-sided), using Fisher's z transformation [1,2]: 85 pairs of observations. An additional 10% was added for expected losses or exclusions, bringing planned recruitment to 95 pairs. The result was checked against the pwr.r.test function of the pwr package [4], which gives 85 pairs. Calculations used the "Sample size for a correlation" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-correlacion), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Fisher RA. Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population. Biometrika. 1915;10(4):507–521. doi:10.1093/biomet/10.4.507 Original source
  2. 02 Fisher RA. On the “probable error” of a coefficient of correlation deduced from a small sample. Metron. 1921;1:3–32. Original source
  3. 03 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
  4. 04 Champely S. pwr: Basic Functions for Power Analysis. R package version 1.3-0. CRAN; 2020. Complementary
  5. 05 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
  6. 06 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading