Tools · Open Biostatistics
Kaplan-Meier curve, median survival and the log-rank test
Paste the follow-up time, the event indicator and (if you have one) the group variable, and get the Kaplan-Meier curve with its confidence band, median survival, survival at the times you care about, and the comparison of two curves with the log-rank test.
https://udgca1190.com.mx/en/herramientas/bioestadistica/kaplan-meier
This link does not include the pasted data: they are too long for a URL.
Results
Patients in the first curve
15
Events in the first curve
12
Median survival (first curve)
6
3 to 9
95% CI · quantile.survfit rule
Survival at the first time (first curve)
38.9%
15.3% to 62.2%
95% CI · Greenwood
Survival at the second time (first curve)
11.7%
0.9% to 37.6%
95% CI · Greenwood
Patients in the second curve
15
Events in the second curve
12
Median survival (second curve)
14
7 to 18
95% CI · quantile.survfit rule
Survival at the first time (second curve)
80.0%
50.0% to 93.1%
95% CI · Greenwood
Survival at the second time (second curve)
44.0%
18.5% to 67.1%
95% CI · Greenwood
Log-rank (χ²)
5.174
log-rank, rho = 0
Degrees of freedom
1
log-rank, rho = 0
Log-rank (p)
0.023
log-rank, rho = 0
Interpretation
Two curves were compared: in group 0 the number of patients was 15 and the number of events 12; in group 1, 15 and 12. At time 7, the estimated probability of still being event-free was 38.9% (95% CI 15.3% to 62.2%) in group 0 and 80.0% (95% CI 50.0% to 93.1%) in group 1. At time 14, the estimated probability of still being event-free was 11.7% (95% CI 0.9% to 37.6%) in group 0 and 44.0% (95% CI 18.5% to 67.1%) in group 1.
Median survival: 6 (95% CI 3 to 9) in group 0 and 14 (95% CI 7 to 18) in group 1. It is the time at which each curve crosses 50%; when a curve ends above 0.5 the median is not reached and must be reported that way, never as the last observed time.
Log-rank test: χ² = 5.174 with 1 degree of freedom, p = 0.023. At α = 0.05 the null hypothesis that the two curves are identical is rejected. The p value says the observed difference would be rare if there were none, but not how large it is: that is what the medians, the survival at the times of interest and the gap between the curves say.
Follow-up was censored in 6 of 30 patients (20.0%); the split by curve was group 0: 3/15 · group 1: 3/15. Censoring is informative only if it is unrelated to prognosis: loss to follow-up concentrated among the patients who were doing worse biases the curve upwards, and the manuscript should say why follow-up was lost.
The curves do not cross within the common follow-up, which is the situation in which the log-rank test and a single hazard ratio can be read without reservations.
Explanation
Survival analysis answers a question a proportion cannot: how long something takes to happen. The "event" need not be death; it can be defervescence, discharge, relapse, seroconversion or catheter failure. What makes these data special is censoring: when the study closes, some patients have not had the event yet, and others were lost before having it. Those patients are neither failures nor successes: they contribute the information that they were event-free up to the day we stopped seeing them, and dropping them would bias the result.
The Kaplan-Meier estimator (1958) uses that information. At each time with at least one event it computes the conditional probability of getting through it event-free, 1 − dᵢ/nᵢ, where nᵢ is the number still at risk just before that moment, and multiplies those probabilities one after another. The result is the step curve: it drops only when events occur and stays flat when there are only censorings, even though every censoring shrinks the risk set and makes later steps larger. That is why the numbers-at-risk table under the plot is not decoration: the tail of the curve, where two or three people are left, is the least reliable part and the most misleading at a glance.
The width of the confidence band comes from Greenwood's formula (1926), which accumulates the uncertainty of every step. It is computed on the log scale and then carried back to the survival scale: by default with the log-log transformation of Kalbfleisch and Prentice, which keeps the limits between 0 and 1 and behaves better in the tails, and alternatively on the log scale, which is R's own default and the one many textbooks show. It is the same difference that in R separates `conf.type = "log-log"` from `conf.type = "log"`, which is why the selector travels into the code.
Median survival is the time at which the curve crosses 50%. It is preferred over the mean because it does not depend on the censored tail, and it may not exist: if the curve ends above 0.5, the median is "not reached" and must be reported that way, never as the last observed time. Its confidence interval follows Brookmeyer and Crowley (1982): the same rule is applied to the lower and upper bands, so the lower limit comes from the band that falls fastest. When the curve sits at exactly 0.5 over a stretch, the median is the midpoint of that stretch; that is the rule of `quantile.survfit`, and this calculator reproduces it step by step.
The log-rank test (Mantel 1966; Peto and Peto 1972) compares two whole curves, not two percentages at one moment. At each event time it counts how many events would have been expected in one group if the two curves were identical, accumulates the difference between observed and expected, and turns it into a χ² with one degree of freedom. It comes with an important reading condition: it assumes the difference between groups points the same way throughout. If the curves cross, the log-rank loses power and a single hazard ratio stops summarising what happens; in that case compare survival at specific times and say so in the manuscript.
Equations
- events occurring exactly at time t_i
- patients at risk just before t_i (those with time greater than or equal to t_i)
- estimated probability of still being event-free at t
- Greenwood sum accumulated up to t
- normal quantile of the chosen confidence level
- standard error on the log-log scale
- standard error of ln S(t), the square root of the Greenwood sum
- median survival
- first time the curve reaches 50%
- first time the curve drops below 50%
- events observed in the first group
- events expected in the first group if the two curves were identical
- patients of the first group at risk just before t_i
R code
# Kaplan-Meier survival curves and the log-rank test - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(survival)
library(jsonlite)
# The three pasted columns, one row per patient:
tiempo <- c(2, 5, 3, 7, 3, 7, 4, 9, 4, 10, 5, 11, 5, 12, 6, 13, 6, 14, 7, 16, 8, 16,
9, 18, 12, 19, 14, 20, 18, 21)
evento <- c(1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1,
0, 1, 1, 0, 0, 0)
grupo <- c(0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1,
0, 1, 0, 1, 0, 1)
nivel <- 0.95
tipo_ic <- "log-log"
t1 <- 7 # first time of interest; 0 = none asked for
t2 <- 14 # second time of interest; 0 = none asked for
if (length(grupo) == 0) grupo <- rep(0, length(tiempo)) # no group column: a single curve
codigos <- sort(unique(grupo)) # group 1 is the lowest numeric code
d <- data.frame(tiempo, evento, grupo = factor(grupo, levels = codigos))
etiquetas <- levels(d$grupo)
k <- length(etiquetas)
# conf.type is explicit: R defaults to "log", papers usually report "log-log"
sf <- survfit(Surv(tiempo, evento) ~ grupo, data = d, conf.type = tipo_ic, conf.int = nivel)
# survfit drops $strata when the grouping factor has a single level
filas <- if (is.null(sf$strata)) list(seq_along(sf$time)) else split(seq_along(sf$time), rep(1:k, sf$strata))
# Median survival with its CI (Brookmeyer-Crowley; the rule of quantile.survfit:
# first time with S < 0.5, midpoint of the plateau when S is exactly 0.5)
q <- quantile(sf, 0.5)
mediana <- function(g) {
if (g > k) return(c(NA_real_, NA_real_, NA_real_))
if (k == 1) c(q$quantile[[1]], q$lower[[1]], q$upper[[1]]) else c(q$quantile[g, 1], q$lower[g, 1], q$upper[g, 1])
}
# S(t*) with its CI. t = 0 means no time of interest was asked for (S(0) = 1
# is trivial), and past the last observed time of the group S(t) is not
# defined, although extend = TRUE would prolong the curve
s_en <- function(g, t) {
if (g > k || t <= 0 || t > max(d$tiempo[d$grupo == etiquetas[g]])) return(c(NA_real_, NA_real_, NA_real_))
s <- summary(sf, times = t, extend = TRUE)
j <- if (is.null(s$strata)) 1 else which(as.integer(s$strata) == g)
c(s$surv[j], s$lower[j], s$upper[j])
}
# Log-rank test (rho = 0, Mantel-Haenszel). It needs two groups, and it is NOT
# defined when the variance of the statistic is 0: survdiff then either stops
# with a singular system (every event tied at one time, or no group at risk
# when the events happen) or reports "on 0 degrees of freedom" with chisq = 0.
# In both situations chi2, the degrees of freedom and p are NA, not 0 and 1.
chi2 <- NA_real_; gl <- NA_real_; p <- NA_real_
if (k == 2) {
lr <- tryCatch(survdiff(Surv(tiempo, evento) ~ grupo, data = d, rho = 0), error = function(err) NULL)
if (!is.null(lr) && lr$var[1, 1] > 0) {
chi2 <- lr$chisq
gl <- length(lr$n) - 1
p <- pchisq(chi2, gl, lower.tail = FALSE)
}
}
n_de <- function(g) if (g > k) NA_real_ else sum(d$grupo == etiquetas[g])
ev_de <- function(g) if (g > k) NA_real_ else sum(d$evento[d$grupo == etiquetas[g]])
# Life table of one group, as survfit reports it (std.err is the SE of log S)
col <- function(g, campo) if (g > k) numeric(0) else sf[[campo]][filas[[g]]]
res <- list(n_1 = n_de(1), eventos_1 = ev_de(1), mediana_1 = mediana(1),
s_t1_1 = s_en(1, t1), s_t2_1 = s_en(1, t2),
n_2 = n_de(2), eventos_2 = ev_de(2), mediana_2 = mediana(2),
s_t1_2 = s_en(2, t1), s_t2_2 = s_en(2, t2),
chi2_logrank = chi2, gl = gl, p_logrank = p,
t_1 = I(col(1, "time")), riesgo_1 = I(col(1, "n.risk")), ev_1 = I(col(1, "n.event")),
cens_1 = I(col(1, "n.censor")), s_1 = I(col(1, "surv")), ee_1 = I(col(1, "std.err")),
lo_1 = I(col(1, "lower")), hi_1 = I(col(1, "upper")),
t_2 = I(col(2, "time")), riesgo_2 = I(col(2, "n.risk")), ev_2 = I(col(2, "n.event")),
cens_2 = I(col(2, "n.censor")), s_2 = I(col(2, "surv")), ee_2 = I(col(2, "std.err")),
lo_2 = I(col(2, "lower")), hi_2 = I(col(2, "upper")))
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio: survminer::ggsurvplot(sf, risk.table = TRUE)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Survival functions were estimated with the Kaplan-Meier method [1] with 95% CIs based on Greenwood's formula [2] on the log-log of Kalbfleisch and Prentice [6] scale; the median and its CI follow Brookmeyer and Crowley [5] and the curves were compared with the log-rank test [3,4]. Calculations used the "Kaplan-Meier and log-rank" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/kaplan-meier), verified against R's survival package [7] (survfit, quantile.survfit and survdiff).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Kaplan EL, Meier P. Nonparametric estimation from incomplete observations. Journal of the American Statistical Association. 1958;53(282):457–481. doi:10.1080/01621459.1958.10501452 Original source
- 02 Greenwood M. The natural duration of cancer. Reports on Public Health and Medical Subjects No. 33. London: His Majesty's Stationery Office; 1926. Original source
- 03 Mantel N. Evaluation of survival data and two new rank order statistics arising in its consideration. Cancer Chemotherapy Reports. 1966;50(3):163–170. PMID: 5910392 Original source
- 04 Peto R, Peto J. Asymptotically efficient rank invariant test procedures. Journal of the Royal Statistical Society. Series A (General). 1972;135(2):185–207. doi:10.2307/2344317 Original source
- 05 Brookmeyer R, Crowley J. A confidence interval for the median survival time. Biometrics. 1982;38(1):29–41. doi:10.2307/2530286 Original source
- 06 Kalbfleisch JD, Prentice RL. The Statistical Analysis of Failure Time Data. 2nd ed. Hoboken, NJ: John Wiley & Sons; 2002. doi:10.1002/9781118032985 Complementary
- 07 Therneau TM. A Package for Survival Analysis in R. R package survival version 3.8-6. CRAN; 2026. Complementary
- 08 Bland JM, Altman DG. Survival probabilities (the Kaplan-Meier method). BMJ. 1998;317(7172):1572–1580. doi:10.1136/bmj.317.7172.1572 PMID: 9836663 Didactic reading
- 09 Bland JM, Altman DG. The logrank test. BMJ. 2004;328(7447):1073. doi:10.1136/bmj.328.7447.1073 PMID: 15117797 Didactic reading