Complete Statistical Evidence Calculator

Calculate p-fr-nb triplets for time-to-event outcomes reported as a hazard ratio

Calculate CSE for: 2x2 Trials | Survival

Survival (SFQ/SRQ) Calculator

Enter a published hazard ratio and its confidence interval to calculate complete statistical evidence (significance, fragility, and robustness).

Hazard ratio Lower limit Upper limit CI level (%)
Example: Immunotherapy Trial

A trial reporting overall survival for an immune checkpoint inhibitor versus standard therapy:

Enter: HR = 0.70, lower = 0.55, upper = 0.89, CI level = 95

Returns p = 0.003674, SFQ = 0.4858, SRQ = 0.2629. The result is significant. Its z-statistic (|z| = 2.90) sits 0.94 beyond the 1.96 significance boundary, which SFQ rescales to 0–1; a higher SFQ means a more stable classification. Its log hazard ratio sits 0.36 from neutrality (HR = 1), which SRQ rescales to 0–1; a higher SRQ means farther from neutrality.

If the paper reports its interval at a level other than 95% (an interim analysis with alpha spending may report 97.73% or 99%), enter that level instead. The standard error is derived from the level you enter.

Why Complete Evidence?

A hazard ratio with a p-value is partial evidence. It says whether an effect cleared the significance threshold, not how close it sits to that threshold, and not how far it sits from no effect at all. Complete statistical evidence requires three complementary dimensions:

Only when all three dimensions align do you have truly convincing, replication-ready evidence. Clinical decisions also need the absolute effect size, such as months of survival gained: the triplet plus effect size is complete evidence.

SFQ needs the hazard ratio and its confidence interval, because the interval is what carries the standard error. SRQ needs the hazard ratio alone. Neither requires raw survival data, Kaplan-Meier curves, censoring information, or patient-level data.

Read Full Documentation

Note: SFQ assumes the published interval is a symmetric Wald interval on the log scale; SRQ uses the hazard ratio alone and makes no assumption about the interval. Compare the recovered p-value against the one printed in the paper: a close match confirms the assumption, and a visible discrepancy means the interval was stratified, exact, or group-sequential and the inputs need checking. SFQ anchors on z = 1.96 by definition, so for a trial stopped at an alpha-adjusted interim boundary it measures distance to the conventional boundary rather than the stricter one the trial had to clear. Delayed curve separation, common in immunotherapy trials, makes the hazard ratio an average over time; SFQ and SRQ inherit whatever the hazard ratio means in that setting. The framework has not published numeric fragility cutoffs (they are under validation), and its numeric robustness bands (near, intermediate, far) are defined for RQ only, so no band labels are applied to SFQ or SRQ.

Citation: Heston TF. Fragility metrics toolkit v6.0.0. Zenodo. 2026. DOI: 10.5281/zenodo.17254763