The Global Fragility Index (GFI) is the minimum number of patients who must be moved between the cells of a clinical trial's contingency table, with the total sample size held fixed, to reverse the trial's statistical significance: a significant result becomes nonsignificant, or a nonsignificant result becomes significant. Unlike the original Fragility Index, which changes outcomes in one treatment arm only, the GFI searches every possible reallocation across all cells, so its value does not depend on the path taken or on how the groups are labeled. Dividing by the sample size gives the Global Fragility Quotient (GFQ = GFI / N). The GFI was defined by Thomas F. Heston in a 2025 SSRN paper.
This Global Fragility Index is a biostatistics measure of clinical trial evidence, not the geopolitical Fragile States Index or any other ranking of fragile countries.
For a 2×2 table with Arm A events a and non-events b, Arm B events c and non-events d, and total N = a + b + c + d:
The moves are a measuring device, not a claim that any error occurred. The GFI is the distance, in patients, between the observed table and the nearest table on the other side of the significance boundary.
Three trials with 100 patients per arm (N = 200):
| Trial {a, b, c, d} | Event rates (A vs B) | Baseline p | GFI | GFQ | Example table after GFI moves | p after |
|---|---|---|---|---|---|---|
| {2, 98, 10, 90} | 2% vs 10% | 0.033 | 1 | 0.005 | {2, 98, 9, 91} | 0.058 |
| {50, 50, 10, 90} | 50% vs 10% | < 0.000001 | 20 | 0.100 | {31, 50, 30, 89} | 0.061 |
| {15, 85, 25, 75} | 15% vs 25% | 0.111 | 2 | 0.010 | {14, 84, 27, 75} | 0.037 |
In the first trial, recording one Arm B event as a non-event removes significance, so 0.5% of the sample separates the verdict from its reversal. The second trial holds until 20 patients, 10% of the sample, are moved. The third is the calculator's home-page example: it is nonsignificant, and moving two Arm A patients into the Arm B event cell makes it significant. Several tables can sit at the same minimum distance; each row shows one of them.
p-values are two-sided Fisher's exact. Each GFI was checked by scoring all 1,373,701 possible 2×2 tables with N = 200.
A p-value alone is partial evidence. The p–fr–nb framework reports three dimensions: significance (p), fragility (fr), and robustness (nb). For 2×2 and multinomial tables the GFQ is the recommended fragility metric, paired with the Risk Quotient (RQ), which measures distance from neutrality. When a very large sample makes the exact GFI search intractable, the Modified-Arm Fragility Quotient (MFQ) is the fallback. See the full documentation.
The GFI measures how far a trial's significance verdict sits from being reversed, counted in patients. A GFI of 1 means that moving a single patient between two cells of the table moves the result across the p = 0.05 threshold; a large GFI means the verdict holds under substantial changes to the data. It is a count, not a probability.
The Fragility Index (Walsh et al., 2014) switches outcomes between event and non-event within one treatment arm only. The GFI allows a patient to be moved between any two cells of the table, including across arms, while the total sample size stays fixed. Because every Fragility Index move is also a GFI move, the GFI is never larger than the Fragility Index when both use the same significance test, and it does not depend on which arm is chosen or how the groups are labeled.
The Global Fragility Quotient is the GFI divided by the total sample size: GFQ = GFI / N. It expresses fragility as the proportion of the sample that must be reallocated to reverse significance, so a GFQ of 0.03 means 3% of patients. The GFQ allows comparison across trials of different sizes and is the recommended fragility (fr) metric for 2×2 and multinomial tables in the p–fr–nb framework.
For 2×2 tables, the GFI always uses the two-sided Fisher's exact test at α = 0.05, whatever test the original trial reported, and the same test scores every rearranged table. For larger tables, the Fisher–Freeman–Halton exact test is preferred, with Pearson's chi-square as the fallback when exact computation is intractable and expected counts are adequate. GFI values computed under different tests should not be compared directly.
Yes. For a nonsignificant result, the GFI is the minimum number of moves that makes the result significant. One definition covers both directions, so no separate reverse fragility index is needed.
No. Resampling fragility, the probability that a new sample of the same size would reverse the verdict, is strongly associated with the p-value. The GFI measures perturbation fragility instead: the number of recorded outcomes that must change before the observed table crosses the threshold. In three pairs of published trials matched on p-value, the GFQ differed 21-fold, 30-fold, and 2-fold within pairs (Heston, Internet Medical Journal, 2026).
No. The Global Fragility Index described here is a statistical measure of clinical trial results. It has no connection to the Fragile States Index, other geopolitical rankings of state fragility, or the U.S. Global Fragility Act.
Enter the four cell counts of a 2×2 table in the free calculator at fragilitymetrics.org. It reports the GFI and GFQ alongside the other p–fr–nb metrics, using an exact search scored by Fisher's exact test. The open-source Fragility Metrics Toolkit on GitHub and Zenodo (doi:10.5281/zenodo.17254763) is the reference implementation.
Thomas F. Heston, MD, of the University of Washington and Washington State University, defined the GFI in “The Global Fragility Index: A Path-Independent Measure of Statistical Fragility” (SSRN, 2025). The paper defines the index for multinomial tables and proves that the global cell-move distance to the significance boundary is path-independent.
Cite the GFI as: Heston TF. The Global Fragility Index: a path-independent measure of statistical fragility. SSRN. 2025. doi:10.2139/ssrn.5709162
Last updated September 12, 2026.