Statistics are a compressed language for facts. The same number carries a different evidentiary weight depending on whether it came from "a pre-specified analysis confirming a single pre-fixed hypothesis" or from "a promising trend discovered after seeing the results." When the definitions of terms are not held precisely, the same words can be used to smuggle in claims of very different strength. This section defines each scientific and statistical term that appears in Chapter 1 — in the order of definition, what it means, and where misreading tends to occur.
The three pillars of the Guide — pillar ① (close off misunderstanding) and pillar ③ (keep the record verifiable) — begin with accurate use of statistical terms. Even a single phrase like "statistically significant" can produce opposite impressions depending on the type of analysis and context behind it.
| Term | Core meaning | Misreading / caution |
|---|---|---|
| Confirmatory analysis 検証的解析 |
An analysis designed to test a single hypothesis fixed before the trial began. The result can serve as the basis for a statistical conclusion. | "We ran the analysis and it looked promising" is not confirmatory. Pre-specification and hypothesis fixation are mandatory conditions. |
| Exploratory analysis 探索的解析 |
An analysis carried out to generate hypotheses. It points toward what should be confirmed in the next trial. | "Exploratory but the trend is clear" is not confirmation. No matter how small the p-value, the result cannot be used to confirm a claim. |
| p-value p値 |
The probability that, assuming the null hypothesis is true, a difference at least as large as observed would arise by chance. | It does not speak to effect size, clinical importance, or causation. A small p-value does not mean a large effect. |
| Nominal p-value 名目上のp値 |
Any p-value obtained from analyses other than the pre-specified confirmatory analysis. | "Unadjusted for multiplicity" or "reference value" are incorrect substitutes. The definition is about the position of the analysis in the trial plan, not about adjustment. |
| 95% confidence interval 95%信頼区間 |
An interval that, under repeated application of the same procedure, would contain the true value 95% of the time. Conveys the precision of the estimate. | It does not mean "there is a 95% probability the true value is inside." It does not guarantee causation and cannot support a non-inferiority claim on its own. |
| Hazard ratio ハザード比 |
The ratio of the rate of event occurrence between groups at any given time. Commonly used in survival analyses. | "Risk reduced by X%" conflates a hazard ratio with an absolute risk difference. Risk reduction language is not permitted when no significant difference exists. |
| Meta-analysis / systematic review メタ解析 / システマティックレビュー |
A method for systematically collecting and integrating multiple studies according to a pre-defined protocol. | Without stating search sources, search terms, and inclusion/exclusion criteria, there is no way to distinguish systematic collection from selective aggregation. Methodological transparency is a requirement. |
| Post-hoc analysis 事後解析 |
An analysis planned and carried out after seeing the results. In principle, exploratory in character and not usable for confirmation. | "p < 0.05 even in the post-hoc analysis therefore significant" does not hold. An analysis designed after seeing the results cannot inhabit a confirmatory framework. |
| Subgroup analysis サブグループ解析 |
An analysis restricted to a subset of the full population. Most are exploratory. | Results must be presented alongside those for the full population. Extracting subgroup results and placing them in the foreground is not permitted. Pre-specification status must be stated. |
| Type of summary statistic 統計量の種別 |
Mean, median, geometric mean, and least-squares mean each rest on different assumptions. | Without stating the type in the legend, the reader cannot tell them apart. Adjusted figures (least-squares means) and crude figures must not be mixed without labeling. |
| Two-sided 5% significance level 両側5%の有意水準 |
The conventional threshold for rejecting the null hypothesis. α = 0.05. | Any deviation from this default must be explicitly stated. "Statistically significant" without a stated significance level is not verifiable. |
01Confirmatory and exploratory analysis — the classification is decided before the trial begins
A confirmatory analysis is one in which a single hypothesis is fixed before the trial starts, and the analysis is designed specifically to test that hypothesis. The hypothesis, the primary endpoint, the testing method, and the significance level are all recorded in the protocol before a single patient is enrolled. The p-value that results from this analysis can be compared to the pre-set significance level to determine whether the null hypothesis is rejected or retained. This is the only category of analysis whose results can serve as grounds for a statistical conclusion.
An exploratory analysis is conducted to generate hypotheses. It is performed either without a prior hypothesis or to answer a question that was not planned in the protocol. It can identify promising directions for future research, but that is its function — pointing toward what should be confirmed in a subsequent trial. No matter how small the p-value, an exploratory result is not confirmation. Confirmation requires a new independent trial in which that hypothesis is designated as the primary pre-specified hypothesis.
The confusion most often arises in the situation of "we did an exploratory analysis and got p < 0.05." The size of the p-value does not alter the classification of the analysis. The classification is determined at the time of trial planning and cannot be changed once results are known.
02The p-value — what it asserts and what it does not
The definition of a p-value is one thing: "the probability that, assuming the null hypothesis is true, a difference at least as large as observed would arise by chance." No claim that falls outside this definition is warranted by the p-value.
Three things a p-value cannot speak to. First, the size of the effect. p = 0.001 does not mean a larger difference than p = 0.04. With a large sample, a clinically negligible difference can produce p = 0.001. Second, clinical relevance. Whether a statistically significant difference matters to patients is a separate question. Third, causation. A p-value from an observational study cannot control for confounders. The reading "p < 0.05 therefore there is an effect" exceeds what the definition supports.
A 95% confidence interval supplements the p-value by conveying a "range" the p-value does not provide. Rather than a point estimate alone, it shows the spread of uncertainty around the estimate. It is not, however, a guarantee of causation, nor a statement of the probability that the true value falls within the interval. The width of the interval reflects the precision of the estimate — wider when data are sparse, narrower when data are abundant.
03Nominal p-value — the definition, word for word
The definition of a nominal p-value is: any p-value obtained from analyses other than the pre-specified confirmatory analysis.
The critical phrase is "other than the pre-specified confirmatory analysis." This is not a question of the type of analysis or the statistical method used. It is a question of where the analysis stands in the trial plan. Secondary endpoint analyses, post-hoc subgroup analyses, sensitivity analyses, exploratory biomarker analyses — all of these lie outside the pre-specified primary confirmatory analysis. Every p-value obtained from them is a nominal p-value.
The reason a nominal p-value cannot support a confirmatory conclusion is multiplicity. When multiple analyses are repeated, the probability that at least one produces a chance significant result accumulates with the number of tests. Without having fixed a single hypothesis in advance as the confirmatory test, a result of "this analysis was significant" cannot be distinguished from a chance finding.
A common mischaracterization deserves specific note. Describing a nominal p-value as "a p-value not adjusted for multiplicity" or "a reference value before multiple comparison correction" is inaccurate. The defining feature of a nominal p-value is not the absence of multiplicity adjustment — it is that the analysis lies outside the pre-specified confirmatory analysis. Applying a multiplicity correction after the fact does not change a nominal p-value into a confirmatory one.
Including a nominal p-value in a promotional material is not prohibited. It may be presented as reference information for hypothesis generation. What it cannot do is serve as grounds for a confirmatory conclusion, support a declaration of "statistically significant difference," or confirm superiority or non-inferiority. When a p-value is reported, the context must make clear which category of p-value it is.
04Hazard ratio — a ratio of relative rates, not an absolute reduction
A hazard ratio is the ratio of the rate of event occurrence between two groups at any given point in time. A hazard ratio of 0.7 means the event rate in the treatment group is 0.7 times that in the control group at any given moment. It is widely used in survival analyses and is typically estimated from a Cox proportional hazards model.
Two misreadings to avoid. The first is the translation "risk reduced by 30%." A hazard ratio of 0.7 is a relative ratio, not an absolute risk difference. The absolute risk difference requires a separate calculation. The second is describing a direction of effect when no significant difference was found. When no statistically significant difference exists, using the numerical value of the hazard ratio to assert a "trend toward risk reduction" or "30% improvement" is not permitted. The result is limited to presenting the value, without evaluative language.
05Meta-analysis and systematic review — transparency as a requirement
A systematic review collects and evaluates studies relevant to a defined question according to a pre-specified search strategy and eligibility criteria. A meta-analysis statistically pools the results from the collected studies to produce an overall estimate, and is often conducted as part of a systematic review.
When citing these in promotional materials, certain elements must be stated. Which databases were searched (search sources), what search terms and combinations were used (search strategy), and on what criteria studies were included or excluded (eligibility criteria with the reasoning) — without all three, the claim that the collection was "systematic" cannot be verified. The transparency of the search is the minimum condition for the reader to assess selection bias.
06Post-hoc analysis and subgroup analysis — declaring the exploratory character explicitly
Post-hoc analysis refers broadly to analyses that were planned and conducted after seeing the trial results. Knowing the outcome before designing the test creates room for conscious or unconscious selection — choosing the analysis that shows the most favorable result. This is the fundamental reason why post-hoc analyses cannot be used for confirmation: the structure has become "form a hypothesis from the results" rather than "fix a hypothesis, then confirm it."
A subgroup analysis restricts the analysis to a subset of the full population. Most remain exploratory. Applying the significance level from the primary analysis to each of several subgroups causes the probability of a chance significant finding to accumulate with each additional comparison. Subgroup analyses may be included in materials only if they were pre-specified in the original trial plan and are scientifically valid. Where the full-population analysis was also conducted, the subgroup results must be presented alongside the full-population results. Selecting a favorable subgroup result and placing it in the foreground — even where the subset showed good numbers — is not permitted.
07Types of summary statistic — a single word in the legend changes the meaning
Even under the umbrella of "average," different summary statistics produce different numbers and carry different assumptions.
The mean is the sum of all values divided by the number of observations. It is sensitive to outliers and can diverge from the typical value when the distribution is skewed. The median is the middle value when data are arranged in order. It is resistant to outliers and, in skewed distributions, often a better representation of the typical value than the mean. The geometric mean is computed by taking the arithmetic mean on the log scale and back-transforming. It is appropriate for data that follow a log-normal distribution, such as pharmacokinetic parameters like AUC and Cmax. The least-squares mean (LS mean) is a model-based estimate after adjustment for covariates. It is an adjusted figure, not a crude observation, and this must be stated clearly — otherwise it is indistinguishable from directly observed data.
Without stating the type of statistic in the figure legend, the reader has no way to tell which is which. Whether the legend reads "Mean ± SD" or "Median (IQR)" changes the interpretation of an otherwise identical bar chart. When reporting the presence or absence of a statistically significant difference, the statistical method used must also be named.
08Two-sided 5% significance level — the default, not the only permissible value
The significance level α = 0.05 (two-sided 5%) is the conventional default widely adopted in pharmaceutical clinical trials. It sets at 5% the probability of observing a result this extreme when the null hypothesis is true. A two-sided test assumes that the treatment group could differ from the control in either direction, and is the standard approach where the possibility of a difference in either direction is not ruled out a priori.
This is a default, not the only permissible setting. Extension trials, safety evaluations, dose-finding studies, and other designs may use a different significance level for clearly stated scientific reasons. In those cases, the significance level actually used must be stated in the material. A bare statement of "statistically significant" without a stated significance level gives the reader no way to check the claim. Stating the level used is the minimum condition for a verifiable record.
Each of the eleven terms defined here is a term that cannot be replaced by a similar-sounding substitute. The nominal p-value is the clearest example: paraphrasing it as "unadjusted for multiplicity" strips away the core of the definition — that the analysis lies outside the pre-specified confirmatory analysis. The boundary between what a statistic can and cannot assert is maintained only when the terms are held to their precise definitions. Presentation that does not mislead begins with the choice of accurate words.