Statistics are colorless on their own. The same number, depending on how it is presented, can read as "a scientifically confirmed effect" or as "an exploratory reference value." The provisions on statistics and figures in Section 2 of Chapter 1 exist to close that gap. More than which numbers to report, the requirement is to state clearly in what context each number appears and what it can and cannot show — that is how the foreword's three pillars, "treat the package insert as the original, inform accurately without misleading, and keep the record verifiable," take concrete form at the point of statistical reporting.
01Reporting statistical results — what to write and how far to go
When statistical analysis results are included in a material, the analytical method and its results — confidence intervals (= the range within which the estimate is likely to fall) and p-values (= a rough measure of how often a difference this large would arise by chance) — must be stated. The conventional assumption is a two-sided 5% significance level (= the probability cutoff for judging a result "not due to chance"); any deviation from this must be explicitly noted. A bare statement of "statistically significant" leaves the reader no way to check the claim. Writing the method and the level is what makes the result verifiable.
Detailed rule a: when the analysis is adjusted, name the adjusting variables in the results
When covariates — baseline characteristics, prognostic factors, and so on — or stratification factors (= the way subjects are divided into layers, such as by age or severity, for comparison) are used to adjust the analysis, the variables used must be named in the results section. This applies to multivariate regression (= a method that compares groups while subtracting the effects of several factors at once), to stratified Wilcoxon or log-rank tests (= comparing survival or time-to-event separately within each layer), and to all comparable approaches. Without the adjustment variables, the reader cannot tell a crude comparison (= a raw, unadjusted comparison) from an adjusted estimate. When both types sit in the same table, distinguishing them is the writer's responsibility.
Detailed rule b: a statistical model without a standard name requires its equation
If the statistical model used does not carry a commonly recognized name, its equation (= the calculation steps written out as a formula) must be stated in the results. An established name such as "Bayesian hierarchical model" (= a method that estimates probabilistically across layers of data) or "modified Poisson regression" (= a regression that estimates proportions or risk ratios) can be supplemented with a reference. A bespoke model or a non-standard extension, by contrast, is impossible to reproduce or verify from a name alone. Providing the equation is the minimum condition for verifiability.
Detailed rule c: when missing values are imputed, state the imputation method
If missing values (= data that could not be measured or was left unrecorded) are imputed — that is, filled in — before aggregation or analysis, the imputation method must be stated. The estimates shift depending on the method: single imputation (= filling a gap with one value), multiple imputation (= filling gaps with several candidate values so the resulting uncertainty is also captured), last observation carried forward (LOCF = continuing to use the last measured value thereafter), or a mixed-effects model for repeated measures (MMRM = a statistical model that handles all the repeated measurements together). Imputed data without a stated method cannot be distinguished from directly observed data. Making clear where observation ends and estimation begins is what accurate communication requires.
02What statistics can and cannot say
A p-value is only "the probability that, assuming the null hypothesis (= the provisional premise that 'there is no difference') is true, a difference at least as large as observed would arise by chance." It does not directly speak to the size of the effect, clinical importance, or reproducibility. A confidence interval conveys the precision of an estimate but does not by itself guarantee causation (= an actual cause-and-effect relationship). And a nominal p-value — any p-value from analyses other than the pre-specified confirmatory analysis — is liable to arise from the repeated testing of multiple analyses and cannot serve as the basis for a confirmatory conclusion.
Conflating these three makes the same number look like it carries three different weights. A confirmatory p-value from a pre-planned, fixed-hypothesis test, a nominal p-value from a post-hoc search for promising signals, and a confidence interval offered for reference are each assertions of a different strength. Conveying that distinction to the reader is the core of what the Guide means by "present without misleading."
| Statistic | What it can assert | What it cannot assert |
|---|---|---|
| Confirmatory p-value (pre-planned, fixed hypothesis) |
Whether the null hypothesis is rejected or retained at the pre-set significance level. | Effect size, clinical relevance, causation. Claims of safety. |
| Nominal p-value (all analyses outside the pre-planned confirmatory analysis) |
Reference information for hypothesis generation. | Confirmatory conclusions. Asserting "statistical significance." Confirmations of superiority or non-inferiority. |
| 95% confidence interval | Precision of the estimate (width of the interval). Point estimate and range of uncertainty around the effect size. | A guarantee that the true value lies inside. Causal inference. Non-inferiority claims standing alone. |
| Hazard ratio, odds ratio, etc. | Relative magnitude of comparison between groups (within the scope of a planned analysis). | Relative risk reduction framing when no significant difference exists. Evaluative language such as "superior." |
When no statistically significant difference was found, or when no statistical analysis was performed, the result is limited to presenting the numerical values alone. A non-significant result may hide a large absolute difference or may reflect a negligible gap that is clinically irrelevant. Placing the numbers and leaving the judgment to the reader is the boundary the writer must not cross.
03Subgroup analyses — representing their exploratory nature honestly
Most subgroup analyses (= analyses limited to a portion of patients, for example by age or disease type) remain exploratory (= at the "looking for clues" stage rather than a conclusion). This is a statistical reality. Applying the significance level set for the primary analysis to each of several subgroups accumulates the probability of a chance significant finding with each additional comparison — the problem of multiplicity (= repeating many tests inflates the chance of an accidental "hit"). A subgroup analysis that was not pre-planned easily becomes a tool for highlighting a favorable number while carrying the appearance of a significance test.
Hence the conditions the Guide sets. Subgroup analyses may be included only if they were specified in the original trial plan and are scientifically valid. Where the full-population analysis has also been conducted, the subgroup results must be presented alongside those for the full population. Selecting a subgroup result and placing it in the foreground — even where that subset showed favorable numbers — is not permitted.
A "post-hoc subgroup with significant difference" is the textbook case of a nominal p-value. The probability of a chance significant finding accumulates with each additional comparison across unplanned subgroups. The resulting number may be shown as an exploratory reference, but it cannot stand as grounds for claiming "particularly effective in this patient group." Presenting it without stating its position is the misleading presentation the Guide is designed to prevent.
04Graphs and tables — keeping visual presentation from overwriting the numbers
Graphs and tables compress data into an impression, but they also add impressions the numbers themselves do not carry. Moving the y-axis baseline away from zero makes a small difference look large. An arrow draws the eye to the gap. Heavier weight or stronger color on one bar causes the reader to set the other aside before forming a judgment. The Guide devotes substantial attention to the detailed rules for figures precisely because visual misleading is harder to detect than a numerical transcription error.
Name which numerical quantity is being shown
Whether the figure shows a mean, a median (= the middle value when the data are ordered by size), a geometric mean (= an average obtained by multiplying values, suited to ratio data), or a least-squares estimate (= a value adjusted for background factors) must always be stated explicitly. Means and medians diverge in skewed distributions; least-squares estimates are adjusted values. Whether the legend reads "Mean ± SD" (= mean ± standard deviation, the width of variation) or "Median (IQR)" (= median and interquartile range, the spread of the middle 50% of the data) changes the meaning of an otherwise identical bar chart. If the presence or absence of a statistically significant difference is shown, the statistical method used must be named.
Five prohibitions — visual operations that inflate differences
(1) Do not merge data from conditions that are not comparable into the same graph or table. Placing results from trials with different doses or different target populations on a single graph creates the appearance of comparison where there is none. (2) Do not manipulate the scale of the axes beyond what is necessary to exaggerate a difference. (3) Do not use arrows or similar devices to visually emphasize the gap in comparisons between the drug and a control drug (including placebo), or in before-and-after administration comparisons. (4) Do not use text size or color to draw attention to one set of numbers over another. (5) Do not add adjectives that describe the size of the difference without a quantitative basis — words like "markedly," "substantially," or "clearly" are additions the numbers themselves do not justify.
All five prohibitions converge on the same principle: a figure is a substitute for the numbers, not a preview of the interpretation. The writer must not use visual means to guide the reader's conclusion before the reader has had the chance to form one. The rule that only numerical values are to be presented — without evaluative language — when no significant difference exists or no analysis was performed is the same principle expressed on the verbal side.
05Citing original papers — where selective excerption distorts intent
When data are cited from an original paper, the material must be recorded so the content is accurately conveyed, must not excerpt only the portion of the conclusion that is favorable to the company's product, must not distort the true intent of the original paper, and must clearly state the source. This single sentence blocks a layered set of common failures.
The first layer is "accurately conveyed." Cutting out only one figure without its context, or presenting follow-up periods and primary endpoints (= the evaluation measure fixed in advance as the most important) from one trial in a way that suggests they belong to another, violates this requirement. The second layer is "not excerpt only the favorable portion." If the same paper contains an unfavorable subgroup result or a discordant secondary endpoint (= a secondary measure that supplements the primary endpoint), those cannot be withheld while the favorable numbers are shown. The third layer is "not distort the true intent." Citing a trial whose authors concluded "this finding remains exploratory" as evidence of a confirmed effect is falsification of intent.
Stating the source is part of verifiability. A full bibliographic reference — title, journal, year, volume, and pages — lets the reader return to the original and check the whole. A citation presents a part while implying the part represents the whole, and without the reference to let that claim be checked, the implied endorsement is borrowed rather than earned.
What the rules on statistics and figures demand is not merely "numerical honesty." They ask the writer to understand which context produced each number, what it can and cannot assert, and how visual presentation can alter meaning — and to build that understanding into the material itself. Distinguishing a confirmatory p-value from a nominal p-value and a confidence interval; noting whether a subgroup analysis was planned; resisting the temptation to amplify differences through axis manipulation, arrows, or adjectives — each of these closes a specific gap where "factually accurate but misleading" becomes possible.
The same number can carry different weight depending on how the material presents it. The foreword's requirement to "inform accurately without misleading" is tested most directly here.