4  Inference and Hypothesis Testing

Author

Mallory L. Barnes

Modified

September 18, 2026

Data and data sets are not objective; they are creations of human design. We give numbers their voice, draw inferences from them, and define their meaning through our interpretations.

Kate Crawford

Inferential statistics is the branch of statistics concerned with drawing conclusions about populations or causal relationships from sample data. It provides the tools to answer questions like: did our manipulation actually cause a change, or could the pattern we observed have arisen by random chance alone?

This chapter does that in two parts. The first part builds the logic by simulation, with no formulas: we work out what chance alone can produce, then compare a real result against it. The rest of the chapter replaces the simulation with the standard hypothesis testing machinery, because once you can assume a known reference distribution you no longer have to rebuild the chance window from scratch every time.

4.1 What chance alone can produce

An experiment asks whether a manipulation causes a change in what we measure. Some manipulations are so overwhelming that no statistics are required to see them. Flip a light switch up and the light comes on; flip it down and it goes off, every time. You don’t need a \(t\)-test to tell you that flipping the light switch is the cause of the change in brightness.

Nothing in environmental science behaves like this. Effects are real signals buried in noisy, variable measurements, where flipping the switch changes the outcome on average. Not in every instance, and not by the same amount every time.

Random chance can produce apparent differences between groups even when no real difference exists. We have already seen this with sampling: two samples from the same distribution come out differently. The same principle applies when we compare group means. This is a fundamental problem for inference.

Consider a scenario where we expect to find no real difference. A researcher surveys 10 caves for white-nose syndrome, a fungal disease that has killed millions of North American bats, capturing 20 bats at each cave and counting how many show visible signs of infection. Figure 4.1 lays out the design.

Study design schematic. IV: group assignment, arbitrary. Two boxes labelled Group A and Group B, each with n = 5 caves. DV: infected bats out of 20 captured.
Figure 4.1: The bat-cave scenario study design. The split into two groups is arbitrary, which is the whole point: there is no real difference for the test to find.

The key feature is that all 10 caves are drawn from the same population. There is no treatment, no environmental gradient, nothing systematic separating them, so the true infection rate is identical everywhere. Splitting them into two groups of 5 is arbitrary. Any difference we see between those groups therefore has to be sampling error, which makes this the cleanest possible look at what chance alone can manufacture.

Figure 4.2 shows one run of this scenario: the mean number of infected bats in each group.

Bar chart with two bars, Group A and Group B, showing the mean number of infected bats. The y-axis is zoomed to run from 6 to 10, so a modest difference between the two bars looks larger than it would on a full 0-to-20 scale.
Figure 4.2: Mean number of infected bats per site for two groups drawn from the same population. Note the zoomed y-axis, which does not start at 0.

The bars are not the same height: Group A and Group B have different mean counts. But both groups were drawn from the very same population of sites, so this isn’t a real difference in infection risk, it’s sampling error.

This is the core inference problem: differences appear even when nothing real caused them. How can we tell whether an observed difference is real or just the result of chance?

Chapter 2 showed that chance alone can produce spurious correlations. The same is true for group means, and it can be pinned down: run the same arbitrary split many times and record what comes out. We summarize each run as a single difference score: the mean count for Group A minus the mean count for Group B. Figure 4.3 shows the difference scores for 10 replications of the survey, each one sampling all 10 caves from the same population, splitting them 5 and 5, and computing the means.

Bar chart of the A-minus-B difference score for each of 10 replications, with a horizontal reference line at zero. Bars extend above and below the line by varying amounts, some replications favoring Group A, others favoring Group B.
Figure 4.3: Difference in mean infected count between Group A and Group B across 10 replications. The true difference is zero, but sampling error produces non-zero differences each time.

A bar at zero means the two groups had the same mean count in that replication. Positive values indicate Group A was higher; negative values indicate Group B was higher. Chance produces a new difference every time, but not an unlimited one. To see which sizes are common and which are rare, we run 10,000 replications and plot the difference scores as a histogram.

Histogram of 10,000 A-minus-B difference scores, tallest at zero in the center and tapering off symmetrically toward both larger positive and larger negative differences, with almost nothing beyond roughly plus or minus 3.
Figure 4.4: Histogram of mean count differences between Group A and Group B over 10,000 replications. All differences arise from chance alone. The most frequent difference is near zero; larger differences occur less often. Chance cannot produce everything.

This histogram is the chance window for this study design. It shows that chance most often produces a difference near zero, the tallest bar in the center. Larger differences in either direction occur less and less frequently. Differences beyond about \(\pm 3\) are possible, but rare.

We can use this window to evaluate any observed difference. If a study found a difference of 1 infected bat between two groups, that falls well inside the chance window: differences of that size occur frequently by random sampling alone. If the study found a difference of 6 bats, that would fall far outside the window; chance produced a difference that large 0 times out of 10,000. When an observed result almost never happens by chance, we have grounds to conclude that something other than chance produced it.

NoteNote: Which direction?

Still probability, model → data. The chance window starts from a world we chose, every cave at the same infection rate, and shows the differences that world produces. The 1-bat and 6-bat results above are hypothetical, used to show what the window can tell us. No real study’s data have entered yet.

4.2 Chance vs. not chance

The window tells us what chance can do. It does not tell us where to draw the line between “chance could have” and “chance didn’t.” That is a judgment, not a calculation.

Figure 4.5 draws lines on the same bat histogram from above. The boundaries are the smallest and largest differences chance actually produced across the 10,000 replications.

The same histogram of 10,000 difference scores, now overlaid with colored zones. A grey band across the middle, labeled CHANCE, covers the full range chance actually produced. Narrow gold slivers sit at the two edges of that grey band, and red zones beyond them on both sides, labeled NOT CHANCE, mark differences chance never produced in any of the 10,000 replications.
Figure 4.5: Decision boundaries applied to the histogram of mean count differences from Figure 4.4. Grey marks the range chance did produce across 10,000 replications; gold marks the edges of that range, which chance reached only rarely; red marks the differences chance never produced at all.

Across the 10,000 replications, chance produced differences from about -5.6 to 5.4 bats.

  • Chance (grey). Chance produced differences like this routinely. No grounds for suspicion.
  • Ambiguous (gold). The edges of what chance produced. Chance can land here, but rarely: probably not chance, but not ruled out.
  • Not chance (red). Chance never produced a difference this large in 10,000 tries. Real grounds to suspect something beyond sampling error.

4.2.1 How rare is rare enough?

Question: How many times does something need to happen for it to happen “a lot”? How rare does something need to be before you stop worrying about it?

Environmental scientists already have a working vocabulary for this: the 100-year flood. A “100-year flood” doesn’t mean floods that size happen once a century on a schedule; it means that in any given year, there’s about a 1-in-100 chance of a flood that severe. Would you build a house in that floodplain?

Many people do, and often regret it, because a 1% annual chance still adds up over a 30-year mortgage (roughly a 26% chance of at least one such flood in that span). A location with a 1-in-10,000 annual flood risk is very different: over the same 30 years, the cumulative chance is under 0.3%.

That’s the intuition we need going forward: rare events still happen, but how rare something is changes what you’re willing to conclude from seeing it.

Where the gold zone ends is a judgment call. A stricter line means fewer false alarms and more missed real effects; a looser line means the reverse. Science standardizes that call with a convention, \(\alpha = .05\), defined later in this chapter. It is about 1 in 20, far looser than the red zone’s never in 10,000.

4.3 Hypothesis tests

4.3.1 The null and alternative hypotheses

The chance window described a world where nothing but chance is at work. A formal test gives that world a name. In the bat example, “nothing but chance” meant no real difference between groups. From here on we compare a single sample against a specific value, such as a regulatory limit or a historical average.

The null hypothesis, \(H_0\), is the claim that the population mean equals that specific value, written \(\mu_0\). If \(H_0\) is true, any gap between the sample mean and \(\mu_0\) is sampling error.

The alternative hypothesis, \(H_a\), is the claim that the population mean does not equal \(\mu_0\). It can also point in one direction only, larger or smaller; Section 4.5 covers when that is justified.

\[H_0:\ \mu = \mu_0 \qquad H_a:\ \mu \neq \mu_0\]

Null value notation. In this chapter, hypotheses are about the population mean \(\mu\), compared with the constant \(\mu_0\). The sample mean \(\bar{x}\) is never part of \(H_0\) or \(H_a\); it is the evidence we bring to test them.

4.3.2 From simulation to a formal test

Everything so far was built by brute force: thousands of replications, piled into a histogram, with hypothetical results laid against the pile. Nobody wants to rebuild a chance window for every question. Chapter 3 already supplies a shortcut: sample means have a predictable distribution of their own: centered on the population mean, with spread \(\sigma/\sqrt{n}\), and close to normal once the sample is large enough (the Central Limit Theorem). That lets us replace the simulated chance window with a curve whose shape we already know. The logic does not change, only the bookkeeping does.

NoteNote: Which curves have to be normal

The tests in this chapter and the next two read their answers off a normal curve, or its close relative, the \(t\) curve. Two different reasons put a normal curve there.

  • For the sampling distribution of the sample mean, it is the Central Limit Theorem at work: whatever the population looks like, sample means pile up in a roughly normal shape once the sample is large enough.

  • For a population, it is not required. A test reads its answer off the sampling distribution of the sample mean, not off the population. The population’s shape only matters when the sample is small, because the Central Limit Theorem needs a large enough sample to promise a normal shape for sample means. With a small sample, the normal curve for \(\bar{x}\) rests on the population itself being roughly normal, and the only check we have is the sample: a histogram with no strong skew and no extreme outliers.

When the sample says the normal assumption is not reasonable, the nonparametric tests in Chapter 7 take over.

4.3.3 The players

A test is about \(\mu\), the true population mean, which we never observe; everything we know about it comes from one sample. The null hypothesis proposes a value for it, \(\mu_0\). To reject \(H_0\), the sample has to be hard to explain if the null were true: its mean, \(\bar{x}\), has to land farther from \(\mu_0\) than sampling error would plausibly carry it.

Several distributions and values take part in that comparison. Table 4.1 lists every one, with the line each is drawn with in this chapter’s figures. They fall into three groups: what we would expect if \(H_0\) were true, what the sample gives us, and the truth we never see.

Table 4.1: The players in a hypothesis test, and the line each is drawn with in this chapter’s figures.
Line Player What it is
If H0 were true \(\mu_0\) The population mean if \(H_0\) were true
Null population distribution The population if \(H_0\) were true: centered at \(\mu_0\), as wide as the population
Sampling distribution of the sample mean under \(H_0\) Where \(\bar{x}\) would land across many samples if \(H_0\) were true: centered at \(\mu_0\), spread = standard error
From the sample \(\bar{x}\) The sample mean: the one value we observe
Estimated population distribution Centered at \(\bar{x}\), with the sample's standard deviation as its spread: our best picture of the true population
Sampling distribution of the sample mean Where \(\bar{x}\) would land across many samples: centered at \(\bar{x}\), spread = standard error
Confidence interval \(\bar{x} \pm 1.96\) standard errors, built from the sampling distribution of the sample mean
The truth True population distribution The population the data actually came from. Never known. \(H_a\) is a claim about where it sits

4.3.4 Seeing them together

The players are easier to hold onto as a picture than as a table. Let’s build that picture a curve or two at a time.

The population curves in these figures are drawn as normal for clarity; the note in the previous section explains why a test does not need them to be.

Notation reminder. Chapter 3’s shorthand for a normal curve: \(X \sim N(\mu, \sigma^2)\) reads “the random variable \(X\) is normally distributed with mean \(\mu\) and variance \(\sigma^2\).” A normal curve needs just those two numbers, a center and a spread. The second number is the variance, the square of the standard deviation.

1. The null population distribution.

Start with the null population distribution: the population we would have if \(H_0\) were true. Its center, \(\mu_0\), comes from \(H_0\) itself. Its spread is the population standard deviation, \(\sigma\). In practice \(\sigma\) is almost never known and the sample’s standard deviation stands in for it; this chapter treats \(\sigma\) as known as a simplification. The \(t\)-tests in Chapter 6 are the adjustment used in practice.

A single blue bell-shaped curve centered on a magenta dashed vertical line. Text reads H sub 0 colon mu equals mu sub 0.
Figure 4.6: Null population distribution: the population we would have if \(H_0\) were true, centered at \(\mu_0\).

2. Null vs. true population.

In an idealized world where \(H_0\) is exactly true, the true population and the null population distribution are identical. In reality, we never know the true population.

A single bell-shaped curve, since the orange true population curve and the blue null population distribution curve sit exactly on top of each other, both centered at the same point marked with a dashed vertical line. Text reads H sub 0 colon mu equals mu sub 0.
Figure 4.7: True population distribution equals the null population distribution (\(H_0\) true): \(\mu=\mu_0\).
NoteNote: Which direction?

The switch to statistics, data → model. Steps 1 and 2 came from \(H_0\), a model we proposed. Step 3 starts from the sample we actually have and works back toward the population it came from.

3. The sample’s picture of the truth. We never see the true population, but we do have a sample. From the sample’s mean, \(\bar{x}\), and its standard deviation we build the estimated population distribution: our best picture of the true population, drawn here as a normal curve. With a representative sample, the two line up closely.

Two wide bell curves nearly on top of one another: the solid orange true population and the light-green dashed estimated population.
Figure 4.8: A representative sample: the estimated population distribution, built from the sample, lines up closely with the true population distribution.

4. From individual readings to sample means. By the Central Limit Theorem, the sampling distribution of the sample mean is close to a normal curve once the sample is large enough, so it needs only a center and a spread. Its center is \(\bar{x}\), the same as the estimated population. Its spread is the standard error: the sample’s standard deviation divided by \(\sqrt{n}\). (Not \(\sqrt{n-1}\): the \(n-1\) already went into computing the standard deviation itself, back in Chapter 3.) In the shorthand, it is \(N(\bar{x}, SE_{\bar{x}}^2)\): same center, much narrower. Figure 4.9 shows the Central Limit Theorem at work: make the population right-skewed, and the sampling distribution of the sample mean keeps the same normal shape and the same spread.

Animation. A wide light-green dashed bell curve centered on a dotted line labeled x-bar, labeled estimated population, center x-bar, spread s. A much narrower dark-green dot-dash curve appears at the same center, labeled sampling distribution of the sample mean, N of x-bar and SE squared. The wide curve then gradually becomes right-skewed, with a long right tail, while the narrow dot-dash curve stays exactly the same; the final label reads same sampling distribution of the sample mean, once n is large enough.
Figure 4.9: From individual readings to sample means. The sampling distribution of the sample mean (dark-green dot-dash) shares the estimated population’s center, \(\bar{x}\), with the standard error as its spread. When the population turns right-skewed, the sampling distribution keeps the same normal shape and spread, once the sample is large enough.

5. The same move on the null side. The sampling distribution of the sample mean under \(H_0\) is centered at \(\mu_0\), and its spread is again the standard error, the null population’s standard deviation divided by \(\sqrt{n}\) (with \(\sigma\) treated as known, the same standard error as in step 4): \(N(\mu_0, SE_{\bar{x}}^2)\). It describes where \(\bar{x}\) would land across many samples if \(H_0\) were true, which makes it the formal version of the chance window. The critical values later in the chapter come from it.

A wide solid blue bell curve and a much narrower, taller blue dot-dash bell curve, both centered on a magenta dashed line.
Figure 4.10: The same move on the null side: the null population distribution (blue solid) and the sampling distribution of the sample mean under \(H_0\) (blue dot-dash) share a center, \(\mu_0\). Dividing the spread by \(\sqrt{n}\) makes the second much narrower.

All of them at once. Put every player on one set of axes and the picture gets crowded fast, which is why the rest of the book shows only the pieces a given question needs. Figure 4.11 adds them one at a time, grouped as in Table 4.1, and ends with all of them together.

Animation. Curves appear one at a time on shared axes, each with a colored label: a magenta dashed line for mu sub 0; a wide blue null population curve; a narrow blue dot-dash sampling distribution under H sub 0; a black dotted line for x-bar; a wide light-green dashed estimated population; a narrow dark-green dot-dash sampling distribution of the sample mean; a wide orange true population. The last frame shows all of them together, crowded.
Figure 4.11: The players, added one at a time: first everything built from \(H_0\), then everything built from the sample, then the true population, ending with all of them on one set of axes.

The two curves a decision uses. Strip the picture back to the two sampling distributions of the sample mean (Figure 4.12). Everything in the next section runs on these two: critical values and the \(p\)-value read off the one built from \(H_0\), and the confidence interval comes from the one built from the sample.

Two narrow dot-dash bell curves of the same width on shared axes: a blue one centered on a magenta dashed line labeled mu sub 0, and a dark-green one centered on a black dotted line labeled x-bar, to its right.
Figure 4.12: The two curves a decision uses: the sampling distribution of the sample mean under \(H_0\) (blue dot-dash), centered at \(\mu_0\), and the sampling distribution of the sample mean from the sample (dark-green dot-dash), centered at \(\bar{x}\). Both have the standard error as their spread.

4.4 Making the decision

4.4.1 Terms for decision rules

When running a hypothesis test, four terms guide our decision. The first is the quantity every test computes.

Test statistic. A test statistic is a ratio. On top goes the effect you observed, the distance between what you measured and what \(H_0\) predicted. On the bottom goes an estimate of how much that distance could move around by chance alone:

\[\text{statistic}=\frac{\text{effect}}{\text{error}}\]

That is the whole idea, and every test in this book is a version of it. A difference of 2 units is not impressive or unimpressive on its own. It is impressive only relative to how much difference chance produces for a study of this size, which is exactly what the chance window measured earlier in this chapter. The test statistic is that comparison written as a single number: how many units of ordinary noise is this effect worth?

Large statistic means the effect is big compared to the noise. Small statistic means it is not. Everything after this is bookkeeping about where to draw the line.

Significance level (\(\alpha\)). Chosen before the test. \(\alpha\) is the long-run probability of a Type I error (rejecting a true \(H_0\); more on error types next chapter). A common choice is \(\alpha = 0.05\). Smaller \(\alpha\) makes tests more cautious and produces wider confidence intervals.

\(p\)-value. Assuming \(H_0\) is true, the \(p\)-value is the probability of seeing a test statistic at least as extreme as what we observed. Small \(p\) means our data look unusual under \(H_0\).

Confidence level. Confidence level is the long-run proportion of confidence intervals that contain the true parameter when we repeat the entire data collection and interval procedure.

\[\text{Confidence Level} = 1 - \alpha\]

With \(\alpha=0.05\), about 95% of intervals constructed this way will capture the true value.

CautionWarning: Don’t misread a confidence interval

Do not say:

“There is a 95% probability this interval contains \(\mu\).”

Instead say:

“We used a procedure that captures \(\mu\) about 95% of the time in the long run.”

4.4.2 Three equivalent ways to test \(H_0\)

All three routes below run on the same quantity: the test statistic defined above.

  1. Test statistic / critical values. Choose \(\alpha\), find the corresponding critical value(s) for your test statistic, and compute your observed statistic. Reject \(H_0\) if the statistic falls in the rejection region.

  2. Confidence interval (two-sided). Build a \((1-\alpha)\) CI for the parameter. Reject \(H_0\) if \(\mu_0\) lies outside the CI; fail to reject if it lies inside.

  3. \(p\)-value. Compute the \(p\)-value for your observed statistic under \(H_0\) and compare to \(\alpha\). Reject \(H_0\) if \(p<\alpha\).

Note. Methods 1 and 2 yield a yes/no at \(\alpha\); method 3 yields a yes/no plus a graded measure of evidence (exact \(p\)).

4.4.3 Route 1: test statistic and critical values

Critical values come from the sampling distribution of the sample mean under \(H_0\), built in Figure 4.10.

Figure 4.13 draws it for a two-tailed test at \(\alpha = 0.05\), marked off in standard errors from \(\mu_0\). The red lines are the critical values, 1.96 standard errors on either side of \(\mu_0\). They cut off the most extreme 5% of the curve, 2.5% in each tail, and those shaded tails are the rejection regions. Here the observed \(\bar{x}\) lands 2.4 standard errors above \(\mu_0\), inside the right-hand rejection region. Counting that distance in standard errors is exactly the test statistic from earlier, \(z = (\bar{x} - \mu_0)/SE = 2.4\).

Here’s the decision rule:

  • If the observed test statistic falls inside the central region, we fail to reject \(H_0\) because the result is consistent with chance variation.
  • If it falls in either rejection region, we reject \(H_0\) because the result is too extreme to attribute to chance alone at the 5% level.

This picture makes clear why \(\alpha\) and the \(p\)-value are linked. \(\alpha\) sets the cutoff for how extreme is “too extreme,” while the \(p\)-value tells us how far into the tails our actual result lies.

A blue dot-dash bell curve centered on a magenta dashed line labeled mu sub 0. The x-axis is marked mu sub 0 minus 2 SE, minus 1 SE, mu sub 0, plus 1 SE, plus 2 SE. Red vertical lines at plus and minus 1.96 SE, each labeled critical value, with the small tail areas beyond them shaded red and labeled rejection region. A black dotted line labeled x-bar sits at 2.4 SE, inside the right-hand shaded region.
Figure 4.13: Two-tailed test at \(\alpha = 0.05\): the sampling distribution of the sample mean under \(H_0\) (blue dot-dash), centered at \(\mu_0\) (magenta) and marked off in standard errors. Red lines are the critical values at \(\pm 1.96\) SE; the shaded tails are the rejection regions. The observed \(\bar{x}\) (dotted) falls 2.4 SE above \(\mu_0\), in the rejection region.

4.4.4 Route 2: confidence interval

The confidence-interval route makes the same decision from the sample’s side. It never uses the null population distribution, only the single value \(\mu_0\).

Start from the sampling distribution of the sample mean built in Figure 4.9: centered at \(\bar{x}\), with the standard error as its spread. A sample mean lands within 1.96 standard errors of the true mean \(\mu\) about 95% of the time, so the middle 95% of this curve, \(\bar{x} \pm 1.96\,SE\), is the 95% confidence interval: the values of \(\mu\) that are consistent with our sample (Figure 4.14). Then ask where \(\mu_0\) falls (Figure 4.15).

A narrow dark-green dot-dash bell curve centered on a black dotted vertical line labeled x-bar, with a thick dark-green bar along the baseline under its middle, labeled 95% CI.
Figure 4.14: The 95% confidence interval: the middle 95% of the sampling distribution of the sample mean, \(\bar{x} \pm 1.96\) standard errors.

Decision rule:

  • If \(\mu_0\) lies inside the CI bar, we fail to reject \(H_0\).
  • If \(\mu_0\) lies outside the CI bar, we reject \(H_0\) at \(\alpha=0.05\).
Two panels, each showing the same narrow dark-green dot-dash curve centered on a dotted line labeled x-bar, with a thick dark-green confidence interval bar underneath. Left: a magenta dashed line labeled mu sub 0 sits just left of x-bar, inside the bar, with a box reading Fail to reject. Right: the magenta line sits further left, outside the bar, with a box reading Reject.
Figure 4.15: The decision, same sample in both panels. Left: \(\mu_0\) (magenta) falls inside the 95% confidence interval, so we fail to reject \(H_0\). Right: a different \(\mu_0\) falls outside it, so we reject \(H_0\).

The right panel is worth a second look. A single reading at that \(\mu_0\) would be unremarkable: it sits well inside the wide estimated population of Figure 4.9. It is simply not a plausible value for the mean, because sample means vary much less than individual readings do.

For common models, a two-tailed level-\(\alpha\) test of \(H_0:\ \mu=\mu_0\) is equivalent to checking whether \(\mu_0\) lies inside the \(100(1-\alpha)\%\) confidence interval. The CI is often the most intuitive way to see this: it shows both the decision (reject or not) and the range of parameter values still consistent with the data.

4.4.5 Route 3: \(p\)-value

The \(p\)-value is an area under the curve, so it comes from pnorm(), which Chapter 3 introduced as the area to the left of a value. For the observed \(z = 1.85\) in Figure 4.16, pnorm(-1.85) gives the area below \(-1.85\), about 0.032, which by symmetry equals the area above \(+1.85\). A two-tailed test counts both tails, so \(p\) = 2 * pnorm(-1.85) \(\approx 0.064\). Reject if \(p<\alpha\); here \(0.064 > 0.05\), so we fail to reject.

A blue dot-dash standard normal curve centered at zero, marked with a dashed vertical line at the center for the null value and a solid black vertical line further right labeled 'observed z' at about 1.85. The areas beyond plus and minus 1.85 in both tails are shaded red. Text on the left reports a two-tailed p-value of 0.064, labeled fail to reject.
Figure 4.16: p-value lens: the two-tailed \(p\)-value is the shaded area beyond the observed statistic (vertical line), in both tails, under the null.

4.4.6 Reject or fail to reject

Test outcomes are binary: we either reject \(H_0\) or fail to reject \(H_0\). We never accept \(H_0\), because sampling variability means we can’t prove it’s true, only that the data are consistent with it. Similarly, we don’t formally accept \(H_a\), but we can say the results provide support for \(H_a\) when \(H_0\) is rejected.

Report the exact \(p\) and an effect size or CI.

Here are our options (for a two-tailed test).

  1. Reject the Null Hypothesis (Two-Tailed Test). The sample mean falls beyond the critical values, in the rejection region. The confidence interval also excludes the null mean. In this example, \(\bar{x}\) is 3.75 standard errors from \(\mu_0\) (\(z = -3.75\)), well past the critical values of \(\pm 1.96\), and the two-tailed \(p\)-value is about 0.0002, far below \(\alpha = 0.05\). All three routes say reject.
A blue dot-dash bell curve centered on a magenta dashed line labeled mu sub 0, with red critical-value lines and shaded tails. A dark-green dot-dash bell curve of the same width sits well to the left, centered on a black dotted line labeled x-bar, which lies inside the left shaded rejection region. A thick dark-green confidence interval bar under the green curve does not reach mu sub 0.
Figure 4.17: Two-tailed test, reject \(H_0\). Blue: the sampling distribution of the sample mean under \(H_0\), with its rejection regions shaded beyond \(\pm 1.96\) SE; \(\bar{x}\) lands in one. Green: the sampling distribution of the sample mean from the sample, with its 95% confidence interval; \(\mu_0\) falls outside it.
  1. Fail to Reject the Null Hypothesis (Two-Tailed Test). The sample mean falls in the non-rejection region, and the confidence interval includes the null mean. Here \(\bar{x}\) is 1 standard error from \(\mu_0\) (\(z = -1.00\)), inside the critical values, and the two-tailed \(p\)-value is 0.32, above \(\alpha = 0.05\). All three routes say the evidence is consistent with chance variation, so we fail to reject \(H_0\).
A blue dot-dash bell curve centered on a magenta dashed line labeled mu sub 0, with red critical-value lines and shaded tails. A dark-green dot-dash bell curve of the same width sits just to its left, heavily overlapping it, centered on a black dotted line labeled x-bar that lies between the critical values. The thick dark-green confidence interval bar under the green curve reaches past mu sub 0.
Figure 4.18: Two-tailed test, fail to reject \(H_0\). \(\bar{x}\) lands 1 standard error from \(\mu_0\), inside the critical values of \(\pm 1.96\) SE, and the confidence interval from the sample includes \(\mu_0\).
CautionWarning: Beyond ruling out chance

Every hypothesis test in this book asks one question: could chance alone have produced this result? It is not the only question available.

Other methods start from two or more candidate explanations, say a steady linear decline versus a sharp threshold response, and ask which one the data support better.

Ruling out chance tells you that something is going on, but not always what. It is worth knowing the difference, because “significant” only ever answers the first question.

4.5 One- and two-tailed tests

\(H_a\) is a claim about where the true population could sit. A two-tailed alternative says it might sit below \(\mu_0\) or above it. A one-tailed alternative says it can only sit on one side (Figure 4.19).

Three panels. Left, titled H sub a colon mu not equal to mu sub 0: a blue bell curve in the center with an orange dashed bell curve of the same shape shifted to each side. Middle, titled H sub a colon mu less than mu sub 0: the blue curve with one orange dashed curve to its left. Right, titled H sub a colon mu greater than mu sub 0: the blue curve with one orange dashed curve to its right.
Figure 4.19: What \(H_a\) claims. The null population distribution (blue) sits at \(\mu_0\). The orange dashed curves are possible true population distributions under \(H_a\): on either side for a two-tailed alternative (left), on one side only for a one-tailed alternative (middle and right).

The test handles those claims through where \(\alpha\) goes. A two-tailed test splits \(\alpha\) between both tails of the sampling distribution of the sample mean under \(H_0\). A one-tailed test puts all of \(\alpha\) in the one tail \(H_a\) points toward (Figure 4.20).

Two panels, each a blue dot-dash bell curve centered on a magenta dashed line. Left panel, titled two-tailed with an arrow pointing both ways: small red shaded areas in both tails beyond red vertical lines at minus 1.96 and plus 1.96. Right panel, titled one-tailed with an arrow pointing right: a single, larger red shaded area in the right tail beyond a red vertical line at 1.645, which sits closer to the center than 1.96.
Figure 4.20: Where \(\alpha\) goes. Both panels show the sampling distribution of the sample mean under \(H_0\) (blue dot-dash), in standard-error units. Left, two-tailed (\(H_a:\ \mu \neq \mu_0\)): \(\alpha/2\) in each tail, critical values at \(\pm 1.96\). Right, one-tailed (\(H_a:\ \mu > \mu_0\)): all of \(\alpha\) in the right tail, critical value at 1.645.

4.5.1 One- or two-tailed: which and why?

Use this rule of thumb:

  • Two-tailed (default): You care whether \(\mu\) differs from \(\mu_0\) in either direction.
  • One-tailed (only if justified): Differences in the opposite direction are scientifically irrelevant and would not change your decision or action.

Default to two-tailed unless your theory rules out the other direction. At \(\alpha = .05\), a one-tailed test’s critical value is 1.645 instead of 1.96, so it rejects on less evidence, but only in the direction you named in advance. If the effect goes the other way, a one-tailed test cannot detect it at all.

That is why the direction has to be chosen before you see the data. Picking the tail afterward, because it gives the smaller \(p\)-value, makes your real chance of a false positive higher than the \(\alpha\) you report.

4.6 Chapter Summary

Why it matters. This chapter is the bridge between describing data and drawing conclusions from it. Everything after this point is a variation on one move: work out what chance alone could produce, then ask whether your result is rare enough under chance, by a standard set in advance (\(\alpha\)), to rule chance out.

Core ideas

  • Chance produces differences on its own. Split one population into two arbitrary groups and their means will differ anyway, from sampling error alone. The question is never “is there a difference,” it is “is this difference bigger than chance can account for.”
  • The chance window is buildable. Repeat a study thousands of times under conditions where you know nothing real is going on, collect the differences, and you have a picture of exactly what chance is capable of for a design your size.
  • Rarity is a judgment, and it needs a convention. The 100-year flood shows the reasoning already exists outside statistics: a 1% annual chance is not zero, but it changes what you build. \(\alpha = .05\) is that same judgment, standardized so researchers aren’t each drawing their own line.
  • A hypothesis test is the practical shortcut. Once you can assume a known reference distribution, the simulated chance window becomes a curve, and the decision becomes a calculation. The logic is the same.
  • Two pairs of curves. From the sample: the estimated population distribution and the sampling distribution of the sample mean, which gives the confidence interval. Under \(H_0\): the null population distribution and the sampling distribution of the sample mean under \(H_0\), which gives the critical values and \(p\)-value. The true population distribution is never known.
  • Three routes, one decision. A test statistic against a critical value, a confidence interval, and a \(p\)-value are three ways of asking the same question. For a two-tailed test at the same \(\alpha\) they always agree; a one-tailed test checked against a two-sided confidence interval can disagree near the boundary.
  • Commit to one- or two-tailed in advance. A one-tailed test is only justified when theory rules out the other direction. Choosing after seeing the data raises your real error rate above the \(\alpha\) you claimed.