Central Limit Theorem, t-test, and z-test
1. Overview
A. Definition
The foundational theory of statistical hypothesis testing, which infers the characteristics of a population from statistics obtained from a sample. The Central Limit Theorem (CLT) guarantees the normal approximation of the sample mean, and on top of it the z-test and t-test probabilistically judge hypotheses about the population mean.
The three concepts are linked in a single logical chain. We cannot know the entire population and see only a sample → the sample mean is a random variable that fluctuates each time it is drawn → to make inferences we must know the distribution of that fluctuation → the CLT guarantees that this distribution is normal → using that normality, the z-test and t-test judge "whether the observed difference is by chance or real." In short, the CLT is the "theoretical license" for inference, and the z-test and t-test are the "judgment tools" that operate under that license.
This triangular structure is not abstract mathematics but a practical engine of data-driven decision-making. Whether a new drug is more effective than a placebo, whether screen design A has a higher conversion rate than B, whether a process has deviated from spec, and whether the performance gap between two machine-learning models is meaningful are all judged with this framework. From a professional engineer's standpoint, the core is statistical thinking that goes beyond "memorizing test-statistic formulas" to checking whether the assumptions (normality, independence, equal variance) hold, interpreting significance together with effect size, and designing the sample size.
To position the three concepts briefly: the CLT is the theoretical foundation explaining "why the sample mean follows a normal distribution"; the z-test is the judgment tool for cases "where the population variance is known (or a large-sample proportion)"; and the t-test is the judgment tool for "a small sample with unknown population variance." The three are not separate topics but form one inference system, so they should be understood and described together.
B. Background and Necessity
In reality we cannot survey an entire population, so we observe only a portion as a sample. A full census is impossible in most cases due to cost, time, and physical constraints (such as destructive testing). The problem is that the sample mean x̄ is a random variable that differs each time it is drawn, and if we cannot quantify this fluctuation (sampling error), we cannot judge "whether the difference observed in the sample is a real population difference or merely chance."
The Central Limit Theorem opens the door to probabilistically computable inference by guaranteeing precisely that this fluctuation of the sample mean follows a known form, the normal distribution. "Knowing the distribution of the fluctuation" means "being able to compute the probability that an observation occurs by chance," and this probability (p-value) becomes the basis for hypothesis judgment. The z-test and t-test are the concrete tools that translate that normal approximation into actual hypothesis judgments.
This triangular structure underlies nearly all data-driven decision-making, including A/B testing, statistical process control (SPC), clinical trials, and model-performance comparison. Even in the era of big data and AI, the importance of hypothesis testing, which asks "is the observed difference statistically significant," is only growing, and misapplied tests lead to false discoveries and irreproducible conclusions.
C. Basic Procedure of Hypothesis Testing
Hypothesis testing follows these steps. (1) Set up the null hypothesis (H₀: no difference) and the alternative hypothesis (H₁: there is a difference). (2) Fix the significance level α (typically 0.05) in advance. (3) Compute the test statistic (z or t) from the data. (4) Obtain the p-value of that value and compare it with α. (5) If p < α, reject H₀ and judge "a significant difference"; otherwise judge "failure to reject." In this procedure, fixing (1) and (2) before looking at the data is the heart of reliability. Changing the hypothesis or significance level after seeing the results collapses the statistical meaning of the test.
2. Central Limit Theorem (CLT)
A. Overall Conceptual Structure
flowchart TD
POP["Population (any distribution form)"] -->|repeatedly draw samples of size n| SAMP["Compute sample mean x-bar"]
SAMP -->|n sufficiently large| CLT["Sample-mean distribution approaches normal"]
CLT -->|population variance known| Z["z-test"]
CLT -->|population variance unknown| T["t-test"]
Z --> DEC["Judge population-mean hypothesis"]
T --> DEC
The structure diagram above shows the relationship of the three concepts at a glance. The starting point is a population of any form; drawing samples of size n from it repeatedly and computing their means makes the distribution of those means converge to a normal distribution by the CLT. On top of this normality, the path splits into the z-test and the t-test depending on whether the population variance is known, and finally the hypothesis about the population mean (equal/different) is judged.
An everyday example is an opinion poll. The support rate of the entire populace (population proportion) is unknown, but if 1,000 people are randomly drawn to compute the support rate, by the CLT the distribution of that sample proportion approaches normal. That is why an interval can be presented, as in "48% support, margin of error ±3%p (95% confidence level)." Here the margin of error is exactly the standard error multiplied by 1.96, and the √n law — that the sample must be quadrupled to halve the margin of error — applies directly.
It is precisely because of this property that the CLT is called "the most important theorem in statistics." Even without knowing the population distribution, as long as the sample is large enough, inference can be unified with the single mathematical tool of the normal distribution. Not only the z-test and t-test but also confidence intervals, coefficient tests in regression analysis, and control charts, and many other techniques, are built upon this one theorem.
B. Definition and Principle
The theorem that, regardless of the population's distribution form, if the sample size n is large enough, the distribution of the sample mean x̄ approaches a normal distribution.
The key is that it holds "even if the population is not normal." For instance, a die's pips are a uniform distribution evenly spread over 1–6, but if you repeatedly compute the mean of 30 die rolls, the distribution of those means converges to a bell-shaped normal distribution. Even from an asymmetric population skewed to one side, such as an exponential distribution or an income distribution, collecting enough sample means makes their distribution gather into a normal form. This is because the sum and mean of many independent factors converge to a normal form as the skews of the individual distributions offset one another.
In this case the expected value of the sample mean stays equal to the population mean μ (unbiasedness), while the standard error, which indicates the dispersion, shrinks to σ/√n. The fact that √n is in the denominator is very important here. Because the sample must be quadrupled to halve the error, doubling the precision costs four times as much. This diminishing-returns structure is the core trade-off of sample-size design and the reason that a "just increase the sample indefinitely" approach is inefficient.
It is clearer in numbers. When the population standard deviation σ=10 and n=100, the standard error is 10/10=1.0, but to halve the error to 0.5 the n must be quadrupled to 400, and to reduce it to 0.25 requires 1600. In other words, no matter how large the sample grows, the speed at which the error converges to 0 keeps slowing. For this reason, in practice it is reasonable to take the approach of "first deciding the required precision and then back-calculating the minimum sample to match it" (power analysis).
The empirical criterion for "large enough" is often n≥30, but this is not an absolute rule and depends on the degree of population skew (skewness). If the population is already close to normal, the approximation is good even with a small n, and if it is severely asymmetric or has many extreme values, n must be in the hundreds to be stable. Therefore, in practice it is advisable to first visualize the data's distribution form (histogram, Q-Q plot) to check the validity of the approximation.
C. Why It Matters — The Foundation of Inference
Without the CLT we could not state in probability how far the sample mean can deviate from the population mean, and neither confidence intervals nor hypothesis tests would hold. For example, it is thanks to the CLT that the intuition "the sample mean is 52, so the population mean is probably around 52" can be turned into a quantitative interval such as "with 95% confidence the population mean is between 50.0 and 54.0." Making an interval by multiplying the standard error σ/√n by 1.96 of the normal distribution is the principle of the confidence interval, which is two sides of the same coin as hypothesis testing (if the interval does not contain μ₀, it coincides with the conclusion that it is significant). In this way the CLT is the foundation of all inferential statistics, binding estimation and testing into one.
| Item | Content | Meaning |
|---|---|---|
| Expected value of sample mean | Same as population mean μ | Unbiased estimation (unbiasedness) |
| Standard error of sample mean | σ/√n | Error decreases as n grows (proportional to √n) |
| Distribution form | Approaches normal | Normal-based inference even with unknown population distribution |
| Empirical condition | Typically n ≥ 30 | Larger n needed if skewness is large |
3. z-test and t-test
A. Selection Criteria
flowchart LR
Q{"population variance sigma known?"} -->|yes| Z["z-test (standard normal distribution)"]
Q -->|no| T2{"is the sample large?"}
T2 -->|small| T["t-test (t distribution, degrees of freedom n-1)"]
T2 -->|large| Z
The fork between the two tests is whether the population variance σ² is known. In reality it is rare to not know the population mean μ while knowing the population variance σ², so most practice involves estimating σ with the sample standard deviation s. A problem arises here. Because s itself is an estimate that fluctuates from sample to sample, the uncertainty of variance estimation is added on top of the uncertainty of the mean, doubly.
The t distribution is a distribution made with heavier tails than the normal distribution precisely to reflect this additional uncertainty. Heavier tails mean assigning a larger probability to extreme values, and as a result the critical value needed for rejection at the same significance level becomes larger, yielding a more conservative (cautious) judgment. The smaller the sample (the fewer the degrees of freedom), the heavier the tails, and as n grows, s approaches σ and the t distribution gradually converges to the normal distribution. So for large samples (n≥30) the results of the t-test and the z-test become practically the same, and in practice, when the population variance is unknown, the t-test is often used as the default regardless of sample size.
When, then, is the z-test used? Cases where the population variance is actually known are rare, but the proportion test is a representative use of the z-test. Proportion data follow a binomial distribution of success/failure with variance determined by the mean as p(1-p), so when the sample is large, the normal approximation (z) is used to test differences in proportions such as conversion rates and defect rates. This is why large-scale A/B test conversion-rate comparisons are commonly handled with the z-test (or chi-squared test). In other words, "mean comparison is mostly t, large-scale proportion comparison is mostly z" is the rough practical layout.
B. Characteristics Comparison
| Category | z-test | t-test |
|---|---|---|
| Distribution used | Standard normal (z) | t distribution (degrees of freedom n-1) |
| Population variance | Known (uses σ) | Unknown (uses sample standard deviation s) |
| Sample size | Assumes large sample (n≥30) | Applies even to small samples (n<30) |
| Distributional feature | Fixed bell shape | Heavy tails → converges to normal as n↑ |
| Representative use | Large-scale quality/proportion tests | Small-sample mean comparison, A/B testing |
C. Test Statistic and Calculation Example
The test statistic can be understood intuitively as "the observed difference divided by the standard error of that difference," i.e., a "signal-to-noise ratio." z = (x̄-μ₀)/(σ/√n), t = (x̄-μ₀)/(s/√n), where the numerator is the difference between the observed mean and the hypothesized mean (signal), and the denominator is the sampling error (noise). The larger this value, the more it means "a large difference hard to regard as chance."
For example, suppose a process has a target mean of μ₀=50, and a sample of 64 has a mean of x̄=52 and a sample standard deviation of s=8. The standard error is 8/√64 = 8/8 = 1, so t = (52-50)/1 = 2.0. If the two-tailed p-value of this value in a t distribution with 63 degrees of freedom is less than the significance level (e.g., 0.05), the null hypothesis (mean=50) is rejected and the judgment is "the mean differs from 50." Conversely, if the p-value is large, the difference is regarded as "within the range of chance" and cannot be rejected.
A common misunderstanding needs to be addressed here. Failing to reject the null hypothesis is not "proving there is no difference" but "lacking sufficient evidence to claim there is a difference." Also, the p-value is not "the probability that the null hypothesis is true" but "the probability of obtaining a value as extreme as the one observed, assuming the null hypothesis is true" — understanding this precisely is necessary for correct interpretation.
D. One-tailed/Two-tailed Tests and the Two Kinds of Error
Tests are divided into two-tailed and one-tailed according to the direction of the alternative hypothesis. When the direction is not specified, as in "is the mean different from 50," a two-tailed test looks at both tails of the distribution. When the direction is fixed, as in "is the mean greater than 50," a one-tailed test looks at only one tail, and at the same significance level rejection becomes easier. However, deciding the direction after seeing the data constitutes p-hacking, so a one-tailed test must always be chosen in advance with justification.
Hypothesis testing has two unavoidable kinds of error. A Type I error (α) is wrongly rejecting "there is a difference" when in fact there is none (false positive), and a Type II error (β) is missing a real difference by claiming "there is none" (false negative). There is a trade-off in which lowering α (being stricter) increases β, so the significance level must be set by comparing the costs of the two errors. For example, for a case like disease diagnosis where missing is fatal, design toward reducing β, and where a false alarm is costly, design toward reducing α. Power (1-β) indicates how well this Type II error is avoided.
4. Types of t-test
The t-test is divided into three types according to the comparison structure, and choosing the type that fits the situation governs the validity of the result. "What is being compared with what" and "are the two groups independent or paired" must first be sorted out before the correct type is determined. The key is that type selection is not a matter of choosing a formula but a matter of understanding the structure of the experiment and data.
The one-sample t-test examines whether the mean of a single group equals a specific reference value. For example, it tests whether the measured average lifespan of a new battery differs significantly from the published spec of 10 hours. It is used in quality control to judge "whether a process maintains its target value."
The independent two-sample t-test compares the means of two unrelated groups and is the standard tool for A/B testing. For example, it exposes web-page design A and B to different visitor groups and compares the average conversion rate (or dwell time). When it is hard to assume that the two groups have equal variance, it is safer to use Welch's t-test, which does not assume equal variance.
The paired t-test tests the mean of the difference by measuring the same subjects twice, before and after treatment. As with the blood pressure of the same patient before and after medication, or the task time of the same user before and after a UI improvement, it removes the noise of individual differences and has high power. Because it can detect an effect with fewer samples than the independent-sample test, it is preferred in experimental designs where before-after comparison is possible.
Why the paired-sample test is advantageous in power can be understood through its variance structure. When people's baseline blood pressure differs greatly, in the independent-sample test those large individual differences inflate the denominator (standard error) and mask the difference. The paired-sample test, by contrast, looks only at "the before-after difference of the same person," so individual differences cancel out and the denominator shrinks, making the same effect easier to detect significantly. However, if some other change (learning effect, passage of time) intervenes between the before and after measurements, it is confounded with the treatment effect, so this must be controlled with a control-group design.
| Type | Use | Example |
|---|---|---|
| One-sample t | One group's mean vs reference value | Measured lifespan vs spec |
| Independent-sample t | Compare means of two unrelated groups | A/B test conversion rate |
| Paired-sample t | Before-after comparison of same subjects | Blood-pressure change before/after treatment |
Confusing the three types distorts the result. For example, mistakenly handling the same person's before-after data as independent samples leaves individual differences as noise and lowers power, while conversely binding data from different people as paired samples commits the error of assuming nonexistent pairs. Also, when comparing three or more groups, repeating the t-test many times creates a multiple-comparisons problem, so in this case analysis of variance (ANOVA) is used to compare all at once and post-hoc tests determine which pair differs. In other words, the principle is to apply the t-test only within the scope of "comparing the means of two groups (or one group)."
5. Deep Dive — Practical Pitfalls and Trends in Statistical Reform
The multiple-comparisons problem is the most common pitfall in practice. If a test is repeated 20 times at a 5% significance level, even if there is actually no difference, on average one will come out "significant" by chance. More generally, when m independent tests are performed, the probability that at least one is a false positive is 1-(1-0.05)^m, which reaches about 40% when m=10. This is why false positives surge when multiple metrics are viewed simultaneously in an A/B test or interim results are peeked at repeatedly. To prevent this, the Bonferroni correction (dividing the significance level by the number of tests), FDR control, a pre-specified analysis plan, and sequential-test design are needed.
Statistical process control (SPC) on the manufacturing floor is a representative practical application of the CLT. A control chart repeatedly measures subgroup means and, on the CLT premise that their distribution is normal, draws ±3σ control limits, alerting a process anomaly when the mean goes outside this range. For example, in a plating process with a target thickness of 50μm and a process standard deviation of 2μm, if the subgroup size is n=25, the standard error of the sample mean is 2/√25=0.4μm, so the control limits are set narrowly at roughly 50±1.2μm. Because the mean, not individual measurements, is managed, more sensitive and stable anomaly detection is possible thanks to the CLT.
p-hacking and the reproducibility crisis are also central topics in the recent data-science community. If one keeps trying by changing variables, subsets, and analysis methods until a significant result appears, the p-value loses its credibility. In response, the American Statistical Association (ASA) recommended in its 2016 statement, "Do not rely on the dichotomous threshold of p<0.05; interpret comprehensively the effect size, confidence interval, and study context." That is, statistical practice is moving in the direction of reporting, beyond "significant or not," also "how large the difference is and how uncertain it is."
Effect size and power design are the heart of that alternative. When n is very large, even a difference so trivial as to be practically meaningless becomes "statistically significant." For example, in an A/B test with millions of users, even a 0.01%p conversion-rate difference yields p<0.05, but whether that difference is meaningful for the business is a separate matter. Therefore an effect size such as Cohen's d must be reported together with the confidence interval to judge whether it is a "meaningful difference." Conversely, one must design in advance with a power analysis to determine "how many samples are needed to detect the desired effect at significance level α and power 1-β," to prevent missing a real effect (Type II error) due to insufficient sample. This is a practical application directly tied to the CLT's σ/√n structure.
Application in AI and Data Fields and Past-Exam Links
In machine learning this framework is used as-is. One measures the accuracy of two models A and B multiple times (per cross-validation fold) and uses a paired t-test to judge "whether the performance difference is significant," or an online experiment platform computes confidence intervals in real time on a CLT basis to decide when to end the experiment. However, if the samples are not independent (time series, repeated measures) or normality breaks down, corrected tests, bootstrap, and Bayesian methods should be considered alongside the standard t-test. This topic is examined in the Professional Engineer of Information Management exam in connection with statistical quality control, data analysis, experimental design, and confidence intervals/regression analysis ([[correlation-causation]]), in short-answer and essay forms, so it is effective to organize it not as formula memorization but as the thought flow "assumptions → test selection → interpretation → decision."
6. Considerations and Implications
- Assumption verification comes first: The t-test and z-test assume the normality, independence, and (when comparing two groups) equal variance of the observations. When violated, they must be replaced with Welch's t (unequal variance) or nonparametric tests such as the Mann-Whitney U or Wilcoxon (non-normal) to prevent result distortion. Applying the formula without checking assumptions is the most common error.
- Report significance and effect size together: The p-value says only "whether the difference is by chance," not "whether the difference is practically large." Present Cohen's d and the confidence interval together, and clearly distinguish "statistical significance ≠ practical importance" in large samples.
- Pre-design sample size and power: By the CLT's σ/√n structure, precision is proportional to √n, so pre-compute the appropriate sample for the cost with a power analysis. An undersized sample misses the effect (Type II error), and an oversized one causes wasted cost and meaningless significance.
- Control multiple comparisons and repeated tests: Control false positives from multiple metrics and repeated looks with corrections such as Bonferroni and a pre-specified analysis plan, and set early-stopping rules for A/B tests in advance.
- Secure the sample's representativeness and randomness: The CLT and tests stand on the premise that the sample represents the population. Biased sampling (self-selection, survivorship bias) is not corrected no matter how large n grows, so random sampling and sample design are a more fundamental task that precedes the computation of statistics.
- Correct interpretation and communication of the p-value: The p-value is merely "the probability of obtaining an extreme value under the null hypothesis," not "the probability that the hypothesis is true," nor "the size of the effect." When reporting to decision-makers, convey the effect size, confidence interval, and uncertainty together instead of the dichotomy of "significant or not" to prevent over- and under-interpretation.
- Practical application and automation: It is used directly for A/B test conversion-rate comparison, detecting quality deviations in manufacturing processes (SPC), and verifying significant performance differences between two ML models; embed test logic into data pipelines and experiment platforms, but assumption checking and interpretation must be the responsibility of people. The more statistical tools are automated, the more important becomes the competence to understand "what is being tested and whether the assumptions hold."
References
- NIST/SEMATECH e-Handbook of Statistical Methods: https://www.itl.nist.gov/div898/handbook/
- American Statistical Association, Statement on p-Values (2016): https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf
- Wasserstein & Lazar, "The ASA Statement on p-Values: Context, Process, and Purpose," The American Statistician (2016): https://doi.org/10.1080/00031305.2016.1154108
In one line: The CLT guarantees that when n is large, regardless of the population distribution, the sample mean approaches a normal distribution (mean μ, standard error σ/√n), and the population-mean hypothesis is judged with the z-test if the population variance is known and the t-test (t distribution, degrees of freedom n-1) if it is unknown (small sample); but reliable results require also considering checks of the normality, independence, and equal-variance assumptions and control of effect size, power, and multiple comparisons.