Version 2 of 2
Introduction
Generated Aksbel book section. · Working · Aug 11, 2026 22:14 · saved by @mujirin
Introduction
Statistics begins with a simple difficulty: we want to know more than we can directly see.
A doctor wants to know whether a treatment helps patients, but cannot observe what would have happened to the same patients under every possible treatment. A city wants to know whether a new traffic policy reduces crashes, but the city cannot rerun the same year with and without the policy. A researcher wants to understand whether income, education, health, or environment are related, but the data come from a world where many things change at once. A polling organization wants to estimate public opinion, but it cannot ask every person.
In each case, there is evidence: observations, measurements, records, survey answers, experimental results, or counts. There is also a desired inference: a conclusion that reaches beyond the exact observations in hand. Statistics is the discipline that studies how to make such inferences responsibly under uncertainty.
This book is about that movement from evidence to inference.
It is not mainly a book about memorizing formulas. Formulas matter, and we will learn them carefully. But a formula is useful only when we understand what question it answers, what assumptions it requires, and what it leaves uncertain. Averages, standard deviations, regression lines, confidence intervals, and p-values are tools. Like all tools, they can clarify or mislead depending on how they are used.
A central theme of this book is that statistical reasoning begins before calculation. The way data are produced often matters as much as the arithmetic performed afterward. This emphasis is shared by influential introductory statistics traditions, including the approach of Freedman, Pisani, and Purves, who repeatedly stress the connection between study design, data quality, chance variation, and the limits of statistical conclusions (Freedman, Pisani, and Purves, 2007).
What makes a question statistical?
A question becomes statistical when it involves variation and uncertainty.
Variation means that observations are not all the same. People have different heights, incomes, blood pressures, test scores, political opinions, and medical outcomes. Manufactured parts differ slightly even when they are made by the same process. Daily temperatures change. Prices fluctuate. Measurements contain error.
Uncertainty means that the evidence does not determine a conclusion with complete certainty. If a survey asks 1,200 adults whom they support in an election, the results may give useful evidence about the larger electorate, but they will not be exactly the same as the result of asking every eligible voter. If patients receiving a new drug recover more often than patients receiving a placebo, the difference may reflect a real treatment effect, but it may also partly reflect chance variation, differences among patients, or problems in how the study was conducted.
For example, suppose 54% of people in a survey say they support a policy. It is tempting to report, “A majority supports the policy.” But a statistical thinker asks more questions:
- How were the people selected?
- Who was left out?
- How was the question worded?
- How many people refused to answer?
- Could random sampling error account for the result?
- Does “support” mean strong support, mild support, or temporary agreement?
The number 54% is not meaningless. But it is not self-explanatory. Statistics teaches us how to connect such numbers to the process that produced them.
Data are not the same as truth
The word data refers to recorded information: measurements, labels, answers, counts, times, locations, images, or other observations. Data are often treated as if they are raw truth, but they are better understood as traces left by a measurement process.
Consider “hours studied last week.” A student may estimate from memory. A learning platform may record time with a browser tab open. A researcher may ask students to keep a diary. These are three different ways to measure something related to studying, but they are not identical. A browser tab can be open while the student is distracted; memory can be inaccurate; a diary can change behavior because the student knows studying is being recorded.
This is why later chapters will pay close attention to measurement. A measurement is a rule-governed way of assigning values to features of the world. Good measurement does not require perfection, but it does require clarity. We must know what was measured, how it was measured, and what kinds of error may have entered.
A simple example makes the point. Suppose two schools report average class size. School A reports 24 students per class; School B reports 28. Before concluding that School A has smaller classes, we should ask how “class size” was defined. Did the schools count special tutorials? Laboratory sections? Online sections? Did they average over courses, teachers, or students? The same phrase can hide different operational definitions.
An operational definition is the exact procedure used to define and measure a concept. The concept may be broad, such as “health,” “poverty,” or “learning.” The operational definition is concrete, such as “self-reported health on a five-point survey scale,” “household income below a specified threshold,” or “score on a particular exam.” Statistical calculations operate on the operational definition, not directly on the broad concept.
Description, prediction, causation, and decision
Statistical work often has one of four purposes: description, prediction, causal inference, or decision-making. These purposes overlap, but they are not the same.
Description summarizes what the data show. If we compute the median rent in a city, draw a histogram of household income, or report the proportion of patients who recovered in a study, we are describing data. Good description is not trivial. It requires choosing summaries that reveal structure without hiding important features such as skewness, outliers, or subgroup differences.
Prediction uses information we have to guess or estimate information we do not yet have. A weather model predicts tomorrow’s temperature. A bank estimates whether a loan applicant will default. A university may try to predict which students need additional support. Prediction can be useful even when we do not fully understand the underlying causes. However, a prediction that works in one setting may fail in another if conditions change.
Causal inference asks whether changing one thing would change another. Does a vaccine reduce infection risk? Does a tutoring program improve exam performance? Does air pollution increase hospital admissions? Causal questions are harder than descriptive or predictive questions because they involve comparison with alternatives that usually cannot all be observed for the same unit at the same time. Modern discussions of causal inference often emphasize the importance of comparing observed outcomes with relevant counterfactual outcomes—outcomes that would have occurred under different conditions (Holland, 1986).
Decision-making uses evidence, values, costs, and uncertainty to choose an action. A public health agency deciding whether to recommend screening must consider not only test accuracy, but also false positives, false negatives, cost, access, and possible harm. Statistics can inform decisions, but it does not by itself determine what society should value.
Confusing these purposes is a common source of error. A study may describe an association without proving causation. A model may predict accurately without explaining why. A statistically significant result may be too small to matter in practice. A decision may be reasonable even when uncertainty remains.
Association is not automatically causation
One of the most important habits in statistics is to distinguish association from causation.
Two variables are associated when their values tend to vary together. For example, ice cream sales and drowning incidents may both be higher in summer. That association does not mean ice cream causes drowning. A third factor—hot weather and seasonal swimming—helps explain why both numbers rise. Such a third factor is often called a confounder when it is related to both the possible cause and the outcome in a way that can distort a causal interpretation.
This does not mean associations are useless. Associations can suggest hypotheses, improve prediction, and sometimes support causal conclusions when combined with strong study design and background knowledge. But association alone is not enough.
Randomized experiments are powerful because random assignment tends, in expectation, to balance both known and unknown background factors between treatment groups. This logic was developed as a foundation of experimental design in the work of R. A. Fisher (Fisher, 1935). If patients are randomly assigned to receive either a new treatment or a control treatment, then large systematic differences in outcomes are harder to explain by pre-existing differences between the groups. Randomization does not make every experiment perfect, but it gives a principled way to reason about chance and comparison.
Observational studies, by contrast, observe what happens without assigning treatments by random mechanism. They are often necessary: we cannot randomly assign people to smoke for decades or to live in polluted air. But observational evidence must be interpreted with special care because people who receive one exposure may differ from others in many ways.
This book will return to this distinction many times. It is one of the central boundaries between what data can show directly and what requires additional assumptions.
Chance is not chaos
In everyday speech, “random” sometimes means patternless or meaningless. In statistics, a random process is one whose individual outcome is uncertain but whose long-run behavior can often be described by a probability model.
A fair coin toss is the simplest example. We cannot predict with certainty whether the next toss will land heads or tails. But if the coin is fair and the tosses are performed under stable conditions, we expect the proportion of heads to be close to one-half over many tosses. The individual result is uncertain; the long-run pattern is regular.
This idea is central to statistical inference. A sample survey gives a result that varies from sample to sample. An experiment gives results that may differ if repeated. A manufacturing process produces parts that are not exactly identical. Statistics uses probability to describe this kind of variation.
A probability model is a mathematical description of chance behavior. It does not claim that the world is literally a set of dice or coins. Rather, it provides a disciplined approximation. The model is useful when its assumptions are close enough to the data-generating process for the question at hand.
For example, if a survey uses a properly conducted simple random sample, probability theory can help estimate how far the sample proportion is likely to be from the population proportion. If the survey is instead an online poll of volunteers, the same formulas may produce a neat number but not a trustworthy inference. The mathematics cannot repair a badly produced sample.
Inference requires assumptions
An inference is a conclusion that goes beyond the data directly observed. If 38 out of 50 sampled parts fail inspection, saying “38 of these 50 parts failed” is description. Saying “about three-quarters of all parts from this process would fail” is inference.
Every inference requires assumptions. Some assumptions are explicit, such as “the sample was randomly selected.” Others are hidden, such as “respondents understood the question in the same way” or “future conditions will resemble past conditions.” A major goal of statistical education is not to eliminate assumptions—that is impossible—but to make them visible and examine whether they are reasonable.
Consider a medical test. Suppose a test for a disease is positive. Many people want to know: “What is the probability that the person has the disease?” The answer depends not only on the test’s accuracy but also on how common the disease is in the tested population. This is a conditional probability problem: we must reason about probability under given information. Without the base rate of disease, the test result cannot be interpreted properly.
This kind of example shows why statistics is a way of thinking, not only a set of calculations. The numerical answer depends on the structure of the question.
Statistical tools have meanings, not just formulas
Throughout the book, we will learn many standard tools. Each will be introduced through its meaning before its formula.
The mean is the arithmetic average, but it is also a balance point of a distribution. The median is the middle value, but it is also a resistant summary that is less affected by extreme values. The standard deviation measures typical distance from the mean, but it is meaningful only after we understand the distribution being summarized. A correlation coefficient measures the strength and direction of a linear association, but it does not detect all forms of relationship and does not prove causation. A p-value measures how unusual the data would be under a specified null hypothesis, not the probability that the null hypothesis is true. The American Statistical Association has emphasized that p-values are often misinterpreted when treated as automatic measures of truth or importance (Wasserstein and Lazar, 2016).
Graphs will be treated as serious statistical tools, not decoration. A histogram can reveal skewness or multiple clusters that a mean hides. A scatterplot can show nonlinearity, outliers, or separate groups that a correlation coefficient compresses into one number. John Tukey’s work on exploratory data analysis helped establish the importance of using visual and resistant methods to discover structure in data before imposing formal models (Tukey, 1977).
The order matters: look, think, model, calculate, interpret, and then check.
A small example: the average alone is not enough
Suppose two neighborhoods report the same average household income: $60,000.
In Neighborhood A, most households earn between $50,000 and $70,000. In Neighborhood B, many households earn below $35,000, while a few earn several hundred thousand dollars. The mean is the same, but the social meaning is different. If we report only the average, we hide inequality and variation.
This example introduces a recurring lesson: a statistic is a summary, and every summary loses information. The question is whether it loses information that matters.
A good statistical analysis does not ask, “What formula can I apply?” It asks:
What is the question?
What data were collected?
How were they collected?
What does the summary reveal?
What does it hide?
What uncertainty remains?
What conclusion is justified?
These questions are simple, but they are not easy. They require practice.
The path of the book
The first chapters build the foundation. We begin with what statistics is for, then study data, measurement, and design. Before formal inference, we learn to describe distributions, compare summaries, and see relationships in graphs. This is deliberate. Statistical inference without descriptive understanding is fragile.
The middle chapters introduce probability. Probability is the language we use to model chance variation. We will study random variables, expected value, binomial models, sampling distributions, and the central limit theorem. These ideas explain why sample results vary and why, under suitable conditions, we can quantify that variation.
The later chapters develop inference: confidence intervals, tests of significance, power, categorical data methods, group comparisons, and regression inference. These tools are widely used in science, business, public policy, medicine, and everyday argument. We will study both how to use them and how they can be misused.
The final chapters return to real-world judgment: how to read statistical claims, how to evaluate evidence, how to communicate uncertainty, and how to use statistics responsibly.
The attitude of this book
Statistics rewards a certain intellectual temperament. Be curious, but not gullible. Be skeptical, but not cynical. Demand evidence, but remember that evidence is almost always incomplete. Use models, but do not worship them. Calculate carefully, but interpret even more carefully.
A statistical conclusion should usually sound like this:
“Given these data, produced in this way, and assuming this model is adequate for this purpose, the evidence suggests this conclusion, with this much uncertainty.”
That sentence is longer than “the data prove it.” It is also more honest.
By the end of this book, you should be able to read a graph, question a survey, interpret a confidence interval, understand a p-value, recognize a weak causal claim, and explain why study design matters. More importantly, you should be able to slow down when faced with a number and ask what kind of evidence it really is.
That is the beginning of statistical thinking.
References
Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
Freedman, D., Pisani, R., and Purves, R. (2007). Statistics (4th ed.). W. W. Norton & Company.
Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945–960.
Tukey, J. W. (1977). Exploratory Data Analysis. Addison-Wesley.
Wasserstein, R. L., and Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. https://doi.org/10.1080/00031305.2016.1154108.