Econometrics is the toolkit economists and social scientists use to make data say something reliable. It serves four broad purposes: estimating economic relationships, testing economic theories, evaluating policy, and forecasting. What sets it apart from the methods of the natural sciences is the kind of data it usually has to work with. A chemist can run a controlled experiment that isolates one variable while holding everything else fixed. A social scientist almost never can. Instead, economists mostly rely on observational data, also called nonexperimental data, which is gathered by passively watching outcomes on people, firms, schools, and cities, without the researcher controlling who receives any treatment. Genuinely experimental data, where subjects are randomly assigned to treatment and control groups, is the rare exception in economics rather than the rule.
The distinction matters because it creates the central difficulty of the whole field. A drug trial that randomly assigns patients to a drug or a placebo produces experimental data, and so does a programme that randomly assigns job training to some unemployed workers. But a survey recording workers’ wages and years of schooling, or an analysis of crime rates and police spending across cities, produces observational data, where the researcher had no control over who got what. Because the observational case dominates economics, the recurring challenge is distinguishing correlation from causation.
The Four Steps of Empirical Analysis
Every empirical study moves through the same four steps. First you pose the question of interest, stating clearly the causal or descriptive relationship you want to study. Second you specify an economic or conceptual model, using theory or intuition to decide which variables are relevant and how they relate. Third you turn that economic model into an econometric one, which means resolving three practical issues: how to measure each variable, what functional form the relationship takes, and how to account for the unobserved factors that make any real relationship inexact. Fourth you collect data and apply statistical methods to estimate the parameters, build confidence intervals, and test hypotheses.
Consider worker productivity and job training. The question is whether job training raises productivity, measured by the hourly wage. Economic reasoning suggests the wage depends on education, experience, and training.
To make this estimable, you measure education in years of schooling, experience in years in the workforce, and training in weeks of job training, then assume a linear form and add an error term to capture everything left out.
Each piece has a role. The intercept beta-zero is the baseline wage when every explanatory variable is zero. The slope parameters describe the direction and strength of each variable’s effect, with beta-three, the effect of training, being the parameter of primary interest. The error term u absorbs all the other determinants of wage that are not modelled explicitly, such as intelligence, motivation, and measurement error. With data on a large sample of workers, regression estimates the betas and tests hypotheses, for instance that training has no effect on wage.
The same four steps work for very different questions. Applied to crime, the question becomes what determines the time an individual spends in criminal activity. Becker’s utility-maximising framework treats crime as an economic choice in which a person weighs the earnings from crime against the costs of being caught, so hours in criminal activity depend on legal wages, other income, the probabilities of arrest and conviction, age, and sentence length. Translated into a linear model, this becomes an equation in observed variables plus an error term that soaks up the unmeasurable factors such as moral character, family background, and membership of a criminal network.
The structure is identical to the wage example; only the substantive theory guiding which variables to include has changed.
The Three Structures of Economic Data
Economic data comes in three main shapes, and the right method depends on which you have.
Cross-sectional data records multiple units, people, families, firms, or cities, at a single point in time, with each row a unit and each column a variable. Its defining convenience is that it can often be treated as a random sample, meaning each unit had an equal chance of selection and the draws are statistically independent, which greatly simplifies the analysis. Qualitative features such as gender or marital status are recorded with binary indicator or dummy variables taking the value zero or one. A useful property is that the order of observations is completely arbitrary: a dataset of 526 workers would be identical in content if the rows were shuffled, because nothing about the analysis depends on their sequence.
Time series data is the opposite in this respect. It records one or more variables repeatedly over time, monthly, quarterly, or annually, as with interest rates, unemployment, GDP, or stock prices, and here the order is essential and must be preserved. Independence cannot be assumed, because knowing GDP in 2009 tells you a great deal about GDP in 2010, so the observations are serially correlated, or autocorrelated, and the methods must account for that. Many series also carry a long-run trend, such as the steady growth of real GDP over decades, and sub-annual series often show seasonality, the regular within-year pattern of, say, housing starts peaking in summer. Both trend and seasonality have to be handled explicitly.
Panel data, also called longitudinal data, combines the two by following the same units across multiple periods. A dataset of 2,000 Michigan schools observed annually for ten years yields twenty thousand observations and two real advantages. It can control for time-invariant unobservables, so a school’s unchanging quality of management can be accounted for even though it is never directly measured, and it can model lagged responses, showing how an event in one period affects outcomes later. Panel methods are more advanced, and introductory work concentrates on the cross-sectional case.
Random Sampling and Its Failures
Random sampling is the formal condition that lets us treat cross-sectional observations as independent and identically distributed, meaning they are independent draws from the same probability distribution. It matters because it makes the sample representative of the population, which is what gives us a genuine chance of learning the population’s true characteristics from a finite sample. A researcher who randomly draws 500 workers and records their wages and schooling has an i.i.d. sample: each worker was equally likely to be chosen, and one worker’s schooling reveals nothing about another’s, which is exactly what standard regression needs to estimate the return to education validly.
Two violations break this. Sample selection occurs when part of the population is systematically excluded for a reason tied to the variable under study, as when wealthy families disproportionately refuse to report wealth, leaving an unrepresentative dataset. A wage dataset built only from volunteers fails the same way if, say, high earners are too busy to participate, because the sample then under-represents them and the estimates are biased. The second violation is non-independence from geographic clustering: when units are large relative to the whole, like US states, neighbours are not economically independent, because business activity in adjacent states is linked through cross-border employment and trade.
Causality and Ceteris Paribus
Causality is the real concern of most economic policy analysis, and correlation is only a hint of it. The causal effect of x on y is defined by a precise question: how does y change when x is changed while all other relevant factors are held constant? That last condition is the Latin ceteris paribus, all other things being equal, and nearly every policy question in economics is a ceteris paribus question.
The need to hold other factors fixed is what makes naive correlations dangerous. Cities with more police officers often have higher crime, but it would be absurd to conclude that police cause crime; larger cities simply have more of both and hire police in response to crime. The causal question is whether, holding city size and everything else fixed, one more officer reduces crime. The difficulty is that in observational data you cannot literally hold all else equal, so the empirical task is always whether you have controlled for enough other factors to make a causal reading defensible. Consumer demand shows the same trap: the causal question is how much quantity demanded falls when price rises by a pound, holding income, the prices of substitutes and complements, and tastes constant. If income happens to fall at the same time as the price rises, as in a recession, a simple price-demand correlation tangles the price effect together with the income effect, and only controlling for income isolates the price effect.
The Logic of Experiments: The Fertiliser Example
The cleanest way to see causal inference is in a setting where the ideas are simple. How much does soybean yield rise when fertiliser increases by one unit? This is a ceteris paribus question, because yield also depends on rainfall, land quality, and parasites, and you cannot literally hold those fixed since land quality varies from plot to plot and cannot even be fully observed. The experimental solution sidesteps the problem: assign fertiliser amounts randomly and independently of every land characteristic. If the assignment is truly random, the plots getting more fertiliser are, on average, similar in all other respects to those getting less, so the remaining differences in yield can be attributed to fertiliser alone, and a simple regression gives a valid causal estimate. The key principle is that genuine random assignment of the treatment, independent of all other factors, is enough for causal inference, even though it is impossible to hold every factor fixed by hand. The contrast makes it vivid: if instead farmers chose their own fertiliser levels and the more fertile land got more, then plots with better land would show higher yields for two reasons at once, and the correlation would overstate fertiliser’s true effect.
Counterfactuals and Self-Selection
Counterfactual reasoning offers another way to understand the same idea. Rather than holding factors fixed across different people, it imagines the same individual in two states of the world, one where they received the treatment and one where they did not. The causal effect for that individual is the difference between the two outcomes. Take private versus public schooling and its effect on adult wages: for individual i, the effect is the wage under private schooling minus the wage under public schooling.
The fundamental problem is that each person is only ever observed in one state, and the other, the counterfactual outcome, is never seen. This makes the individual causal effect impossible to measure directly, but the framework is valuable because it clarifies exactly what we wish we could observe and why a simple comparison falls short. Comparing adults who attended private schools with those who attended public ones does not hold other factors fixed, because the two groups also differ systematically in family income, parental education, and ability, and those differences drive part of the wage gap even if schooling itself had no causal effect at all.
That last point is the essence of self-selection, one of the most common threats to causal inference in observational data. Self-selection happens when individuals or firms choose their own level of the treatment based on characteristics that also affect the outcome. If higher-ability people choose more schooling, and ability raises earnings independently, then the correlation between wages and education reflects both the true return to schooling and the hidden effect of ability, so none of it can be cleanly attributed to education. A truly randomised experiment, assigning years of schooling at birth, would solve it, but that is neither feasible nor ethical, which is why observational work has to lean on statistical methods that control for confounders like ability as far as they can be measured. The same logic appears in a classroom: if stronger students attend lectures more often, the positive correlation between attendance and grades partly reflects ability rather than attendance, because students self-select into how much they attend.
Bringing It Together: Firm-Level Job Training
A firm-level question ties all of this together. Suppose Ohio manufacturing data records hours of job training per worker and output per worker hour. The ceteris paribus thought experiment frames it correctly: if two firms were identical in every respect except that one provided an extra hour of training per worker, how much higher would its output per worker be? The trouble is that training is not assigned independently of worker characteristics. Firms decide how much training to offer based partly on their workers’ schooling, experience, and tenure, and on unmeasurable qualities like ability, motivation, and management’s judgment of potential, and workers with particular traits may self-select into firms that train more. Other factors muddy it further: the capital and technology available to workers, and the quality of management, both of which are hard to measure and both of which affect output. So a positive correlation between training and output establishes causality only if training was randomly assigned. In the observational setting it does not, because that correlation is confounded by worker characteristics, capital, and management. Even if job training had no causal effect whatsoever, you could still see a positive correlation, simply because firms that invest in training also tend to have more skilled workers, better technology, and stronger management.
The Short Version
Econometrics exists to extract reliable answers from data that is usually observational rather than experimental, and that fact shapes everything. Empirical work moves from a question to an economic model to an econometric model with an error term, and finally to estimation. The data arrives as cross-sections, where random sampling lets us treat observations as independent, as time series, where order and serial correlation matter, or as panels, which combine the two. The deepest challenge throughout is causality, defined by the ceteris paribus condition of holding all else equal, which observational data can never do literally. Random assignment solves the problem when it is available, the counterfactual framework clarifies what we are really after, and self-selection and confounding explain why naive correlations so often mislead. The discipline of econometrics is, in the end, the discipline of asking whether you have controlled for enough to believe a number means what you want it to mean.
[…] The Nature of Econometrics and Economic Data […]