Glossary
Every technical term defined in plain language.
Plain-language definitions of the data-science terms used in this project, listed A–Z.
A
Adding background variables to a model so the main relationship is less distorted by measured differences between cases. Adjustment strengthens a comparison, but it does not guarantee a causal result.
Data summarized at a group level — for example, a state's average score rather than individual student scores.
A relationship in which two variables tend to move together. Association can support description or prediction, but by itself it does not prove that one variable causes the other.
B
The measurement taken before an event or intervention, used as the reference for comparison. In this project, 2015–2019 is the baseline period.
A reference score or standard against which performance is judged (for example, NAEP's achievement levels).
The straight line drawn through a scatterplot to summarize the overall linear trend. In ordinary least squares regression, it is chosen to minimize the squared distances between the points and the line.
A systematic error that pushes an estimate off in one direction, rather than random error. Sampling bias, selection bias, and measurement bias all mean the estimate is consistently off-target.
A computational method that repeatedly resamples your data (random samples drawn with replacement, thousands of times) to estimate how much a statistic — like a mean or a regression coefficient — would vary across new samples.
C
The part of data science that asks whether changing one thing would change another. It is stronger than prediction because it tries to estimate the effect of an intervention rather than only describe or forecast a pattern.
Measuring every member of a population, as opposed to a sample.
In machine learning, a model that assigns each case to a category (for example, "at risk / not at risk").
A machine-learning method that groups similar items together without being told the categories in advance.
The number in a regression output that tells you how much the outcome changes when a predictor increases by one unit, holding all other predictors constant.
A group of individuals who share a defining characteristic or time period (for example, "the cohort of 8th graders tested in 2022").
An approach in which groups, times, or places are contrasted. Comparison is a way of doing analysis, not automatically a separate research aim. (This project is built on comparison: pre-COVID vs. post-COVID, national vs. Texas.)
A range computed from the data that would capture the true value in a fixed percentage of repeated samples (for example, 95%).
A variable that is correlated with both the predictor and the outcome, creating a false impression that the predictor caused the outcome. (Poverty is a classic confound when comparing scores across groups.)
A control group is the comparison group that does not receive the treatment, and a control variable is a variable held constant in a model.
A measure of how two variables move together. Correlation does not imply causation.
A number, often written as r, that summarizes the direction and strength of a linear relationship between two variables. Values near 1 or −1 indicate a stronger linear pattern; values near 0 indicate a weak one.
The unobserved "what would have happened instead" outcome under a different treatment, policy, or exposure. Causal inference tries to estimate this missing comparison.
A background variable you measure and include in your analysis so it does not distort your main finding. Sometimes called a control variable.
A table that displays the relationship between two or more categorical variables. For example, a crosstab of race/ethnicity × English-learner status shows the average score for each combination.
D
Recorded observations or measurements. In statistics, "data" is technically plural, and a single observation is a datum.
A table-like data structure used in pandas, where rows represent cases and columns represent variables. It is a common working format for cleaning, merging, and analyzing data in Python.
A question focused on what is happening in the data — the level, pattern, or distribution of something — rather than why it happened or what would happen under an intervention. (This project asks a descriptive question.)
A design that compares the change over time in a treated group to the change over time in an untreated group. For example, when a state changes a policy, you compare the before-and-after change there to the change in states that did not change the policy.
An effect that is larger for some groups than for others. If a measurement problem affects one group more than another, it is a differential effect.
How the values of a variable are spread across their range (for example, a bell-shaped or skewed distribution).
E
A standardized measure of how large an effect is (for example, a drop of 0.3 standard deviations), used to compare effects across studies.
The uncertainty or variability in a measurement or estimate (sampling error, standard error, margin of error).
ESSER (Elementary and Secondary School Emergency Relief)
A federal funding program that distributed roughly $190 billion to U.S. schools in three waves between 2020 and 2024 to address pandemic-related needs.
F
In machine learning, an input variable (a column of data) fed into a model.
How closely a model matches the observed data ("goodness of fit").
H
When an AI confidently generates plausible-sounding but false information.
I
IEP (Individualized Education Program)
A legal document under U.S. federal law (IDEA) that describes the specialized instruction and services a student with a qualifying disability will receive. NAEP reports scores separately for students with IEPs.
Two technical meanings: statistical inference is drawing conclusions about a population from a sample; in machine learning, inference is running a trained model to produce predictions on new data.
A variable that affects the predictor but has no direct effect on the outcome, used to estimate causal effects when random assignment is not possible.
The predicted value of the outcome when all predictors equal zero — the point where the regression line crosses the y-axis.
A variable created by multiplying two predictors together, included in a regression to test whether the effect of one predictor depends on the level of another. For example, income × year tests whether the achievement gap between income groups changed over time.
The degree to which a study's design supports the conclusion that the predictor caused the outcome, rather than some other factor.
A research design that compares the trend in an outcome before and after a specific event or interruption (like a policy change, or COVID), looking for a change in the level or the slope at the interruption point.
L
In machine learning, the target value or category a model is trained to predict.
Following the same units over time, with repeated measurements.
M
The arithmetic average: the sum of all values divided by the count.
The middle value when data are sorted — the 50th percentile.
A statistical method that tests whether a third variable (the mediator) explains the relationship between a predictor and an outcome.
A table operation that joins rows from two datasets using a shared key, such as a state name or ID number.
The most frequently occurring value in a dataset.
A change in test scores caused by the format of the test rather than by what students know. (When NAEP switched from paper to digital in 2017, part of any score change could be a mode effect.)
A mathematical or statistical summary of the relationships in data, used to describe or predict.
N
NAEP (National Assessment of Educational Progress)
Often called "The Nation's Report Card," the largest nationally representative assessment of what U.S. students know, administered since 1969 and scored on a 0–500 scale. Its data is public through the NAEP Data Service API at nationsreportcard.gov.
NCES (National Center for Education Statistics)
The federal entity within the U.S. Department of Education that collects and analyzes education data, including NAEP.
Random variation in data that obscures the underlying pattern.
A bell-shaped, symmetric distribution that many statistics assume.
O
Data collected without randomly assigning people, schools, or places to different conditions. Useful for description and prediction, but causal claims are harder because other differences may be mixed in.
OLS regression (ordinary least squares)
The most common form of regression. It fits a straight line through the data by minimizing the sum of the squared vertical distances from each point to the line.
The result a model is trying to describe, predict, or explain. (In this project, the NAEP mathematics and reading means are the outcomes.)
A data point that falls far from the rest.
P
A number that describes a population (for example, the true mean score of all U.S. 8th graders).
The entire group you want to draw conclusions about (for example, all U.S. 8th graders).
Statistical power is the probability that a study will detect a real effect if one exists.
A model-based forecast that uses known information to estimate an unknown outcome.
A question focused on whether known information helps forecast an outcome. Predictive questions rely on useful associations, so a predictor does not have to be a proven cause.
An input variable used to help describe or forecast an outcome in a model. (For example, a student's prior-year score is a common predictor of their current score.)
The text instruction given to an AI model to produce a response.
A variable that stands in for something you cannot measure directly. (Free or reduced-price lunch eligibility is a common proxy for family income.)
R
Each unit equally likely to be selected, with no systematic pattern.
The difference between the largest and smallest values (a measure of spread).
A family of methods for modeling the relationship between variables. Related: regression to the mean, the tendency for extreme values to move toward the average on re-measurement.
Consistency of measurement: would the same instrument give a similar result again?
The difference between an observed value and the value a model predicts.
S
A subset of a population, drawn so you can estimate something about the whole.
The measurement scale or range of an instrument (NAEP's 0–500 scale).
A graph that places one variable on the x-axis and another on the y-axis so each case appears as a point. Useful for seeing direction, clustering, outliers, and whether a linear pattern might exist.
A check on your main finding: you rerun the analysis under different plausible assumptions to see whether the conclusion holds up. (For example, recalculating the COVID-era score decline after subtracting the estimated mode effect.)
"Statistically significant" means a result is unlikely to be due to chance alone (usually a p-value below 0.05). A result can be statistically significant but too small to matter, or large but not statistically significant.
The real, meaningful pattern in the data, as opposed to the noise that surrounds it.
Asymmetry in a distribution: the values pile up on one side with a tail on the other.
In a straight-line model, the amount the outcome is expected to change when the predictor increases by one unit. A negative slope means the outcome tends to go down as the predictor goes up.
A measure of spread: roughly the typical distance of values from the mean.
A number computed from a sample (for example, a sample mean), used to estimate a population parameter.
A structured method of collecting data from a sample, usually via a questionnaire.
T
A way of stating the ideal randomized study you wish you could run, then asking how closely an observational dataset can imitate it.
A unit of text (roughly a word or part of a word) that a large language model processes.
Fitting a model's parameters to data so it can make predictions (the "learning" in machine learning).
The intervention or condition being studied (for example, a tutoring program).
V
Whether an instrument measures what it claims to measure.
A measured characteristic that can take different values (each column in a dataset is a variable).
A measure of spread: the average squared distance of values from the mean.
W
In machine learning, a parameter that scales the influence of a feature on the model's output.
References: Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction. Journal of Personality and Social Psychology, 51(6), 1173–1182. | Hernán, M. A., Dahabreh, I. J., Dickerman, B. A., & Swanson, S. A. (2025). The target trial framework for causal inference from observational data: Why and when is it helpful? Annals of Internal Medicine, 178(3), 402–407. | Ito, C., Al-Hassany, L., Kurth, T., & Glatz, T. (2025). Distinguishing description, prediction, and causal inference: A primer on improving congruence between research questions and methods. Neurology, 104(4), Article e210171. | Kamper, S. J. (2020). Types of research questions: Descriptive, predictive, or causal. Journal of Orthopaedic & Sports Physical Therapy, 50(8), 468–469. | Lopez Bernal, J., Cummins, S., & Gasparrini, A. (2017). Interrupted time series regression for the evaluation of public health interventions: A tutorial. International Journal of Epidemiology, 46(1), 348–355. | Preacher, K. J., & Hayes, A. F. (2004). SPSS and SAS procedures for estimating indirect effects. Behavior Research Methods, Instruments, & Computers, 36(4), 717–731. | Shmueli, G. (2010). To explain or to predict? Statistical Science, 25(3), 289–310.