Sampling with replacement allows each resample to differ from the observed dataset while retaining the dataset as the source of information. Repeating the calculation across these resamples produces a sampling distribution for the chosen statistic. The spread and shape of that distribution provide evidence about how variable the sample-based result may be.
Bootstrapping estimates uncertainty directly from repeated resampling rather than depending on strong assumptions about the population distribution or on an available analytical formula. This distinction matters when the statistic is complex, the sample is small, or the data are not normally distributed. It can therefore extend uncertainty analysis to situations where conventional calculations are difficult.
The resulting sampling distribution can support estimates of standard errors, confidence intervals, and bias. Standard errors describe the estimated uncertainty around a statistic, while confidence intervals express a range based on the resampling results. Examining bias helps researchers assess whether the sample-based estimate systematically differs from the values suggested by the resampled analyses.
First, begin with the observed dataset and repeatedly draw new samples with replacement. Next, calculate the selected statistic for every resample, such as a mean, median, regression coefficient, or correlation. Finally, use the collection of calculated values as a sampling distribution to estimate uncertainty, reliability, standard errors, confidence intervals, or bias.
The approach can be applied to means, medians, regression coefficients, and correlations, rather than being limited to a single type of summary. For each statistic, repeated resampling produces a corresponding set of estimates. Researchers can then evaluate the statistic’s uncertainty or reliability, including cases where a direct analytical calculation is difficult or unavailable.
Bootstrapping is particularly useful when researchers need to quantify uncertainty for complex statistics, small samples, or data that do not follow a normal distribution. Its applications span experimental, biomedical, social, and computational studies. In each setting, the method helps characterize the reliability of sample-based findings when conventional analytical formulas may be difficult to use.