Small Area Estimation Unlocks Local Health Data from State Surveys
Author
Principal Statistician
Statistics & Data Science
Associate Director
Health Care Programs
August 2026
NORC produced reliable county- and subcounty-level health estimates in Colorado using a statewide survey designed to measure only 21 larger regions.
State health surveys designed to provide a comprehensive statewide picture may not necessarily provide reliable data for every county or subcounty. The problem is especially acute for small, rural counties where getting enough survey responses can be impractical and expensive. But county health departments and other local organizations need local data to make decisions.
NORC faced this challenge for the 2025 Colorado Health Access Survey (CHAS). Using small area estimation techniques, we produced improved health access estimates for the 64 counties in Colorado, even though the survey was designed to only obtain direct estimates for 21 larger Health Statistics Regions (HSRs).
Since 2021, NORC has conducted CHAS biennially for the Colorado Health Institute, interviewing more than 10,000 Coloradoans about health insurance coverage, health services availability and use, and barriers to health care access. The survey targets at least 400 interviews in each HSR, which are geographic areas consisting of large counties or groups of smaller counties. The county-level estimates for this year are new, and this data product represents a major expansion of the survey’s utility for local decision-makers.
What is small area estimation, and how does it work?
Small area estimation (SAE) is a set of statistical techniques that fall in the broad category of regression and prediction modeling. It helps researchers make accurate predictions about groups when their sample sizes are insufficient for reliable inference using the survey data alone. Colorado’s 25 smallest counties each have a population of less than 8,000 people, and their sample sizes in CHAS are 10 or fewer respondents, not nearly enough data for analytic tasks. SAE uses statistical techniques to fill the gaps in data and provide trustworthy insights even when survey sample sizes are tiny, and sometimes zero.
The term “small area” comes from early uses of the method to study small geographic areas, such as individual counties, places, and school districts. Today, the models are equally applicable in research of other small samples, including demographic groups in household surveys or companies in establishment surveys broken down by detailed economic categories such as size, ownership, industry, and location.
NORC used statistical models to blend survey responses with trusted data sources to estimate health care access for even the smallest counties.
We combined survey responses with other data available at fine geographic levels. NORC drew on nearly 600 variables from the Agency for Healthcare Research and Quality (AHRQ) Social Determinants of Health database, which combines American Community Survey (ACS) data with information from Medicare, health facility databases, and other trusted sources. We sifted through a wide range of statistical models to identify which factors best predicted health outcomes and then used the strongest predictors to produce more precise county-level estimates.
Our final estimates blend the survey data from each county with predictions based on these external data sources, weighted according to their reliability. Counties with more survey responses rely more heavily on their own data; counties with fewer responses draw more from the statistical model predictions.
NORC incorporated four cutting-edge techniques to maximize the accuracy and usefulness of these county estimates.
- Multi-level modeling: We built statistical models at the census-tract level (Colorado’s 1,250-plus neighborhoods), then aggregated results to produce county estimates. This approach captures local variation while ensuring county-level totals are reliable. It also allows the model to recognize that tracts within the same county share certain characteristics.
- Automated variable selection: From nearly 600 available predictors, we used a technique called lasso regression—originally developed for genomics research—to improve prediction accuracy by automatically identifying the small set of factors that best predict each health outcome.
- Benchmarking to state totals: We calibrated county estimates so they aggregate to match the statewide CHAS estimates. This ensures consistency across geographic levels and builds confidence in the county data.
- Bayesian modeling for uncertainty: We used Stan, a Bayesian statistical programming language, to properly calculate uncertainty intervals around each county estimate. As models become more sophisticated, traditional statistical formulas for uncertainty become inadequate. Bayesian approaches handle this complexity naturally, giving users appropriate measures of precision for each county estimate.
Policy Implications
The Colorado Health Institute distributed the estimates to all 64 counties. Health officials and other local health professionals now have access to CHAS data at a more granular level, with some data available for the first time ever at the county level. The local health decision-makers are enthusiastic about these new data. While there are assumptions in the estimates, having them is still much better than having no or minimal data, or only data at the regional level of the HSRs.
These more granular data allow county health departments to:
- Identify specific access barriers in their communities
- Make data-driven decisions about resource allocation
- Compare their performance to similar counties
- Bolster funding applications with local evidence
For the Colorado Health Institute, the county estimates strengthen relationships with local government partners by providing actionable intelligence tailored to their specific constituencies.
The broader lesson for public health researchers: With carefully designed statistical models, state and regional surveys can reliably produce estimates at more detailed geographic levels than their original sample designs targeted. This approach is especially valuable when designing a survey in which directly sampling all small areas is cost prohibitive.
Suggested Citation
Kolenikov, S. & Fernandez, B. (2026, August 18). Small Area Estimation Unlocks Local Health Data from State Surveys. [Web blog post]. NORC at the University of Chicago. Retrieved from www.norc.org.