Skip to main content

Big Data Classifier Boosts Representation of Key Groups in Colorado Survey

Innovation Brief
A breathtaking overlook showing lush trees and rugged walls

Author

Patrick Coyle

Statistician

Statistics & Data Science

 

Stas Kolenikov

Principal Statistician

Statistics & Data Science

 

Barbara Fernandez 

Associate Director

Health Care Programs

September 2026

NORC’s predictive tool identified and oversampled key populations, ensuring enough responses to analyze health access across all groups.

The 2025 Colorado Health Access Survey (CHAS) needed to do more than broadly survey Colorado residents about health insurance coverage, health services availability and use, and barriers to health care access. It needed to make sure the subpopulations most central to the research—including uninsured residents, Black and African American residents, and Hispanic and Latino residents—were represented in large enough numbers to analyze. That’s where NORC’s Big Data Classifier (BDC) (Dutwin et al 2024) came in to build a reflective sample of these subpopulations that traditional survey methods can miss. 


Colorado Health Access Survey

Explore the Project


How It Worked

A BDC is a machine learning tool designed to predict characteristics of sampled addresses. It helps researchers achieve more effective oversampling and stratification and supports survey designs that can improve representativeness (Dutwin et al 2024). For the CHAS, we used the BDC in two phases to ensure population groups of specific research interest were accurately represented in the Colorado Health Institute’s (CHI) biennial survey.


A Big Data Classifier helps researchers achieve more effective oversampling and stratification and supports survey designs that can improve representativeness.

The research process started with a large pool of Colorado addresses. We built models predicting characteristics of interest to CHI using data from the NORC AmeriSpeak® panel, U.S. Census data, and third-party commercial marketing data. The predictive models made it possible to identify ahead of time whether a household was likely to belong to one of CHI’s several subpopulations of interest. Based on those predictions, we sorted households into the six groups CHI most needed to understand:

  • Ages 18-29
  • Households with children
  • Ages 65 and older
  • Uninsured
  • Black and African American
  • Hispanic and Latino

All other households were grouped into “Residual” (commercial data available but no positive predictions) or “No match” (no commercial data available).

For the second phase, we drew a stratified random sample, designed to account for what we learned from previous years of the CHAS: different groups respond to surveys at different rates. We also adjusted the second-phase sample to oversample the populations of highest interest.

Colorado is divided into 21 health statistics regions (HSRs), which are essentially geographic areas used for health planning. To ensure we had enough responses from each of our eight population groups in each region, we created 176 sampling strata (21 regions × 8 groups = 168). This allowed us to draw the right number of households from each combination—for example, likely uninsured households in HSR 1 and likely Hispanic and Latino households in HSR 2—ensuring we could analyze health access patterns both by population and by region. 

The Numbers Tell the Story

Comparing the first-phase sample frame to the second-phase sample shows how much oversampling shifted the sample composition among the hard-to-reach groups:

  • Likely Black and African American households represented 2.9 percent of the frame and 9.8 percent of the final sample. In contrast, the 2019-2023 American Community Survey Public Use Microdata Sample can be used to estimate that 4.7 percent of households in Colorado have a Black or African American member (U.S. Census Bureau 2025).
  • Likely uninsured households represented 3.9 percent of the frame and 11.6 percent of the final sample.
  • “Residual” and “no match” households, by contrast, dropped as a share of the sample (to 25.5 percent and 9.8 percent, respectively). Reducing their sample sizes allowed the sampling team to increase the sample sizes for the groups the survey needed most, such as Black and African American households.

All told, NORC sampled 74,120 addresses. We conducted the survey in three waves and used Wave 1 results to fine-tune sampling rates in later waves.

Why Bother with a BDC in the First Place?

Traditional ways of finding specific populations, such as geographic clustering or licensing vendor “flags”—pre-labeled indicators that households might belong to certain groups (English et al 2017, Barron et al 2015)—only go so far. A BDC provides better results by using machine learning­—gradient-boosted decision trees, specifically XGBoost (Chen and Guestrin 2016)—to combine public and commercial data to predict household characteristics ahead of time.

Testing across multiple datasets has shown BDCs generally outperform geographic clustering and hold their own against, or do better than, vendor flags when it comes to finding low-incidence and hard-to-reach populations (Dutwin et al 2024). This translates into more efficient samples and often lower costs per completed interview.


Testing across multiple datasets has shown BDCs generally outperform geographic clustering and hold their own against, or do better than, vendor flags when it comes to finding low-incidence and hard-to-reach populations.

Oversampling Doesn’t Break Representativeness—If You Weight for It

While some groups were sampled at higher rates, we weighted the data to ensure the final numbers accurately reflected the population. We also adjusted the data for response patterns and matched them to population benchmarks at both the HSR and state levels. The study included 10,922 completed interviews and provided enough responses from key subpopulations to make meaningful conclusions about their health access.

The Fine Print

Commercial data do not cover every household equally. Broader research has found lower vendor match rates for Spanish-speaking, Hispanic, and Black households compared to the general population (DiSogra et al 2010). This can introduce bias if such underrepresentation is not accounted for. The fix is to use benchmarks such as the American Community Survey for sample planning and to correct through raking—an algorithm that ensures the survey weights for groups (such as Hispanic respondents) match population totals derived from a census or another survey.

The Bottom Line

Stratifying with the BDC before sampling and then weighting the data enabled us to increase representation of subpopulations while still accurately reflecting the state of Colorado as a whole. Using the BDC process helped serve the client’s analytical objectives and the interests of survey stakeholders such as county health departments.

The BDC-assisted sampling process uniquely positions NORC in the survey industry. No other contract research organization has the functional capacity to draw samples of mailing addresses and augment them with predictions for a rich set of potential household characteristics.


Colorado Health Access Survey

Explore the Project


References

Barron, M., Davern, M., Montgomery, R., Tao, X., Wolter, K., Zeng, W. K., Dorell, C. and Black, C. (2015). Using Auxiliary Sample Frame Information for Optimum Sampling of Rare Populations. Journal of Official Statistics, 31, 545–557.

Chen, T. and Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, California, USA: ACM, pp. 785–794.

DiSogra, C., Dennis, J. M. and Fahimi, M. (2010). On the Quality Of Ancillary Data Available For Address-Based Sampling. Proceedings Of The American Statistical Association, Section On Survey Research Methods (Vol. 417483).

Dutwin, D., Coyle, P., Lerner, J., Bilgen, I. and English, N. (2024). Leveraging Predictive Modelling from Multiple Sources of Big Data to Improve Sample Efficiency and Reduce Survey Nonresponse Error. Journal of Survey Statistics and Methodology, 12, 435–457.

English, N., Allen, M. and O’Muircheartaigh, C. (2017). Using Commercial Data to Enhance Survey Eligibility: The AmeriSpeak Experience. 2017 Proceedings of the American Statistical Association, Survey Research Methods Section, Alexandria, VA: American Statistical Association.

U.S. Census Bureau (2025). 2019-2023 American Community Survey 5-year Public Use Microdata Sample (PUMS).


Suggested Citation

Coyle, P., Kolenikov, S. & Fernandez, B. (2026, September 21). Big Data Classifier Boosts Representation of Key Groups in Colorado Survey. [Web blog post]. NORC at the University of Chicago. Retrieved from www.norc.org.


Tags

Research Divisions



Solutions

Explore NORC Research Science Projects

Analyzing Parent Narratives to Create Parent Gauge™

Helping Head Start build a tool to assess parent, family, and community engagement

Client:

National Head Start Association, Ford Foundation, Rainin Foundation, Region V Head Start Association

America in One Room

A “deliberative polling” experiment to bridge American partisanship

Client:

Stanford University