Praxis Variable Guide · PVG-002

Core NHANES Demographic Variables

A connected guide to participant status, age, gender, race and Hispanic origin, income-to-poverty, weights, and design

Audience
Researchers using NHANES demographics, weights, and design variables
Reviewed
Format
One-page PDF · US Letter

Guide overview

The essential idea

The NHANES demographics file is more than a list of participant characteristics. It supplies the participant key, release and examination status, common demographic variables, full-sample weights, and masked variance variables that connect file construction to valid population inference. Definitions and coding can vary by cycle, so the cycle-specific codebook remains authoritative.

Practical framework

What to apply

Identity and participant status

Which record, release, and participation level are represented?

  • SEQN is the respondent sequence number used to connect participant-level files.
  • SDDSRVYR identifies the NHANES data release cycle; DEMO_L uses value 12 for August 2021-August 2023.
  • RIDSTATR distinguishes interviewed-only participants from those both interviewed and MEC examined.
  • Use these fields to retain provenance and define participant flow before analyzing component measures.

Age and gender

How are common personal characteristics released?

  • RIDAGEYR is age in years at screening; in DEMO_L, people aged 80 and older are top-coded at 80.
  • RIDAGEMN provides age in months for participants aged 24 months or younger under the documented conditions.
  • RIAGENDR is labeled Gender in DEMO_L and is released as 1 Male and 2 Female.
  • Preserve official labels and coding in the source map, then document any analytical categories or terminology separately.

Race, Hispanic origin, and income context

Which recode supports the intended cycle comparison?

  • RIDRETH3 includes a non-Hispanic Asian category and has been used since 2011-2012.
  • RIDRETH1 groups non-Hispanic Asian participants with other non-Hispanic races and supports linkage to the 1999-2010 recode structure.
  • INDFMPIR is family income divided by the applicable poverty guideline and is top-coded at 5.00 or greater.
  • Missingness, category meaning, disclosure recodes, and historical comparability must be considered explicitly.

Weights and variance design

Which fields make national inference possible?

  • WTINT2YR is the full-sample interview weight and WTMEC2YR is the full-sample MEC examination weight for the cycle.
  • Subsample components can supply more restrictive weights in their own files.
  • SDMVSTRA is the masked variance pseudo-stratum and SDMVPSU is the masked variance pseudo-primary sampling unit.
  • Select the weight for the smallest applicable sample and specify strata and primary sampling units in survey-capable software.

Before you finish

A short quality review

  1. SEQN, SDDSRVYR, and RIDSTATR are retained through data construction.
  2. Age top-coding and any age-group derivation are documented.
  3. RIAGENDR is described according to the source codebook and analytical relabeling is transparent.
  4. RIDRETH1 or RIDRETH3 is selected deliberately for the cycles and categories required.
  5. INDFMPIR missingness and the 5.00-or-greater top-code are addressed.
  6. The correct weight, SDMVSTRA, and SDMVPSU are used for the final variable set.

Tie demographics to design. The same file that describes participants also carries the weights and masked design variables needed to describe the population. Demographic interpretation and survey inference should therefore be planned together.

Authoritative sources

Continue with the primary guidance