Methodology
Data sources
The statistics on this site are drawn from three Social Security Administration datasets: the national baby names file, which records name, sex, and count by year back to 1880; the state-level version of that file, available from 1910 onward; and the SSA's period life table, which provides survival probabilities by age and sex. The Social Security Administration excludes any name given to fewer than five babies in a given year, so a sufficiently rare name will display a “limited data” notice rather than a full profile.
Living-age distribution
This is the point most visitors misunderstand: a name's age distribution today is not the same as it was during the name's period of popularity. Each birth-year cohort is weighted by its survival probability to the present, using the number-of-lives column (lx) from the SSA life table, calculated separately by sex. A name that peaked in 1950 has had 75 years for its cohort to age and, actuarially, thin out; a name that peaked in 2015 has not. The median and 15th/85th percentile ages shown come from this survival-weighted distribution, rather than from raw birth counts.
Popularity arc & trend archetype
Popularity is measured as a name's share of all births in a given year rather than as a raw count, which allows for fair comparison across eras with very different total birth volumes. The archetype badge is generated through k-means clustering on several shape features of each name's popularity curve: the recency of its peak, the duration it remained active, the proportion of peak strength that persists today, and whether the curve shows a decline followed by a resurgence. Each resulting cluster is matched to the closest of six predefined archetypes.
Name neighbors
Name neighbors are computed using cosine similarity between full-resolution yearly popularity curves, identifying names whose rise and fall tracked one another closely over the decades. Simple spelling variants are filtered out deliberately; without this step, most results would consist of alternate spellings of the same name rather than genuinely related names.
Geographic fingerprint
For each state and Census region, a name's share of local births is compared to its share of national births over the same period, from 1910 to the present, matching the availability of state-level data. An index of 2.0 indicates that a name appears twice as frequently in that state as would be expected based on the national baseline.