home / fivethirtyeight / most-common-name/adjusted-name-combinations-matrix

Menu
  • GraphQL API

most-common-name/adjusted-name-combinations-matrix: 14

This directory contains the code and data behind the story Dear Mona, What’s The Most Common Name In America?

The main script file is most-common-name.R

There are four input files:

  • state-pop.csv - Total population and Hispanic population by state.
  • surnames.csv - Data on surnames from the U.S. Census Bureau, including a breakdown by race/ethnicity.
  • aging-curve.csv - Data from the Social Security Administration on the chances that someone born in the decade shown was still alive in 2013: http://www.ssa.gov/oact/NOTES/as120/LifeTables_Tbl_7.html
  • adjustments.csv - Taken directly from Lee Hartman's article: http://mypage.siu.edu/lhartman/johnsmith.html.

And five output files:

  • adjusted-name-combinations-list.csv - Adjusted estimates for the most common full names.
  • adjusted-name-combinations-matrix.csv - The same data from the file adjusted-name-combinations-list.csv but in matrix form. These are the estimates presented in the second (and final) table of the article.
  • independent-name-combinations-by-pop.csv - Matrix of estimates for the top 100 most common first names by top 100 most common surnames. These were calculated using independent odds, and displayed in the first table presented in the article.
  • new-top-firstNames.csv - Final estimated ranking of top first names.
  • new-top-surnames.csv - Final estimated ranking of top surnames.

Data source: https://github.com/fivethirtyeight/data/blob/master/most-common-name/adjusted-name-combinations-matrix.csv

This data as json, copyable

rowid Unnamed: 0 FirstName SMITH JOHNSON WILLIAMS BROWN JONES GARCIA RODRIGUEZ MILLER MARTINEZ DAVIS HERNANDEZ LOPEZ GONZALEZ WILSON ANDERSON THOMAS TAYLOR LEE MOORE JACKSON
14 42 Jennifer 12786.8094162309 9587.85619148557 6760.17833327087 6890.91455180973 7942.80482425029 3734.85198512614 3493.64115496868 6383.10044846552 3407.51448804857 5433.59450854971       4168.96851485375 4103.42502933492 3317.17710838062 3845.88070289281 4069.44260540816 3611.28988724838 2994.82488550718
Powered by Datasette · Queries took 1447.886ms · Data source: https://github.com/fivethirtyeight/data/blob/master/most-common-name/adjusted-name-combinations-matrix.csv