Home » Posts tagged 'data and statistics' (Page 8)

Tag Archives: data and statistics

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

The 1967 Census of the West Bank and Gaza Strip: A Digitized Version

Here is an interesting set of census data, digitized from print. The dataset documentation (The 1967 Census Of The West Bank And Gaza Strip: A Digitized Version Dataset Documentation, Levy Economics Institute of Bard College, Joel Perlmann, Project Director, November 2011) also presents a detailed description of the methods used to create accurate, usable, digital statistical tables from print.

  • The 1967 Census of the West Bank and Gaza Strip: A Digitized Version, Levy Economics Institute of Bard College. In the summer of 1967, just after the Six-Day War brought the West Bank and Gaza Strip under Israeli control, the Israeli Central Bureau of Statistics (ICBS) supervised a census in these territories. The census included an impressive array of questions about individuals, households, and the quality of residences—about age, sex, religion, place of residence, educational attainment, occupation, industrial sector, income, household structure, health, female fertility, and housing conditions. Moreover, it asked two crucial questions about refugee status: Had the individual lived prior to the 1948 War in the area that became the State of Israel? And, Was the individual living in or outside of a refugee camp at the time the census was taken? The ICBS prepared seven volumes of reports based on this enumeration—the first modern census reports on the Palestinian population. Yet these volumes have not been used extensively in the writing on the evolution of the occupied territories. One reason is that they are not widely available, and even when at hand they are subject to all the limitations of older volumes published as quickly as possible in order to be of use at the time. For this reason, the Levy Economics Institute is offering, for the first time and free of charge, the content of six of these volumes (the seventh to be added shortly) in machine-readable form, in the hope that the data can be exploited by researchers interested in a fuller understanding of the social history of the Palestinian people in the occupied territories. Bard student volunteers contributed appreciably to this project.
  Continue reading

Continue Reading →

Health Indicators from HHS Community Health Data Initiative

"The Health Indicators Warehouse (HIW) is a new resource serving as the data hub for the HHS Community Health Data Initiative. It contains standardized health outcome and health determinant indicators along with associated evidence-based interventions, which can be easily displayed, and will benefit a broad variety of users."

  • Health Indicators Warehouse (healthindicators.gov). The HIW is a collaboration of many Agencies and Offices within the Department of Health and Human Services. The HIW is maintained by the CDC's National Center for Health Statistics. Data, support and funding are provided by the following: * Centers for Medicare & Medicaid Services * Department of Health and Human Services: o Office of the Deputy Secretary o Office of Adolescent Health o Office of Disease Prevention and Health Promotion o Office of Minority Health o Office of the Assistant Secretary for Planning and Evaluation * Health Resources and Services Administration. New partners are expected to be added over time.
  • "Health Indicators Warehouse Provides Data Accessibility." In: COSSA Washington Update, Consortium of Social Science Associations, Volume 30, Issue 3 (February 7, 2011) [NHS's Linda] Bilheimer noted that the Health Indicator Warehouse includes Healthy People 2010, a new set of Medicare indicators, national and state indicators. She emphasized that this is the first time that these indicators have been put in the public domain. She noted that HIW has extensive metadata. HIW is oriented toward large scale developers who want to use the data. [NHS's Amy] Bernstein reported that there are 1,130 indicators listed from 170 data sources
Continue reading

Continue Reading →

Statistical reality check

As noted here earlier, Sunlight Labs has announced the three finalists in its "Apps for America 2" competition. One of those is DataMasher which enables users to "have a little fun" with government data "by creating mashups to visualize them in different ways and see how states compare on important issues. Users can combine different data sets in interesting ways and create their own custom rankings of the states." A post about this on slashdot prompted a reply, Lies, Damned Lies, and DataMasher that worries that "in practice DataMasher would end up mostly generating a lot of bad information." The reply continues:

The site as it exists now seems to encourage you to think about issues in a really simplistic way (with a simple arithmetic combination of two numbers on a state by state basis) that's going to mislead more often than inform. The devil is always in the spurious correlations, and DataMasher just doesn't give you ability to get at that sort of thing (nor do most people have the understanding of statistics anyway). ...Statistics are extremely useful in determining public policy, but only if used carefully. There's already so much bad use of statistics in our public policy debates, and DataMasher seems perfectly designed (unintentionally, I'm sure) to exacerbate the problem.
I am very sympathetic to this argument but would add an additional caveat to it. Any tool can be misused or used badly. Decades ago, some statisticians were upset when commercial software like SAS and SPSS were being introduced because it allowed anyone to run a regression without knowing what it was or how to do the math or whether they were regressing variables that made sense. While it is certainly true that the design of tools can encourage misuse or bad use, it is also true, I think, that even well-designed tools can be used badly and even bad tools may be better than no tools because they can encourage imagination and exploration and curiosity. Those can lead to better, more informed questions and analysis. For libraries and service providers there is another side to this story. As tools like DataMasher become more available and easier to use, it actually creates new challenges for information service providers. Rather than making our jobs easier, the availability of these kinds of tools actually makes our jobs more complex. Rather than pointing at a reliable book of statistics, created by government statisticians and published by the government, we now have 'raw' sources and sources that require more understanding and skill to use and interpret accurately and responsibly. Where once we tried to make sure that the people we helped looked at footnotes and table headers so they understood statistics, now we are faced with helping people use raw data and helping them produce their own statistics. Every library will have to decide on what level of service to provide in situations like this. No library should avoid addressing the service implications of the availability of new sources of information -- no matter how good or bad they are. Continue reading

Continue Reading →

Job losses in recessions visualized

Recently this graphic was posted on the "The Gavel" blog of the Speaker of the House (What 3.6 Million Jobs Lost Over 13 Months Looks Like, by "Karina," February 6th, 2009). It shows the number of jobs lost (and recovered) in the recessions of 1990, 2001, and 2008. By juxtaposing the three time periods over each other starting with the peak job month and showing employment change by month, it gives a startling comparison that highlights the severity of the current situation. It also implies that we will have a long time to wait before we reach our previous peak month.

I was curious about this graph and did a little follow up that I share below. Data librarians may find this a bit tedious, but for those who have never used raw data, it may be useful as an illustration of the difference between "data" (the raw numbers that you put into statistical software) and "statistics" (the human-viewable tables and graphs that we see in publications).

Unfortunately, as is often the case with statistical information, the source given for the graphic is incomplete: simply "Bureau of Labor Statistics." I could not find the graphic itself on the bls.gov site, so I assume that the chart was constructed from BLS data, specifically, the Current Population Survey or the Current Employment Statistics Survey. These two surveys count employment differently -- one is a survey of individuals and the other is a survey of employers.

There is a similar, but not identical, chart ("Percent change in total nonfarm employment, from beginning of recession) in the January 2009 (released February 6, 2009) Current Employment Statistics Highlights, Monthly (Bureau of Labor Statistics), so my guess is that someone at the Speaker's office built the chart from the raw CES data.

Just out of curiosity, I went to the CES "Most Requested Statistics webpage and downloaded "Total Nonfarm Employment - CES0000000001" for 1990 through the end of 2008. Raw data suitable for analysis even look "raw," not even like a statistical table:

1990,Jan,109151
1990,Feb,109396
1990,Mar,109611
1990,Apr,109651
1990,May,109800
1990,Jun,109817
1990,Jul,109775
1990,Aug,109567
1990,Sep,109485
1990,Oct,109324
1990,Nov,109180
1990,Dec,109120
1991,Jan,109001
1991,Feb,108695
1991,Mar,108535
1991,Apr,108324
1991,May,108196
1991,Jun,108283
...

Of course, it is relatively easy, using statistical software, to construct tables and graphs from raw data. Here, for example, is a published statistical table with essentially the same raw information (but from CPS, not CES) that I downloaded. (See the full table from Employment from the BLS household and payroll surveys: summary of recent trends, February 6, 2009).

By using the raw data to create a graph, one can tell a story that has more impact than just a table of numbers. It is relatively easy to get these data into statistical software. I used Excel and Stata to create a small time-series data file. I organized it by month (from month "1" to month "48") with each row of the data file having data for 3 recessions. The first row has data for the first month of the three recessions, the second row has data for the second month, etc. The CES data has employment totals in millions. For example, the employment for 2008:

 
138152 
138080 
137936 
137814 
137654 
137517 
137356 
137228 
137053 
136732 
136352 
135755 
135178 

I had to compute a new variable for each recession: the cumulative number of jobs lost. So, for example, 2008:

138152	 0
138080	-72
137936	-216
137814	-338
137654	-498
137517	-635
137356	-796
137228	-924
137053	-1099
136732	-1420
136352	-1800
135755	-2397
135178	-2974

The first 12 months with all three recessions (v1, v2, v3) and the computed variables (1990, 2001, 2008) look like this:

month   v1  1990    v2      2001    v3       2008
1   109817    0     132530    0     138152   0
2   109775  -42     132500  -30     138080  -72
3   109567  -250    132219  -311    137936  -216
4   109485  -332    132175  -355    137814  -338
5   109324  -493    132047  -483    137654  -498
6   109180  -637    131922  -608    137517  -635
7   109120  -697    131762  -768    137356  -796
8   109001  -816    131518  -1012   137228  -924
9   108695  -1122   131193  -1337   137053  -1099
10  108535  -1282   130901  -1629   136732  -1420
11  108324  -1493   130723  -1807   136352  -1800
12  108196  -1621   130591  -1939   135755  -2397

Here is a complete tab-separated-values version of the data file I constructed. Then I used Stata to build a graph and it looks very much like the one at the Speaker's Blog.


Of course, when one tells one story, one leaves out other stories. This graphic doesn't show that the starting points of the recessions were different: 1990: 109 million 2001: 132 million 2008: 138 million

Open re-usable government information

One could use the raw data to tell a lot of different stories and analyze the data in many different ways. And that brings me to the connection between all this and why we need to be sure that government information is not just "free as in beer" but also "free as in open."

It is important for statistical agencies to publish statistics to help us understand their raw data. But, it is also essential that they provide us with the raw data so that we can better understand their statistics and do our own analyses. Most of the statistical agencies of the U.S. government do an excellent job of making their raw data easily available. In fact, the rest of government would do well to use statistical agencies as a model for instantiating their information in usable and re-usable formats (in addition to any publishing and presentation of their data/information) so that the information, whether it is text or images or video or sound or numbers, can be used, reused, analyzed, stored, and preserved.

Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM