NARA looks to privatizing 1940 Census

The National Archives is apparently preparing to reverse a long standing policy of providing free public access to Census Schedules when it releases the 1940 Census next year. (See "Update" below for additional information.)

Early next year the National Archives and Records Administration (NARA) can make the 1940 Census Schedules available to the public for the first time. (See "Background" below.) NARA has digitized these files and created metadata for them in preparation for making this valuable trove of information accessible on the web. It now only remains for NARA to decide who should provide access at what cost to users. Should NARA provide free access? Or should it contract this service out to a private company that will imposes fees on users and make a profit by providing access to this public information?

The answer should be obvious. For decades NARA has provided free access to Census Schedules at regional archive facilities and has sold microfilm to libraries that provide free access to their users. Now that online digital access is possible, NARA can provide better access online without having to maintain physical access at its regional facilities. It can distribute digital files to libraries for little or no cost so that libraries could further increase access and functionality for all of the information or subsets of it.

It seems, however, that NARA is choosing privatization instead of free public access. In an eight page RFI (Services Request for Information (RFI) NAMA-11-RFI-0004, 1940 Census [Microsoft Word .docx] or see the PDF version] NARA is seeking "industry input" for a "no-cost contract" to provide managed hosting and online access to the 1940 Census. The RFI is intended to explore options and "may or may not lead to a solicitation" for an actual contract. This means that NARA could, apparently, make a decision to do this work itself, but it is exploring the privatization route first. (Presumably, libraries could respond to the no-cost RFI as well. Responses were due on June 22.)

According to NextGov, the "no-cost contract" means that "the vendor would do the work for free and then charge the public a fee to access the records." (Archives Wants to Put 1940 Census Online, by Joseph Marks, NextGov TechInsider, July 15, 2011

The RFI does not explain why NARA is pursuing this path or what advantages it sees to privatization. I would guess the most likely reason is that NARA does not anticipate that it can get adequate funding to host the data online itself. But has it asked for funding? Has it made the case for continuing its historic provision of free public access to Census Schedules? Has it justified privatizing public information?

We have seen NARA follow this path before. NARA partnered with footnote.com to digitize selected holdings. This resulted in access restrictions including membership fees, per-page charges for downloading, and age restrictions to these digitized public documents. NARA partnered with Ancestry.com to make public records available for a fee. At the time, a spokesperson said that budget constraints and other priorities kept the Archive from making this information available itself.

"In a perfect world, we would do all this ourselves and it would be up there for free," she said. "While we continue to work to make our materials accessible as widely as possible, we can't do everything." -- Ancestry.com unveiled more than 90 million U.S. war records, New York Times (May 24, 2007).

In 2008, NARA contracted with TGN to digitize and provide access to some of NARA's holdings. The contract restricted free public access for five years. We've written about this here at FGI before (The NARA/TGN contract as a bad precedent) and believe that deals like this remain bad for NARA and bad for the country. We believe these kinds of deals set a bad precedent -- a precedent that is now being unnecessarily followed with the 1940 Census.

Those past deals were different from this one in one key way, however. They involved digitization of materials by the private contractors. In the case of the 1940 Census, the materials are already digitized, according to the RFI. The arguments we heard in the past were that digital access was so much better that it was worth privatizing access in order to get the digitization done. Without privatization, it was argued (even by some librarians and archivists), the materials could not be digitized and we'd be stuck with analog access. This is not the case for the 1940 Census since the materials are already digitized. The decision for online digital access has been made. The only question now is whether to make the existing digital files freely available or available through privatization.

Of course, the cost of any project providing access to all the 1940 Census Schedules and maps will not be insignificant. According to the RFI, NARA has created 3.8 million JPEG images, comprised of 20 terabytes of data.

Twenty terabytes is a lot of data, but it is fast becoming an almost modest size for a digital library. For comparison, the HathiTrust has over 3 billion pages and over 400 terabytes of data, OCLC has an over 600 Terabyte capacity, the Wayback machine contains 100 terabytes, the Library of Congress web archive is 235 terabytes, the University of California Curation Center has 70 TB in its Merritt digital preservation repository, and NARA's own Electronic Records Archives (ERA) has more than 90 terabytes. These terabyte-scale digital libraries are virtually the new norm and petabyte-scale digital libraries are already being built. Some of these are the Shoah Foundation Institute's digital library (8 petabytes), the Stanford Digital Repository anticipating a capacity of petabytes, and the Digital Hammurabi Project which is building a petabyte-scale digital library and museum of virtual 3D cuneiform tablets.

But the cost of providing access to microfilm at 13 regional offices was not cheap either. To me the question is whether or not the government is willing to continue its historic mission of providing free access or if it is ready to abandon that mission to the private sector.

There have always been those who argue that the fee-based private sector should take precedence over the public sector free-access. But, privatization of access to Census schedules would represent a reversal of long-standing policy. What was once an unquestioned government function is now, apparently, being considered a commercial function. Where, in the past, it was the government that provided free, public access to census schedules, now, when access can be improved, the government is abrogating its role and turning access over to private companies that will provide the information for a fee. The issues are not new. The precedents for providing free public access exist and have a long and respectable history. The only thing new is that NARA seems to have accepted privatization as inevitable.

Will there be funding for NARA to provide access to 1940 Census Schedules? There may not be. We have argued here at FGI for years that relying solely on Congressional funding for permanent, free public access to government information is risky because there is always the chance that Congress will not fund it. In these highly-politicized, economically troubled times it is easier to imagine a lack of any funding that to imagine adequate funding for the long term.

But this does not mean that privatization is the only option. There are precedents for government projects that are supported by donations and public-private partnerships. The American Memory Project is one notable example. And individual libraries or groups of libraries could step in and offer to provide free public access.

Now is the time for NARA, supported by researchers, libraries, and archivists to actively promote and pursue free public access solutions. There is no reason to accept privatization as inevitable.

Update: Note that the NARA website 1940 Census page says that "The digital images will be accessible at NARA facilities nationwide through our public access computers as well as on personal computers via the internet." Additionally, a comment on the Ancestry World web site said: "NARA will make the digitized copies of the 1940 Census population schedules available to the public, free of charge, on April 2, 2012 through our new Online Public Access search (http://www.archives.gov/research/search/)." It is not clear from the above if the policy on free public access has changed with issuance of the RFI or not.

Background:
The Census Bureau conducts the Decennial Census every ten years. The Bureau summarizes its findings in reports that contain no information on individuals. The raw information collected, including names and addresses of those surveyed (sometimes called the "manuscript census" or the "census schedules"), is protected by law and is kept confidential for 72 years. After 72 years, that raw information is released by the National Archives and Records Administration. This information is invaluable to genealogists and other researchers. Typically, this raw information has been made available on microfilm at regional National Archives offices (Availability of Census Records About Individuals). This 72 year period for the 1940 Census expires on April 2, 2012.

Continue reading

Continue Reading →

Two Books on Control of the Internet

FGI volunteer ShinJoung Yeo reviews two books about global political struggles to govern the world's distributed communication infrastructure.

  • Access Controlled: The Shaping of Power, Rights, and Rule in Cyberspace; Networks and States: The Global Politics of Internet Governance, by ShinJoung Yeo, JASIST, Volume 62, Issue 8, pages 1647-1649, August 2011. Article first published online: 9 MAY 2011. [subscription required] A recent series of events--Google's dispute with China, Secretary of State Hillary Clinton's speech on Internet freedom, the Egyptian government shutting off nearly all Internet services during the 2011 pro-democracy revolution, .xxx domain approval by ICANN after much political controversy--is indicative of the heightening global politics surrounding the Internet. Two books--Access Controlled: The Shaping of Power, Rights, and Rule in Cyberspace, edited by Ronald J. Deibert, John G. Palfrey, Rafal Rohozinski, and Jonathan Zittrain, and Networks and States: The Global Politics of Internet Governance, by Milton L. Mueller--have drawn attention and provide context to exactly these global political struggles to govern the world's distributed communication infrastructure, increase governments' efforts to control, and reassert sovereignty rights over cyberspace by nation states. Both books contribute to explicating the complex tensions between nation states and the extraterritorial nature of the Internet. ...If one has not yet been convinced that the Internet is far from value neutral, once again these books corroborate and stress the fact that cyberspace has grown ever more tightly intertwined with global political economy and has become a site of political, economic, and cultural struggle among nation states.
Continue reading

Continue Reading →

The Internet as Memory

Earlier today I posted a couple of favorite quotes about the role of libraries as institutions that hold in our collective memory things that would otherwise be forgotten. A short article in Scientific American notes the importance of "the internet" as "external memory" or "transactive" memory:

  • Piece of Mind: Is the Internet Replacing Our Ability to Remember?, by Larry Greenemeier, Scientific American (July 14, 2011). ...[T]he Internet has become a primary form of external or "transactive" memory ... where information is stored collectively outside the brain. This is not so different from the pre-Internet past, when people relied on books, libraries and one another ... for information. Now, however, besides oral and printed sources of information, a lion's share of our collective and institutional knowledge bases reside online and in data storage.
The researcher says that "Information is much more available than it was." An interesting article, but it misses, I think, the key point that if no one preserves "the internet" (or the parts of it that we want to preserve), this transactive memory won't be there for us to use. (The article even makes this amazingly naive statement: "And if our gadgets were to fail due to a planet-wide electromagnetic pulse tomorrow, we would still be all right.") This is important because, if we think our "external" memory is safe and we rely on commercial interests to preserve that information, then we are leaving our very memory at the commercial mercy of those companies. An alternative is for communities of interest to rely on libraries to preserve important information. A related article about the same research: Continue reading

Continue Reading →

NSA wins FOIA battle over request for information on any work with Google

In a recent court decision the National Security Agency acknowledged working "with a broad range of commercial partners and research associates" but it obtained from the court the right to keep secret documents that may, or may not, show a working relationship with Google.

  • NSA spooks win fight to keep secret possible ties to Google, by Mike Doyle, Suits & Sentences legal affairs blog, McClatchy Newspapers (July 13, 2011). In a decision made public Wednesday, U.S. District Judge Richard J. Leon denied a Freedom of Information Act request filed by the curious souls at the Electronic Privacy Information Center. EPIC sought documents relating to NSA's possible relationship with Google following news of an alleged cyber attack by hackers in China and of a subsequent cooperation agreement between Google and NSA.
  • EPIC v. NSA: Agency Can "Neither Confirm Nor Deny" Google Ties, EPIC (July 13, 2011).
Continue reading

Continue Reading →

It Takes a Village…to Archive the Internet

From the Library of Congress:

  • It Takes a Village…to Archive the Internet, guest post by Abbie Grotke, Web Archiving Team Lead at the Library of Congress (posted by Mike Ashenfelder), The Signal, digital preservation blog, Library of Congress (July 14th, 2011). Despite the tremendous amount we’ve preserved, we know we can’t do it alone. We often collaborate to build web archives with other libraries, archives and organizations in the United States and around the globe. We do this when events unfold quickly on the Internet and the Library can’t react as quickly as we’d like for whatever reason, but also when the scope of a collection is so big that we must work with others to ensure that breadth of content is preserved. ...So while each institution has its own collection policies to follow, there is an obvious recognition by web archivists that the Internet is a global place, not neatly wrapped by physical boundaries. Personally, I value these collaborations and know that my colleagues do too. If we work together, we can ensure that more of the Web is preserved, and often we can act more quickly, particularly when the risk of loss of content is so great.
Corrected attribution above. Continue reading

Continue Reading →

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM