Building the 2008 End of Term Web Archive

This is the first in a series of guest posts from members of the End of Term (EOT) project. We look forward to blogging here this month and telling you more about our efforts to archive and preserve U.S. government websites. The 2008-2009 End of Term Web Archive contains over 3,300 U.S. Federal Government websites harvested between September 2008 and November 2009 during the transition from the Bush to Obama administrations. The recent DttP article “It Takes a Village to Save the Web” (mentioned in the introduction post yesterday) focuses in detail on the collaborative work that took place at every level, including site nomination, harvesting, data transfer, preservation, analysis and access. This post will briefly highlight some of the more remarkable aspects of building the 2008-2009 archive. The first is that the collaboration came together so quickly and effectively. The group formed in response to the announcement in 2008 that NARA would not be archiving the .gov domain during this transitional period. With no funding and very little notice, the organizations involved were able to set up a means to identify and archive an extremely large body of content. The harvesting process, for example, was carried out across institutions over the course of the year, and while the bulk of the crawls were run by the Internet Archive (IA), the California Digital Library, the University of North Texas and the Library of Congress were able to fill in the gaps at times when the IA could not run crawls, and were also able to target selected areas of content in particular depth. The 16 terabyte body of data that you’re able to browse, search and display at the End of Term Web Archive represents the entirety of what all four organizations harvested, and is just one of three copies of that content. Once the capture phase of the project was complete, each organization transferred its share of content to other project partners to build the complete archive. An August 2011 FGI post Archiving the End of Term Bush Administration points to an article that details the content transfer aspect of this project. The access copy is held at the Internet Archive (with an access gateway provided by CDL), a preservation copy is held at the Library of Congress, and a copy of the data for research and analysis is held at the University of North Texas. Another noteworthy aspect of this project is the number of new, emerging and experimental technologies it either generated or made use of. The Nomination Tool, which will be described in more detail in a future post, was built by the University of North Texas to support selection work for this project, and remains a valuable resource for the web archiving community. The most significant technical work on the End of Term data was conducted by the University of North Texas and the Internet Archive, as they used link graph analysis and other methods to explore the potential for automatic classification of the content in the Classification of the End-of-Term Archive: Extending Collection Development to Web Archives (EOTCD) project. The research identified particular algorithms that hold promise for automatically detecting topically related content across disparate agency sites. This project also evaluated what kinds of metrics might be meaningful as libraries continue to expand their collections to include web harvested material. New reports and findings continue to be posted to the EOTCD project site as of June 2012. Finally, the public access version of the End of Term Web Archive has also drawn from innovative work at the Internet Archive and experimentation with integrating web harvested content with more traditional digital library tools. The Internet Archive has developed a means to extract metadata records from multiple sources of data around this archive. In this case Dublin Core records were generated and loaded into XTF, the digital library discovery system developed at CDL. These records provide the faceted browse access via the archive site list. Not only were these organizations able to act quickly and collaboratively to respond to a significant transition in the government information landscape, but the lessons learned are informing other broad collaborations, advances in web archive collection development, and opportunities to integrate the discovery of content from multiple archives. Tracy Seneca Web Archiving Service Manager California Digital Library Follow us on Twitter: @eotarchive Continue reading

Continue Reading →

Welcome End-of-term-archive (@eotarchive) as FGI guest bloggers for July 2012

[UPDATE 7/2/12: I've added all the names of the people blogging as "EOT archive". jrj] It's been a while since we've had a guest blogger, but this month's turn at the podium will surely make up for it. Our guest bloggers for July, 2012 the members of the End of Term (EOT) Web Archiving project -- that's @eotarchive on twitter. Group members contributing blog posts include:

  • Andrea Goethals: Digital Preservation and Repository Services Manager - Harvard Library
  • Abbie Grotke: Web Archiving Team Lead - Library of Congress
  • Cathy Hartman: Associate Dean - University of North Texas Libraries
  • Michael Neubert: Supervisory Digital Projects Specialist – Library of Congress
  • Kris Carpenter Negulescu: Director, Web Group - Internet Archive
  • Tracy Seneca: Web Archiving Service Manager – California Digital Library
The EOT collaboration began in the summer of 2008, when the project partners, all members of the International Internet Preservation Consortium (IIPC) and partners in the National Digital Infrastructure and Preservation Program (NDIIPP), agreed to join forces to collaboratively archive the U.S. Government web at the end of the Bush administration. The goal of the project team was to execute a comprehensive harvest of the Federal Government domains (.gov, .mil, .org, etc.) in the final months of the Bush administration, and to document changes in the federal government websites as agencies transitioned to the Obama administration. The 2008-2009 Archive includes over 16 terabytes of data collected from Federal Government websites in the Legislative, Executive, and Judicial branches of government and is available for public access. Partners for the 2008-2009 capture included the Internet Archive, the California Digital Library, the Library of Congress, the University of North Texas Libraries, and the U.S. Government Printing Office. Harvard Library has joined the partnership for the 2012-2013 work. The partners are again beginning an End of Term capture for 2012-13. Additionally in 2012, a capture of elections-related websites began in January and will run through the November elections. For information about the 2012-2013 End of Term project, see an upcoming post on this blog. For an in-depth discussion of the 2008-2009 Webarchive, see the article “It Takes a Village to Save the Web: The End of Term Web Archive” recently published in DttP: Documents to the People, Spring 2012, Volume 40, no. 1, pages 16-23 (That issue is not yet online, but IS available in many libraries around the country). Welcome End-of-term archive! Continue reading

Continue Reading →

Smithsonian: the Vice Presidents that time forgot

For all you Presidential historians out there, the Smithsonian has a funny/sad/strange article about the history of the vice-presidency -- a job that John Adams, the first vice-president, described as "the most insignificant office that ever the invention of man contrived" and John Nance Garner, the 32nd VP from 1933-1941, said "wasn’t worth a bucket of warm spit." Read on. It may make you want to visit Huntington, Indiana and the Quayle Vice Presidential Learning Center (yes THAT Quayle :-)). Read more:The Vice Presidents That History Forgot: The U.S. vice presidency has been filled by a rogues gallery of mediocrities, criminals and even corpses. Tony Horwitz. Smithsonian magazine, July-August 2012

The Constitution also failed to specify the powers and status of vice presidents who assumed the top office. In fact, the second job was such an afterthought that no provision was made for replacing VPs who died or departed before finishing their terms. As a result, the office has been vacant for almost 38 years in the nation’s history. Until recently, no one much cared. When William R.D. King died in 1853, just 25 days after his swearing-in (last words: “Take the pillow from under my head”), President Pierce gave a speech addressing other matters before concluding “with a brief allusion” to the vice president’s death. Other number-twos were alive but absentee, preferring their own homes or pursuits to an inconsequential role in Washington, where most VPs lived in boardinghouses (they had no official residence until the 1970s). Thomas Jefferson regarded his vice presidency as a “tranquil and unoffending station,” and spent much of it at Monticello. George Dallas (who called his wife “Mrs. Vice”) maintained a lucrative law practice, writing of his official post: “Where is he to go? What has he to do?—no where, nothing.” Daniel Tompkins, a drunken embezzler described as a “degraded sot,” paid so little heed to his duties that Congress docked his salary.
[HT to BoingBoing!] Continue reading

Continue Reading →

State Agency Databases Activity Report 7/1/2012

It has been awhile since the last activity report for the State Agency Databases project at http://wikis.ala.org/godort/index.php?title=State_Agency_Databases, but work continues. Much of the work has been in the necessary but less showy work of link fixing. We aim to check our pages at least four times a year and our most recent check took place last month. In the past two weeks we've had an outright deletion and a few additions: GEORGIA (Chris Sharpe) Securities & Business Regulation Division Database was deleted. It provided information about registered cemeteries, charities, and securities. BIOGRAPHICAL DATABASES - Karen Kitchens added resources on Wyoming Territorial and State governors. Continue reading

Continue Reading →

Sourcebook of Criminal Justice Statistics: another defunded publication

Add the Sourcebook of Criminal Justice Statistics to the growing list of defunded federal publications. The Sourcebook, published since 1973, is a project of the University at Albany, School of Criminal Justice's Hindelang Criminal Justice Research Center and funded by the U.S. Department of Justice, Bureau of Justice Statistics. The Sourcebook offers a wide variety of statistics regarding the characteristics of criminal justice systems, public attitudes toward crime, nature and distribution of offenses, Characteristics and distribution of persons arrested, Judicial processing of defendants, and Persons under correctional supervision. Due to budget cuts, the Department of Justice is terminating funding for the Sourcebook as of 8/1/12. Please take a few moments to answer the survey they're conducting as part of their effort to secure alternative funding sources.

As a result of the substantial budget cut that has affected the Bureau of Justice Statistics, the Sourcebook of Criminal Justice Statistics Online will cease being funded by that agency as of August 1st, 2012. For over 40 years, Sourcebook has served as a standard reference tool in the field of criminal justice. We are proud to have offered reliable data on a variety of crime and justice issues to a wide scope of users. We are actively seeking ways to maintain the services offered by Sourcebook and any development in that direction will be posted here. We thank you for your support throughout the years and encourage you to respond to our brief survey. Your feedback will be kept anonymous and will be helpful in our efforts to procure other funding sources.
Continue reading

Continue Reading →

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM