Home » Posts tagged 'End of term archive' (Page 4)
Tag Archives: End of term archive
2016 End of Term (EOT) crawl and how you can help
[Editor’s note: Updated 12/15/16 to include updated email address for End-of-Term project queries (eot-info AT archive DOT org), and information about robots.txt (#1 below) and databases and their underlying data (#5 below). Also updated 12/22/16 with note about duplication of efforts and how to dive deeply into an agency’s domain at the bottom of #1 […]
Nominations sought for the U.S. Federal Government Domain End of Term Web Archive
Some of our readers may have seen this announcement already. But in case not, we need your help to preserve the .gov domain. See the announcement below to find out how. How would YOU like to help preserve the United States federal government .web domain for future generations? But, that’s too huge of a swath […]
Election 2012 Web Archive
This is the final post in our series on the collaborative End-of-term and Election web archives. This post focuses on the Election 2012 Web archive, particularly on the challenges we are facing in building an engaged community of web site nominators. The idea for the 2012 Election Web Archive grew out of conversations with the End of Term project partners (Internet Archive, California Digital Library, Library of Congress, the University of North Texas Libraries and the U.S. Government Printing Office) about Harvard’s potential participation in the project. Librarians at the Harvard Kennedy School (HKS - Harvard’s graduate school of government), suggested that a logical focus for them would be the upcoming presidential election. In the past, presidential elections generated a lot of enthusiasm among faculty and students, who frequently requested both printed and online election-related resources from the library. They hoped to harness some of this enthusiasm. Since the Library of Congress has been collecting in this area for many years, the group decided to collaboratively collect web sites about the 2012 election as a sister collection to the 2012-2013 End-of-term collection. The partners decided to distribute selection responsibilities so that they could each focus on areas of particular interest at their institution. Curators from the Library of Congress would focus on official campaign sites produced by presidential, congressional and gubernatorial candidates, using their in-house tools for tracking nominations. The Harvard curators would focus on web sites produced by non-profit organizations, academic institutions, fact-checking organizations and some individuals, including blogs, tweets and YouTube videos, using the nomination tool created by UNT for the 2008-2009 project. Examples of these web sites include http://factcheck.org, http://dailykos.com and http://campaignmoney.com/. Initial selections were made by library staff and plans were made to engage faculty, students and staff in relevant academic areas such as Democracy, Politics and Institutions and from HKS Research Centers such as the Institute of Politics; the Shorenstein Center on the Press, Politics and Public Policy; the Center for Publish Leadership and the Ash Center for Democratic Governance and Innovation. Nomination of web sites began in December 2011 and will continue on an ongoing basis until the election. The Internet Archive began crawling these sites in January 2012, and will continue crawling these sites on a weekly basis until sometime after the 2012 election. After each crawl, detailed reports are distributed to the entire group, highlighting any problematic web sites, for example sites that couldn’t be collected because of robot exclusion files. The campaign sites nominated by Library of Congress curators will be crawled separately as a part of the Library of Congress Web Archives with the ultimate goal of providing a shared interface for researchers to access all of the campaign and other election-related sites. As soon as crawls were underway, HKS librarians focused on soliciting help in nominating web sites to include in the collection. Previous efforts to engage the community of faculty and students included direct emails, in-person conversations, an advertisement in the student newspaper, and posts to the HKS LinkedIn page, but produced few nominations. The campaign also included contacting staff and librarians at other public policy schools. So far this hasn’t resulted in additional nominations. At the start of the semester in September, HKS librarians will publicize the project on the Journalist’s Resource web site, and possibly on The Monkey Cage web site. In addition, the group made the decision to publish the URL to the nomination tool (which does not require an account to use) directly in public articles about the project. As was learned in the 2008-2009 project, it was more important overall to make the tool as accessible as possible than to lock it down because of the risk of misuse. In the event that any inappropriate URLs are submitted, they can be removed from the list of sites to crawl. Although to date it has been challenging to engage a broad group of nominators for this project, we remain optimistic that as we move closer to the election it will be easier to spark interest and participation in the project. How can you help? If you would like to nominate a web site for the Election 2012 Web Archive, visit the nomination tool and start entering URLs. If you have suggestions for us or any questions, please contact us at eotproject@loc.gov, here on this blog, or on Twitter @eotarchive. Andrea Goethals and Wendy Marcus Gogel Harvard Library Keely Wilczek Harvard Kennedy School Library Continue reading
US Executive Branch Closure Crawl
The State of the Federal Web Report issued in late 2011 noted that Federal agencies planned to eliminate or merge several hundred domains, as part of the President's Campaign to Cut Waste. The goal was to reduce outdated, redundant, and inactive domains. As part of this work, the .gov Task Force overseeing the process asked members of the National Digital Stewardship Alliance (NDSA) to archive and preserve all .gov Executive branch domains slated to be decommissioned or merged. NDSA members immediately agreed that an important step in this process was to preserve the content of these sites as part of our national digital heritage - instead of simply eliminating them. Rather than start a separate, standalone project, we chose to launch a collaborative crawl under the auspices of the End of Term Web Archive project (EOT). Although the EOT project has primarily focused on transitions occurring at the end of administrative terms, part of the goal of the project is to document changes in all online presences of the US Federal government during key periods of transition, regardless of when or under what circumstances they occur. So, a comprehensive harvest, using a targeted list of domains supplied by the .gov Task Force and a general list of all Executive branch domains downloaded from data.gov, began on Saturday, October 8, 2011. The crawl concluded on November 5, 2011 and encompassed 46,278,384 captures and ~13TBs of data compressed. Here's a general outline of the sequence of events of the Fall 2011 crawl:
- Agencies identified recommended actions for domains in their Interim Progress Reports and Web Inventory
- The .gov Task Force collected a list of outgoing .gov domains and shared those with the NDSA
- Internet Archive crawled outgoing sites and the full suite of Executive branch domains (note: for some resources it took several weeks to crawl sites in their entirety)
- GSA eliminated domains after they were archived
Challenges of Site Identification for the 2012 End of Term Web Archive
This post in our series is about the difficulty of selecting webs sites and building a list of "seed URLs." Seed URLs are the starting points that crawlers use to capture the web content you want to capture. Part of the difficulty of building a seed list for the End of Term capture is that the federal web space is large. How large? In June 2011, the Office of Management and Budget made federal websites a target for improving transparency in providing government information, particularly reducing "duplicative" websites that create confusion. OMB’s Jeffrey Zients wrote that, "There are nearly 2,000 top-level Federal .gov domains; within these top-level domains, there are thousands of websites, sub-sites, and micro sites, resulting in an estimated 24,000 websites of varying purpose, design, navigation, usability, and accessibility." A "State of the Web" survey published in December 2011 reported that, "The .gov Web Inventory self-reported 1,489 domains and an estimated 11,013 websites from 56 agencies." This report goes on to describe the terminology used: Domains are registered .gov (or .mil, or even .com as the case may be) names on the Internet (in the form www.agencyname.gov). Most agencies (and some much more than others) use sub-domains that vary from the domain by containing a different root domain (for example, project.agencyname.gov). While domains are registered through the General Services Administration and easily tracked, sub-domains are not. The term "website" is even more nebulous, described as "hosted content … which has a unique homepage and global navigation." As a result, the .gov website numbers are considered a "general estimate." It isn’t just the "bigness" of the federal web space that makes the End of Term effort a challenge – there are also variants in how the different branches of the federal government are managed and tracked. The Library of Congress archives the legislative branch websites through a leg branch crawl run on a monthly basis, so for that effort a list of seed URLs (which may be anything from a domain to a sub-domain to a particular website or part of website) for the leg branch is assiduously maintained – in other words, the situation for that branch of the federal government is in good shape. (It doesn’t hurt that it is a relatively small branch of government.) There is no such regular effort organized for the judicial branch sites and they aren’t under GSA or OPM, so a reliable seed list for the judicial branch is not so easy to come by and why judicial branch seed URL nominations are a priority for the EOT project. The Executive branch runs into problems because the OMB lists do not include most .mil, .org, .com, or other top level domain types sometimes used by federal agencies. The executive branch .gov domains are closely tracked and available at data.gov in a list. However, those sub-domains with different roots added to domain names are not tracked here. Crawlers can get derailed and not realize "xyz.govagency.gov" is part of "www.govagency.gov" and won’t capture it, thus xyz.govagency.gov should have its own seed. It can be particularly important on large sites, such as NASA.gov, to identify these sub-domains as separate seeds. Much more common now are social media or quasi social media .com sites where the federal agency represents itself – the State Department, for example, has a "presence" in Facebook, Flickr, Google+, Tumblr, Twitter, and YouTube. All of these can and should be scoped separately. Complicating things further, federal agencies of all sizes, but particularly smaller bodies, can use third party hosting solutions of varying types. Some House committees use a commercial company to provide their streaming and downloadable video. An example of this is the House Ways & Means Committee use of Granicus, (which is linked to from the House Ways and Means Committee website, along with links to their Facebook, Twitter, and YouTube pages). When I first began doing some research for this blog post, my impression was that the situation is getting easier as the GSA leads an effort for federal "web reform." However as one sees the extent of social media and as third party hosting increases, this optimism is likely misplaced. For now, the End of Term project can use your assistance! Michael Neubert Supervisory Digital Projects Specialist Library of Congress Continue reading
Latest Comments