Home » Posts tagged 'Government websites' (Page 2)
Tag Archives: Government websites
Challenges of Site Identification for the 2012 End of Term Web Archive
This post in our series is about the difficulty of selecting webs sites and building a list of "seed URLs." Seed URLs are the starting points that crawlers use to capture the web content you want to capture. Part of the difficulty of building a seed list for the End of Term capture is that the federal web space is large. How large? In June 2011, the Office of Management and Budget made federal websites a target for improving transparency in providing government information, particularly reducing "duplicative" websites that create confusion. OMB’s Jeffrey Zients wrote that, "There are nearly 2,000 top-level Federal .gov domains; within these top-level domains, there are thousands of websites, sub-sites, and micro sites, resulting in an estimated 24,000 websites of varying purpose, design, navigation, usability, and accessibility." A "State of the Web" survey published in December 2011 reported that, "The .gov Web Inventory self-reported 1,489 domains and an estimated 11,013 websites from 56 agencies." This report goes on to describe the terminology used: Domains are registered .gov (or .mil, or even .com as the case may be) names on the Internet (in the form www.agencyname.gov). Most agencies (and some much more than others) use sub-domains that vary from the domain by containing a different root domain (for example, project.agencyname.gov). While domains are registered through the General Services Administration and easily tracked, sub-domains are not. The term "website" is even more nebulous, described as "hosted content … which has a unique homepage and global navigation." As a result, the .gov website numbers are considered a "general estimate." It isn’t just the "bigness" of the federal web space that makes the End of Term effort a challenge – there are also variants in how the different branches of the federal government are managed and tracked. The Library of Congress archives the legislative branch websites through a leg branch crawl run on a monthly basis, so for that effort a list of seed URLs (which may be anything from a domain to a sub-domain to a particular website or part of website) for the leg branch is assiduously maintained – in other words, the situation for that branch of the federal government is in good shape. (It doesn’t hurt that it is a relatively small branch of government.) There is no such regular effort organized for the judicial branch sites and they aren’t under GSA or OPM, so a reliable seed list for the judicial branch is not so easy to come by and why judicial branch seed URL nominations are a priority for the EOT project. The Executive branch runs into problems because the OMB lists do not include most .mil, .org, .com, or other top level domain types sometimes used by federal agencies. The executive branch .gov domains are closely tracked and available at data.gov in a list. However, those sub-domains with different roots added to domain names are not tracked here. Crawlers can get derailed and not realize "xyz.govagency.gov" is part of "www.govagency.gov" and won’t capture it, thus xyz.govagency.gov should have its own seed. It can be particularly important on large sites, such as NASA.gov, to identify these sub-domains as separate seeds. Much more common now are social media or quasi social media .com sites where the federal agency represents itself – the State Department, for example, has a "presence" in Facebook, Flickr, Google+, Tumblr, Twitter, and YouTube. All of these can and should be scoped separately. Complicating things further, federal agencies of all sizes, but particularly smaller bodies, can use third party hosting solutions of varying types. Some House committees use a commercial company to provide their streaming and downloadable video. An example of this is the House Ways & Means Committee use of Granicus, (which is linked to from the House Ways and Means Committee website, along with links to their Facebook, Twitter, and YouTube pages). When I first began doing some research for this blog post, my impression was that the situation is getting easier as the GSA leads an effort for federal "web reform." However as one sees the extent of social media and as third party hosting increases, this optimism is likely misplaced. For now, the End of Term project can use your assistance! Michael Neubert Supervisory Digital Projects Specialist Library of Congress Continue reading
Building the 2008 End of Term Web Archive
This is the first in a series of guest posts from members of the End of Term (EOT) project. We look forward to blogging here this month and telling you more about our efforts to archive and preserve U.S. government websites. The 2008-2009 End of Term Web Archive contains over 3,300 U.S. Federal Government websites harvested between September 2008 and November 2009 during the transition from the Bush to Obama administrations. The recent DttP article “It Takes a Village to Save the Web” (mentioned in the introduction post yesterday) focuses in detail on the collaborative work that took place at every level, including site nomination, harvesting, data transfer, preservation, analysis and access. This post will briefly highlight some of the more remarkable aspects of building the 2008-2009 archive. The first is that the collaboration came together so quickly and effectively. The group formed in response to the announcement in 2008 that NARA would not be archiving the .gov domain during this transitional period. With no funding and very little notice, the organizations involved were able to set up a means to identify and archive an extremely large body of content. The harvesting process, for example, was carried out across institutions over the course of the year, and while the bulk of the crawls were run by the Internet Archive (IA), the California Digital Library, the University of North Texas and the Library of Congress were able to fill in the gaps at times when the IA could not run crawls, and were also able to target selected areas of content in particular depth. The 16 terabyte body of data that you’re able to browse, search and display at the End of Term Web Archive represents the entirety of what all four organizations harvested, and is just one of three copies of that content. Once the capture phase of the project was complete, each organization transferred its share of content to other project partners to build the complete archive. An August 2011 FGI post Archiving the End of Term Bush Administration points to an article that details the content transfer aspect of this project. The access copy is held at the Internet Archive (with an access gateway provided by CDL), a preservation copy is held at the Library of Congress, and a copy of the data for research and analysis is held at the University of North Texas. Another noteworthy aspect of this project is the number of new, emerging and experimental technologies it either generated or made use of. The Nomination Tool, which will be described in more detail in a future post, was built by the University of North Texas to support selection work for this project, and remains a valuable resource for the web archiving community. The most significant technical work on the End of Term data was conducted by the University of North Texas and the Internet Archive, as they used link graph analysis and other methods to explore the potential for automatic classification of the content in the Classification of the End-of-Term Archive: Extending Collection Development to Web Archives (EOTCD) project. The research identified particular algorithms that hold promise for automatically detecting topically related content across disparate agency sites. This project also evaluated what kinds of metrics might be meaningful as libraries continue to expand their collections to include web harvested material. New reports and findings continue to be posted to the EOTCD project site as of June 2012. Finally, the public access version of the End of Term Web Archive has also drawn from innovative work at the Internet Archive and experimentation with integrating web harvested content with more traditional digital library tools. The Internet Archive has developed a means to extract metadata records from multiple sources of data around this archive. In this case Dublin Core records were generated and loaded into XTF, the digital library discovery system developed at CDL. These records provide the faceted browse access via the archive site list. Not only were these organizations able to act quickly and collaboratively to respond to a significant transition in the government information landscape, but the lessons learned are informing other broad collaborations, advances in web archive collection development, and opportunities to integrate the discovery of content from multiple archives. Tracy Seneca Web Archiving Service Manager California Digital Library Follow us on Twitter: @eotarchive Continue reading
Latest Comments