Home » Posts tagged 'Web harvesting' (Page 4)
Tag Archives: Web harvesting
Can we rely on trying to ‘harvest’ the web?
Dr. David S.H. Rosenthal, who is Chief Scientist at LOCKSS, and Kris Carpenter Negulescu of the Internet Archive recently organized a workshop on the problems of harvesting and preserving the Web as it evolves from a collection of linked HTML documents to a programming environment whose primary language is Javascript. David and Kris, with help from staff at the Internet Archive, put together a list of 13 problem areas already causing problems for Web preservation:
Database driven features Complex/variable URI formats Dynamically generated URIs Rich, streamed media Incremental display mechanisms Form-filling Multi-sourced, embedded content Dynamic login, user-sensitive embeds User agent adaptation Exclusions (robots.txt, user-agent, ...) Exclusion by design Server-side scripts, RPCs HTML5Read more about this on David's blog:
- Harvesting and Preserving the Future Web, by David Rosenthal, DSHR's Blog (May 7, 2012).
NBII goes dark. Libraries do what they do: harvest and preserve it for future access #opendata
Many of us in the government documents world woke up to 2012 with the following message posted on the Web site of the [w:National Biological Information Infrastructure] (NBII) and distributed around to various library listservs:
In the 2012 President's Budget Request, the National Biological Information Infrastructure (NBII) is terminated. As a result, all resources, databases, tools, and applications within this web site will be removed on January 15, 2012.NBII has been a critical program since 1994 (See Bill Clinton's Executive Order 12906 which created the "National Spatial Data Infrastructure" ("NSDI")). NBII was set up to coordinate a broad array of information at the federal level about biodiversity and ecosystems. Todd Carpenter, director of National Information Standards Organization NISO, put it nicely and succinctly when he tweeted:
What is particularly sad about NBII shutting down is it's precisely the thing we need MORE of not less=>trusted data repositories #opendataWell have no fear, the Library of Congress, Internet Archive and Stanford Libraries have all harvested (separately) the NBII Website -- Stanford harvested twice between January 5 and January 13, 2012for its Fugitive US Agencies collection. Continue reading
Archiving .Gov: Your Help Requested!
As the inauguration ceremony begins tomorrow, we can be assured that the Library of Congress and other partners in the End of Term Harvest project have captured much of the Bush administration's online presence. Many of these websites will be re-captured at later dates, providing an interesting look at how these websites will change over time, through different administrations. On a related note, there will undoubtedly be changes in the coming days, weeks, months, that will eliminate some government agencies. We are trying to archive as many of these "dead" websites as possible in the CyberCemetery, to preserve them in their final form. Please, if you know of a website that is disappearing, email or call me. I'm keeping my eyes and ears open, but there is a lot of content out there, and I welcome your help. After all, this information is for all of us! Thanks, and I wish you all joy as we witness history tomorrow. Continue reading
Harvesting .gov
Harvest time, By William Jackson, GCN, 10/27/08. A nice article about the end-of-administration web harvest. See also: Library Partnership Saves Government Sites. Continue reading
Can we rely on trying to ‘harvest’ the web? part 2
June 3, 2012 / Leave a comment
Recently, we posted here a link to David Rosenthal's list of problems of we have with harvesting and preserving the Web. Here is more on the same topic.
- IIPC Future of the Web Workshop - Introduction & Overview (May 17, 2012)
It is a 22 page PDF that presents in some detail an overview of challenges to capturing web content. It was presented at The Future Web workshop, which was held in May as part of the 2012 International Internet Preservation Consortium General Assembly meeting (IIPC GA) hosted by the Library of Congress. The purpose of the paper was to provide a shared context for participants. The problems:- Database driven features and functions
- Complex/variable URI formats and inconsistent/variable link implementations
- Dynamically generated, ever changing, URIs
- Rich Media
- Scripted, incremental display and page loading mechanisms
- Scripted, HTML forms
- Multi-sourced, embedded material
- Dynamic login/auth services: captchas, cross-site/social authentication, & user- sensitive embeds
- Alternate display based on user agent or other parameters
- Exclusions by convention
- Exclusions by design
- Server side scripts & remote procedure calls
- HTML5 "web sockets"
- Mobile publishing
The paper also lists "Current Mitigation Strategies" but, as Rosenthal pointed out, all of these are aimed at capturing a "user experience" -- and our ability to meet even that goal is limited: A different question libraries should be asking is, How can libraries capture the content behind the user experience? The presentation is important, but, even more important is the raw data that sites use to provide those experiences. This kind of information used to be instantiated in books and magazines and maps and pamphlets and newspapers. Today that "raw data" is stored in databases, XML files, GIS applications, and other data stores. Web harvesting can do little more than capture a snapshot of how that information was presented at a given time in the past by a particular information provider. Libraries should be capturing those raw data sources. By doing that, libraries will ensure that current and future users of libraries will be able to actually use, analyze, and mine the data in new and interesting ways. Seeing how a user in the past might have seen a web page at a particular point in time will be of interest to some cultural historians and is therefore certainly important. But it is only a very small part of what future users will expect from their libraries. As the report says, in passing, "the classical model of web archiving is no longer sufficient for capturing preserving, and re-rendering all the bytes of interest we care about." There's a quick overview of the workshop and lots more links here:- Harvesting and Preserving the Future Web: Content Capture Challenges, by Nicholas Taylor, The Signal (June 1st, 2012).
Continue reading →Continue Reading →