Home » Posts tagged 'Web harvesting' (Page 3)

Tag Archives: Web harvesting

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

“What the Web said yesterday” and IRC chat on digital collection development

The New Yorker has an interesting piece by Jill Lapore, “What the Web Said Yesterday” which explores the issue of internet preservation and highlights the important work being done by the Internet Archive. Vint Cerf, Chief Internet Evangelist at Google, says we need “digital vellum” or the “twenty-first century will become an informational black hole.” […]

Continue Reading →

Lunchtime Listen: Born Digital Government Information

Back in April I gave a brief (18 minutes!) talk at the Center for Research Libraries Forum, Leviathan: Libraries and Government Information in the Era of Big Data. Here are the slides and the audio recording of that presentation: Government Records and Information: Real Risks and Potential Losses. The presentation gave me an opportunity to […]

Continue Reading →

Examples of challenges of web-harvesting for digital preservation

Have you ever wondered why preserving web-published information is a complex task? Here are three examples of what makes “web harvesting” a difficult and inexact method of digital preservation. Keystone XL Pipeline: Final Supplemental Environmental Impact Statement (SEIS) To preserve this page intact, one would have to collect a total of 13 urls: 7 images, […]

Continue Reading →

Internet Archive’s Wayback Machine now includes auto-nominate feature

I’m a long-time supporter of the Internet Archive and user of the IA’s wayback machine. I’ve even got the wayback machine bookmarklet installed on all of my browsers — and highly recommend all of our readers doing so as well! I just went to check the wayback machine for a site that was 404 and […]

Continue Reading →

Its Not Your Grandfather’s Web Any Longer

David Rosenthal gave another fascinating talk about the state of the web and whether or not we can expect to preserve it by harvesting it. This talk was at the 2013 Spring CNI Membership Meeting in San Antonio, TX. David presents an edited text of his talk with links to the sources on his blog:

David and co-presenter Kris Carpenter Negulescu note, among other things, that the days of a document-centered web are long over and that today, what most web pages do "is download and run programs in the current Web's primary language, Javascript. Javascript is a programming language, not a document description language. Your browser is only incidentally a document rendering engine, its primary function is as a virtual machine." This presents problems for those wishing to preserve information. Among these problems:
  • Database driven features & functions
  • Complex/variable URI formats & inconsistent/variable link implementations
  • Dynamically generated, ever changing, URIs
  • Rich Media
  • Scripted, incremental display & page loading mechanisms
  • Scripted, HTML forms
  • Multi-­sourced, embedded material
  • Dynamic login/auth services: captchas, cross-­site/social authentication, & user-­sensitive embeds
  • Alternate display based on user agent or other parameters
  • Exclusions by convention
  • Exclusions by design
  • Server side scripts & remote procedure calls
  • HTML5 "web sockets"
  • Mobile publishing
For more about these problems, see also: IIPC Future of the Web Workshop -- Introduction & Overview, International Internet Preservation Consortium (May 17, 2012). Read David's complete post for a rich discussion of the issues. Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM