Home » Posts tagged 'digital preservation' (Page 24)

Tag Archives: digital preservation

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

New report on the challenges of digital preservation

The Blue Ribbon Task Force on Sustainable Digital Preservation and Access -- created to seek answers to some challenging questions about sustainability of long-term digital preservation -- released its iterim report yesterday -- "Sustaining the Digital Investment: Issues and Challenges of Economically Sustainable Digital Preservation, December 2008." This should add even more to the substance of the debates and consideration about the future directions of the digital depository library. I should point out Chris Greer served on the Task Force and is a member of the Federal Depository Library Council. Very interesting substance to add to our rhetoric about the civic purpose of depository libraries. Continue reading

Continue Reading →

Scarcity or Abundance? : Preserving the Past in a Digital Age

While researching articles for a bibliography on digitial preservation, I ran across this article that seemed worth sharing: Rosenzweig, Roy. Scarcity or abundance? Preserving the past in a digital era. American Historical Review, v. 108, no. 3, p. 735, June 2003. (28 pages) This article provides a good overview of the stakes involved in digital preservation, why web docs disappear, explains why historians can't fully count on the Internet Archive, and worries about the problems historians might face if digital preservation efforts are successful. Article also insists that digital preservation is much too important to be left to the technologists and that archivists and librarians need to be involved. My thanks to the Center for History and New Media for making this article available online. Continue reading

Continue Reading →

Don Waters on the responsibility of libraries

This, I think, is a must read for government information professionals:

Why? Because there are explicit parallels between the role libraries will play in the scholarly communications process (which Waters addresses) and the role libraries will play in access to and use of government information. In this paper, Waters (Program Officer of Scholarly Communication at The Andrew W. Mellon Foundation) presents another of his excellent overviews of the problems and opportunities for digital libraries. The paper was part of the ARL forum, Managing Digital Assets Strategic Issues for Research Libraries.

Waters minces no words and faces the issues honestly and creatively.

Here are a few excerpts, viewed from the perspective of government information:

  • Waters notes the problem when libraries simply "use content stored on remote systems controlled by publishers."
  • He notes that the libraries are responsible for "collecting, preserving, providing access to, and disseminating content" and that this mission differs from those who take on digital information projects for "business reasons" (e.g., the Google-Library project, Yahoo and Microsoft joining the Open Content Alliance, and now GPO with its intent to distribute electronic documents "on a cost recovery basis"). They do not do so for philanthropic purposes and these "for-profit competitors" bring their resources and set their sights "squarely on key parts of the higher education business."
  • While a shift to electronic versions of government information can eliminate many library costs (ordering, receiving, processing, shelving, and circulating physical copies), many libraries see "covering the costs of preserving digital assets for the long term is a responsibility for someone else" and that this is a "jump-off-the-cliff shift" in responsibility.
  • The advantages of digitization are not ends unto themselves. What makes digital information useful is that "the material becomes 'processable,' or subject to computational processing." In other words, if we have fully-functional copies of government digital information, then the information can be used and re-used, mixed and re-mixed.
  • If we don't have copies of the information, libraries and citizens won't be able to do this for ourselves; our only avenue of use will be what we can purchase. Waters foresees that resulting in a bleak scenario in which "(1) libraries will not own the publications that form the scholarly record; (2) libraries will not own the archive of the scholarly record; and (3) publishers will charge whatever the market can bear for data-mining services because they control all the underlying resources." For "publishers" read "GPO" and "government agencies" and "private sector re-packagers."
  • Waters sees that we need more than one useful database (e.g., FDsys) or one kind of indexing (e.g., Google). "The sheer volume of digitized material ... [will] require implementation of much more sophisticated indexing, searching, and filtering techniques, including broad application of computational linguistic and related statistical techniques as well as sophisticated techniques for filtering based on markup and thesauri, which would relate results to disciplinebased concepts and concerns. Above all, there will be growing demand for mechanisms to link search results flexibly across systems in ways that resemble but will be fundamentally different from metasearching across catalogs." And: "Solutions that the large search engines cannot supply will have to come from search applications developed within and for the academy, and finding these solutions should be a high priority for the academy, its libraries and publishers, to address."

A personal note: Several years ago, I was talking with a colleague in the University of California about the need for digital deposit of government information. She did not see the need for this and to support her point she said that we don't get copies of electronic journals so why should we get copies of government information? I tried to explain that we should get copies of electronic journals and that if we couldn't solve the same issue for non-copyrighted, legally-deposited government information how could we hope to solve the same problem for copyrighted journal articles? Now, we have come full circle and ARL and the Mellon foundation and others are saying that libraries must take responsibility and preserve scholarly journals. The Don Waters article complements the recent report that makes that case: Urgent Action Needed to Preserve Scholarly Electronic Journals . (See also New Mellon Report on Digital Deposit and ARL Endorses Call for Action to Preserve E-Journals.)

Continue reading

Continue Reading →

Digital library technologies

Here at Free Government Information, we're extremely interested in the collection, dissemination and preservation of digital government information (for more background we point you to our manifesto of sorts) and feel that libraries have a vital role to play in this area. Cornell University has been at the forefront of digital preservation and has created a Digital Preservation Tutorial that gives a broad overview of the issues and challenges involved in digital preservation in general. With that in mind, we are creating a list of technologies that will be of interest and importance as we move toward a digital FDLP. We're focusing on digital technologies that get at the basic needs for digital preservation: collection or capturing of digital information, description (metadata creation), dissemination (as opposed to simple availability!), and long-term preservation. This list is a work in progress. If you know of other technologies, software, hardware, clients etc that you have used and would recommend to the community, please contact admin at freegovinfo dot info or leave a comment on this page so we can add to the growing list of useful resources.

  • Archive-IT: Web archiving service from the Internet Archive. The service allows institutions to build, manage and search their own web archive through a user friendly web application, without requiring any technical expertise or hosting facilities. Check out a list of their collections. (added 2/7/07)
  • Capturing Electronic Publications (CEP): A web site archiving system developed with Open Source software for Unix/Linux. CEP makes it possible for organizations to periodically download and retain archival copies of their evolving web site(s). CEP uses a web spider, wget, to traverse and download a target website's pages and CVS to archive the pages and their subsequent changes. CEP uses a variety of software packages to create, maintain historical data and provides summary statistics about the website's content. The packages used to create CEP include: Fedora, Apache, CVS, Perl, GD graphic tools, TreeTagger and Wget.
  • CONTENTdm OCLC Digital Collection Management Software. "CONTENTdm® makes everything in your digital collections available to everyone, everywhere. No matter the format — local history archives, newspapers, books, maps, slide libraries or audio/video — CONTENTdm can handle the storage, management and delivery of your collections to users across the Web."
  • cURL. A command line tool for getting or sending files using URL syntax. Curl is targeted at single-shot file transfers. Curl is not a web site mirroring program. Curl is not a wget clone.
  • Del.icio.us: del.icio.us is a social bookmarking site that allows users to bookmark and share Web sites. It also allows for collaborative collection projects like FGI's IAdeposit project where digital govt documents that are tagged "IAdeposit" in delicious are then uploaded to and preserved in the Internet Archive's US govt documents collection. So even if a library can't afford to build its own digital architecture, it can still participate in a digital collection project.
  • DSpace. Open source software that enables open sharing of content that spans organizations, continents and time. "DSpace is the software of choice for academic, non-profit, and commercial organizations building open digital repositories. It is free and easy to install "out of the box" and completely customizable to fit the needs of any organization. DSpace preserves and enables easy and open access to all types of digital content including text, images, moving images, mpegs and data sets." See also: duraspace.org.
  • EPrints. "EPrints is the most flexible platform for building high quality, high value repositories, recognised as the easiest and fastest way to set up repositories of research literature, scientific data, student theses, project reports, multimedia artefacts, teaching materials, scholarly collections, digitised records, exhibitions and performances."
  • Fedora Commons Repository software "The Fedora Repository software has been installed by institutions, worldwide, to support a variety of digital content needs. The Fedora Repository is extremely flexible and can be used to support any type of digital content. There are numerous examples of Fedora being used for digital collections, e-research, digital libraries, archives, digital preservation, institutional repositories, open access publishing, document management, digital asset management, and more." See also: duraspace.org.
  • Greenstone: Greenstone is a suite of software for building and distributing digital library collections. It provides a new way of organizing information and publishing it on the Internet or on CD-ROM in the form of a fully-searchable, metadata-driven digital library..... The aim of the Greenstone software is to empower users, particularly in universities, libraries, and other public service institutions, to build their own digital libraries.
  • HTTrack. HTTrack allows you to download a World Wide Web site from the Internet to a local directory, building recursively all directories, getting HTML, images, and other files from the server to your computer. HTTrack arranges the original site's relative link-structure. Simply open a page of the "mirrored" website in your browser, and you can browse the site from link to link, as if you were viewing it online. HTTrack can also update an existing mirrored site, and resume interrupted downloads.
  • [w:institutional repository] (IR): Many libraries (especially academic libraries) are building digital institutional repositories to collect and manage the content created by their members/professors/researchers. These software infrastructures offer soup-to-nuts tools for collecting, describing, accessing and preserving digital content. Examples include Dspace, Fedora, and EPrints.
  • Lots of Copies Keep Stuff Safe (LOCKSS): Open source software to collect, store, preserve, and provide access to local copies electronic documents. Using Peer-to-peer (P2P) architecture, libraries can compare, share and repair digital content. A Low cost digital library tool! Contact James Jacobs (jrjacobs AT stanford DOT edu) if you're interested in joining the USdocs private LOCKSS network.
  • Lucene. Apache Lucene is a high-performance, full-featured text search engine library written entirely in Java. Lucene is the guts of a search engine - the hard stuff. You write the easy stuff, the UI and the process of selecting and parsing your data files to pump them into the search engine, yourself.
  • Open Archives Initiative The Open Archives Initiative develops and promotes interoperability standards that aim to facilitate the efficient dissemination of content. OAI has its roots in the open access and institutional repository movements. Projects include the Protocol for Metadata Harvesting (OAI-PMH) and the Object Reuse and Exchange (OAI-ORE) standard.
  • Solr: open source enterprise search server based on the Lucene Java search library, with XML/HTTP and JSON APIs, hit highlighting, faceted search, caching, replication, a web administration interface and many more features.
  • SWISH-E: Simple Web Indexing System for Humans.
  • Teleport Pro: Web-based spidering tool. This one's NOT open-source and NOT free ($40), but has been recommended by an FGI volunteer as easy-to-use
  • Wget: An open source UNIX command line spidering tool used to retrieve files automatically off of web servers.
While not a specific technology package, people working in government-funded institutions should be aware of the Digital Preservation Network. According to their web site, the Network is "dedicated to forging a community of practitioners who are focused on the issues of preserving the digital records and publications of government. This online forum will be a repository for the exchange and discussion of ideas, research, strategy and documents that can be used by other practitioners in their organization. Membership into this community will be open to any practitioner employed in a government funded institution that is currently researching or participating in appraisal, acquisition, preservation or access of government records or publications." Most features require registration, but the front page features a good training calendar. Another initiative of note is the Digital Library Federation (DLF). DLF seeks to:
  • define, clarify, and develop prototypes for digital library systems and system components;
  • scan the larger technical environment for and encourage the development of potentially important trends and practices;
  • encourage technology transfer and information sharing between and among DLF members, and between DLF and appropriate commercial sectors; and
  • communicate technical directions and accomplishments of the DLF to a wider audience.
Continue reading

Continue Reading →

Digital Preservation Costs – Migration vs. Emulation

The current issue of RLG DigiNews discusses two possible approaches to preserving digital objects: migration and emulation in terms of costs over the life of the objects.

Migration is moving electronic files from one application to another - say from WordPerfect to the latest edition of word. Emulation is designing software and sometimes hardware that will mimic a desired operating system - say a Macintosh with SoftWindows. Both approaches have merit, and both have problems. Some of the problems arise from proprietary formats protected by trade secrets. If you can't describe the specifications of a WordPerfect file exactly, you'll have trouble rendering it in Word or future word processors (the pink-strikethrough effect).

Another thing that the two approaches have is ongoing costs, discussed in the RLG Diginews article. You can put a box of microfilm in storage for 40 years and have it come out readable with no further investment. But if you did the same thing with WordStar disks even 15 years ago, you are out of luck today.

Any preservation initiative for government information must take the CHRONIC lack of funding resources of federal information agencies. That's why we at FGI believe in a distributed electronic future and not the mostly centralized model currently favored by the government. It's also why many in the documents community, myself included believe in a multiformat model (i.e. print, electronic, microfilm) as well. Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM