Home » Posts tagged 'digital preservation' (Page 20)

Tag Archives: digital preservation

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

GPO Establishes First Preservation Librarian Position

FOR IMMEDIATE RELEASE: July 14, 2010 No. 10-23 MEDIA CONTACT: GARY SOMERSET 202.512.1957, 202.355.3997 cell gsomerset@gpo.gov GPO ESTABLISHES FIRST PRESERVATION LIBRARIAN POSITION WASHINGTON-The U.S. Government Printing Office (GPO) is continuing its commitment to preserving the documents of our democracy by establishing the agency's first preservation librarian position. GPO's preservation librarian will be tasked with updating the Federal Depository Library Program (FDLP) collection management plan for the preservation of federal government documents. David Walls will serve as GPO's first preservation librarian; he is a member of the American Library Association (ALA) and comes to the agency from Yale University where he worked as a preservation librarian for 12 years. While at Yale, Walls established practices for the digital conversion of library and special collection materials. Digital preservation is an ongoing initiative for GPO. In 2009, the agency launched GPO's Federal Digital System (FDsys), a content management system, preservation repository and advanced search engine that provides the public with permanent public access to federal government information. GPO is also a member of LOCKSS (Lots of Copies Keep Stuff Safe), a worldwide digital preservation alliance that collaborates with libraries and organizations on preservation initiatives. Link to FDsys: www.fdsys.gov "David's experience and expertise in preservation will be an asset to GPO and its mission of Keeping America Informed," said Acting Superintendent of Documents Ric Davis. "This is an important position for the agency as we work with the library community on the continuing transition to a primarily electronic FDLP, and ensure that the content can be migrated in the future to guarantee current and permanent public access to federal government information." The GPO is the federal government's primary centralized resource for gathering, cataloging, producing, providing, authenticating, and preserving published U.S. government information in all its forms. GPO is responsible for the production and distribution of information products and services for all three branches of the federal government. In addition to publication sales, GPO makes government information available at no cost to the public through GPO's Federal Digital System (www.fdsys.gov) and through partnerships with approximately 1,220 libraries nationwide participating in the Federal Depository Library Program. For more information, please visit www.gpo.gov. Follow GPO on Twitter http://twitter.com/USGPO and on YouTube http://www.youtube.com/user/gpoprinter. Continue reading

Continue Reading →

A New Push for the Office of Technology Assessment

Some good news for those who value public policy based on well informed science, "the possibility of reconstituting OTA itself is gaining new momentum."

Steven notes that there is a comprehensive archive of OTA publications from 1972-1995 available on the Federation of American Scientists web site. There is also, of course, the " OTA Legacy" collection at the University of North Texas Libraries, "CyberCemetary." Continue reading

Continue Reading →

How Much Digital Information?

Since 2007, on behalf of EMC Corporation, IDC has been sizing what it calls the Digital Universe, or the amount of digital information created and replicated in a year. The newest report is now available:

These reports estimate the size of everything digital. IDC looks at the installed base of devices or applications that could capture or create digital information and estimates (based on their research and "other sources") how much information was created in a year. They also estimate the number of times a a unit of information is replicated. The include devices such as mobile phones and bar code readers and video games as well as cameras, scanners, email, office applications, databases, GPS, medical imaging, and lots more. A lot of this is estimates and I found it hard to tell how much was gathered evidence and how much was speculation (see their methodology in the first IDC Digital Universe paper, published in 2007). To me this means that the figures they come up with may not be very accurate. Predictions of the future based on these estimates are, I think, very speculative. Nevertheless, I've been following these ever since I noticed that Fran Berman quoted an earlier report and referred to 2007 as the "cross-over" year: the year in which more digital data was created than there was data storage to host it. (Berman, Francine, Got data?: a guide to data preservation in the information age, Commun. ACM, 51 (2008), 50-56.) Even if you don't believe that the IDC numbers are 100% accurate, the general ideas that they promote are probably not that far off the mark. Some of those ideas:
  • Last year, despite the global recession, the Digital Universe set a record growing by 62% to nearly 800,000 petabytes.
  • The average file size is getting smaller. The number of things to be managed is growing twice as fast as the total number of gigabytes.
  • The growth of the Digital Universe is like a perpetual tsunami. How will we find the information we need when we need it?
  • How will we know what information we need to keep, and how will we keep it?
That last item is my favorite. Regardless of exactly how much digital information is created each year, regardless of how much storage space we have, regardless of the fact that a lot of the "digital universe" that IDC describes is throw-away information that no one would think is worth keeping, we are still faced with Lots of Stuff and we need to figure out What to Preserve. That, I believe, is the next big challenge for digital preservation. One way to face that challenge is to rely on producers to decide what to save. If a government agency produces something digital, allow that agency (or GPO, or LoC, or NARA, or OMB or OPM, or your favorite TLA) to decide for you if that information is worth saving. Another way to face the challenge is to rely on a few big organizations. That is: pool our resources and outsource preservation to a few big organizations that will do this for us. Some of the same players pop up here: LoC and NARA, for example, but there are also organizations like Portico, and the Internet Archive, and ICPSR. Both of the above solutions hope that someone else will take into account the needs of all possible users and make the right decisions. That model can work for some classes of information with appropriate governance and decision-making structures in place. But, I believe, the lesson from the IDC report is that the "digital universe" is so large that we should not assume that any single solution will be enough. There is just too much information and there are too many decisions to make about what is worth saving. While information producers and a few big preservation organizations can do a lot, they cannot do everything. And, their size alone will constrain their decisions. It will be harder for big organizations to respond to the needs of smaller communities of interest. What is the alternative? I think that we need (what shall we call them...?) Libraries. Public Libraries, Special Libraries, College and University Libraries, and School Libraries. These can work together or independently. They can address the needs of their particular communities of interest. This will accomplish three things:
  1. It will aid preservation by making the preservation community bigger. This will not only increase redundancy, but will also help ensure that there is less chance that a single system or financial or governance failure will mean a loss of all information.
  2. It will help deal with the scale of the preservation problem (as identified by the IDC report). With more players and more stake holders, there will be more voices and more variety in the decision making process when we collectively decide what to save. This will mean, for example, that a group of School Libraries working together on digital preservation could ensure that an item of essential use to K-12 will be saved even if no university saves it. And vice versa.
  3. It will help users find and use the information they need. Today, it seems that everyone understands what librarians have always known: that there is a lot of information in the world. It used to take a library degree to get an appreciation of all the sources of information in the world. Today, everyone that uses the Web has that same appreciation. It seems like every day there is another newspaper article or blog posting about how great it is to have access to "everything." But the "everything" people see on the Web is really only a subset of everything and it only appears to be "everything" because there is so much in this very large subset of everything. And, when your only option is to search "everything" you quickly discover that that is not always the best way to find just what you want. (Even Google has segmented information into categories like movies, blogs, books, and scholarly information.) Having community-of-interest collections will enable libraries to build user-interfaces that work best for those communities and that provide access to the information those communities most want.
Libraries won't replace "everything" collections. They will complement each other and unfocused "everything" collections. They will enrich us all and help ensure that we will preserve what needs to be preserved as the "digital universe" expands more rapidly than we could otherwise deal with. See also: Forecast of Worldwide Information Growth. How much Information. Citizens in the Dark? Government Information in the Digital Age. Continue reading

Continue Reading →

Open Library redesign and proposal for collaborative digitizing of documents

The Open Library announced yesterday that their redesigned site is now available with lots of new features and functionalities. As I suggested in my tweet a few minutes ago, wouldn't it be great if lots of depository libraries bought cheap book scanners like the Decapod (A Mellon funded project), digitized government documents and uploaded them to the Open Library? There are tons of records for government documents just waiting for the attachment of a digital file. And GPO could help by sharing their records from the Catalog of Government Publications (CGP) with the Open Library where librarians and others could enhance to make more robust metadata (which could be fed back in to the CGP!). Lots of libraries with Decapods make light work! (Full disclosure: I'm on the board of QuestionCopyright, a 501(c)(3) non-profit which has its own book scanning hardware/software project called Book Liberator. BL developers are in close contact with Decapod folks. But I get no economic benefit from either Book Liberator or Decapod.) Continue reading

Continue Reading →

Special issue on Technology and digital preservation

Library Hi Tech, Volume 28, issue 2 (2010) is a special issue on "Technology and digital preservation." Preprints of articles are now available, though subscription may be required to access some. A few that you may find interesting:

  • Economics, sustainability, and the cooperative model in digital preservation, by Mr. Tyler O. Walters, Dr. Katherine Skinner. "The authors provide an examination of the emerging field of digital preservation and its economics. They consider in detail the cooperative model and the path it provides toward sustainability as well as how it fosters participation by cultural memory organizations and their administrators, who are concerned about what digital preservation will ultimately cost and who will pay."
  • "Land of the lost": a discussion of what can be preserved through digital preservation, by Mr. David Pearson, Mr. Nicholas del Pozo, Mr. Andrew Stawowczyk Long. "...proposes the concept of preservation intent: a clear articulation of a commitment to preserve an object, the specific elements of that object that should be preserved, and a clear time line for the duration of preservation. It investigates these concepts through simple and practical examples."
  • Keeping It Simple: The Alabama Digital Preservation Network (ADPNet), by Mr. Aaron Trehub, Mr. Thomas C. Wilson. "The purpose of this paper is to present a brief overview of the current state of Distributed Digital Preservation (DDP) networks in North America and to provide a detailed technical, administrative, and financial description of a working, self-supporting DDP network: the Alabama Digital Preservation Network (ADPNet). The authors view ADPNet in a comparative perspective with other Private LOCKSS Networks (PLNs) and argue that the Alabama model represents a promising approach to DDP for other states and consortia."
There is much more. A very rich and informative issue. Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM