Digitization does not magically preserve paper
When we think about the historical paper-and-ink collections that FDLP libraries have built over the last 200 years, we often wish we could make them more accessible through digitization. But we have to be careful when we think this way. One thing I have learned repeatedly as I have worked with digital information over the last twenty five years is that, in the digital world, "access" and "preservation" have to go together. When we neglect either, we lose both. Some recent writings have reinforced this old idea and are worth remembering:
- All Digital Objects are Born Digital Objects, by Trevor Owens, The Signal (May 15th, 2012). There is no large red button that says "digitize" on it, we make decisions about what significant properties we want to record from a physical object and we work to ensure that those properties are recorded in the newly created digital object. When we talk about the scanner "digitizing" it's all too easy to forget the history of the creation of the digital object and we can easily forget that there are a range of individual and institutional authorial intentions that go into deciding what and how to digitize.
- Digitization is Different than Digital Preservation: Help Prevent Digital Orphans!, by Kristin Snawder, The Signal (July 15th, 2011). Many institutions see the immediate value of having materials available electronically. This is valid reasoning. Many researchers no longer want to come and see the materials. They want access from the comfort of their own couch and fuzzy slippers. But, in the hurry to meet user expectations, institutions may scan large quantities of materials without having a solid plan for preserving the digital images into the future.
- Approaching Digitisation Through A Digital Preservation Perspective. by Alenka Kavčič-Čolić. Presented at the SEEDI (South-Eastern European Digitisation Initiative) 2012, Ljubljana, Slovenia. Most libraries still conceive digitisation as a digital reproduction aimed to provide access to library materials only. The master files resulted from digitisation are usually not digitally preserved and the digital collections run the risk of being lost for the future.
For years, preservation simply meant collecting. The sheer act of pulling a collection of manuscripts from a barn, a basement, or a parking garage and placing it intact in a dry building with locks on the door fulfilled the fundamental preservation mandate of the institution. In this regard, preservation and access have been mutually exclusive activities often in constant tension. "While preservation is a primary goal or responsibility, an equally compelling mandate--access and use--sets up a classic conflict that must be arbitrated by the custodians and caretakers of archival records," states a fundamental textbook in the field (Ritzenthaler, Mary Lynn. Preserving Archives and Manuscripts. Chicago: Society of American Archivists, 1993. p. 1). Access mechanisms, such as bibliographic records and archival finding aids, simply provide a notice of availability and are not an integral part of the object. In the digital world, the concept of access is transformed from a convenient byproduct of the preservation process to its central motif. The content, structure, and integrity of the information object assume center stage; the ability of a machine to transport and display this information object becomes an assumed end result of preservation action rather than its primary goal. Preservation in the digital world is not simply the act of preserving access but also includes a description of the "thing" to be preserved. In the context of this report, the object of preservation is a high-quality, high-value, well-protected, and fully integrated version of an original source document. -- Paul Conway Head, Preservation Department Yale University Library. Preservation in the Digital World Council on Library and Information Resources, Pub62 (March 1996).Continue reading
Thoughts on White House Digital Government Strategy
Building off of last week's post on the Obama Administration's new digital government strategy, I came across this analysis over at TechPresident: "White House Rolls Out New Plan for Digital Government".
Among the changes called for in the plan:Open government activists, including Sunlight Foundation's John Wonderlich and Clay Johnson, writer and former director of Sunlight Labs, expressed "meh" for the new plan. While we're excited that the White House is continuing to espouse the importance of open government principles, our concern is that the plan (PDF) does not address digital preservation or authenticity, two critical issues for librarians in guaranteeing long-term FREE access to government information -- and issues we addressed in a 2010 letter to then deputy CTO for Open Government Beth Noveck. It's all well and good to talk about IT reform, shared IT infrastructure and services, APIs etc, but who's going to manage all of this cool digital stuff for the long-term? And where will the funding (or RE-funding) come from to keep Data.gov afloat in order to manage all of the APIs? In an era where GPO's FY2012 request for $6million to fund continuing development of their Federal Digital System (FDsys) is met with $0 funding by the House and only slightly less catastrophic $500,000 by the Senate, talk is all well and good. Digital infrastructure and services, and more importantly the staff to manage them, costs $$ -- arguably much more $$ than distribution and preservation of paper collections in the FDLP. We need a government and politicians who won't short-change open government and transparency. We need them and the public to realize that "online" does NOT equal "free beer" but "free kittens!" Continue reading
- Within six months, the Office of Management and Budget will release new government-wide standards for open data, content, and web application programming interfaces. Agencies will have another six months to make sure they are following those policies. They are also going to be asked to take two customer-facing online services and expose the information it delivers through APIs to "appropriate audiences," meaning some set of developers will be able to build applications around them without necessarily working in close concert with the agency providing the data.
- Agencies will be asked to publish ever more data through APIs and as structured data, which are the building blocks of modern web design and mobile-ready websites. The White House line on this is that it will also encourage outside developers to build new businesses on top of government data.
- The General Services Administration will establish a Digital Services Innovation Center to work with agencies to modernize how they interact with citizens on the web.
- The White House will begin releasing its own source code on GitHub and launch a "presidential innovation fellowship" program to bring developers from the private sector into government for six-to-12-month projects.
- The federal government will work to develop "MyGov," a prototype central hub for citizens to access all the services and information they're looking for from government online.
- Through programs like one intended to encourage small businesses to compete for government business, the White House will work to change IT procurement practices and cut down on the number of high-dollar, low-output contracts. Other procurement-related initiatives include a government-wide vehicle for mobile device and wireless service contracting and government-wide guidance on bring-your-own-device policies.
- Data.gov, the federal repository for government data available online, will transition away from being a hub for data files and towards a central clearing house of government APIs that developers can incorporate into web applications.
Spread the news and sign the petition to save Library and Archives Canada (LAC)
Here's more news from our Canadian colleagues regarding the ongoing erosion of library services and Library and Archives Canada (LAC). The announced cuts to the LAC include:
- Elimination of 30% of archivists and archival assistants;
- Reduction of digitization and circulation staff by 50%;
- Reduction of preservation and conservation staff;
- Closure of the interlibrary loans unit;
- Elimination of the National Archival Development Program (NADP) which supports -programming at provincial, regional and university archives across Canada.
- CAUT Campaign on Library and Archives Canada
- Updates on Facebook
- Twitter : @CCA_Archives
- Save NADP On-Line petition has reached over 4600 signatures - let's keep going - please share this link with archives supporters
- Resources are also being posted on CCA’s website
National Archives Releases John Huston’s Controversial WWII Documentary
Thanks to Gary for posting about this!
- View Online: National Archives Releases Restored Version of 3rd Film in John Huston’s WWII Documentary Trilogy, by Gary Price, InfoDocket (May 29, 2012).
The National Archives and Records Administration's restoration of Let There Be Light (1946), John Huston's controversial World War II documentary about the rehabilitation of psychologically scarred combat veterans can now be downloaded online.
The third in the World War II trilogy commissioned from Academy Award-winning director John Huston by the US Army Signal Corps, Let There Be Light follows the treatment of emotionally traumatized GIs from their admission at a racially integrated psychiatric hospital to their reentry into civilian life.
...The War Department pulled the film shortly before its premiere at the Museum of Modern Art and commissioned a replacement in which white actors took all the speaking roles and the GIs upbringing was blamed for their psychological condition instead of war trauma. Let There Be Light was first shown publicly in December 1980, after a chorus of Hollywood leaders, joined by Vice President Walter Mondale, persuaded the Secretary of the Army, Clifford Alexander, Jr., to authorize its release....
Can we rely on trying to ‘harvest’ the web? part 2
June 3, 2012 / Leave a comment
Recently, we posted here a link to David Rosenthal's list of problems of we have with harvesting and preserving the Web. Here is more on the same topic.
- IIPC Future of the Web Workshop - Introduction & Overview (May 17, 2012)
It is a 22 page PDF that presents in some detail an overview of challenges to capturing web content. It was presented at The Future Web workshop, which was held in May as part of the 2012 International Internet Preservation Consortium General Assembly meeting (IIPC GA) hosted by the Library of Congress. The purpose of the paper was to provide a shared context for participants. The problems:- Database driven features and functions
- Complex/variable URI formats and inconsistent/variable link implementations
- Dynamically generated, ever changing, URIs
- Rich Media
- Scripted, incremental display and page loading mechanisms
- Scripted, HTML forms
- Multi-sourced, embedded material
- Dynamic login/auth services: captchas, cross-site/social authentication, & user- sensitive embeds
- Alternate display based on user agent or other parameters
- Exclusions by convention
- Exclusions by design
- Server side scripts & remote procedure calls
- HTML5 "web sockets"
- Mobile publishing
The paper also lists "Current Mitigation Strategies" but, as Rosenthal pointed out, all of these are aimed at capturing a "user experience" -- and our ability to meet even that goal is limited: A different question libraries should be asking is, How can libraries capture the content behind the user experience? The presentation is important, but, even more important is the raw data that sites use to provide those experiences. This kind of information used to be instantiated in books and magazines and maps and pamphlets and newspapers. Today that "raw data" is stored in databases, XML files, GIS applications, and other data stores. Web harvesting can do little more than capture a snapshot of how that information was presented at a given time in the past by a particular information provider. Libraries should be capturing those raw data sources. By doing that, libraries will ensure that current and future users of libraries will be able to actually use, analyze, and mine the data in new and interesting ways. Seeing how a user in the past might have seen a web page at a particular point in time will be of interest to some cultural historians and is therefore certainly important. But it is only a very small part of what future users will expect from their libraries. As the report says, in passing, "the classical model of web archiving is no longer sufficient for capturing preserving, and re-rendering all the bytes of interest we care about." There's a quick overview of the workshop and lots more links here:- Harvesting and Preserving the Future Web: Content Capture Challenges, by Nicholas Taylor, The Signal (June 1st, 2012).
Continue reading →Continue Reading →