Home » Posts tagged 'Digital deposit' (Page 7)
Tag Archives: Digital deposit
Public Printer’s Letter to President Obama Regarding Open Government
The Public Printer recently released GPO's letter to the President regarding open government (PDF) (Robert C. Tapella, Public Printer, March 9, 2009). Since it specifically mentions FreeGovInfo, we feel the need to comment and contextualize a bit.
On the one hand, it's great that GPO is reaching out publicly to offer infrastructural help with the government transparency initiative. We're happy to assist in any way we can. We hope FDLP libraries will join GPO in such efforts.
On the other hand, FGI has always argued for a geographically dispersed system of local, official digital repositories, so we cannot support GPO’s goal 1 to make FDsys the official repository for Federal Government publications -- unless it includes a network of distributed repositories modeled after the Federal Depository Library Program (FDLP). What we can support is FDSys as the official distribution channel for federal government publications.
It's not a trivial distinction. "Repository" means that GPO assumes sole responsibility for preservation, a role not specified in legislation. "Distribution channel" means GPO continues its solid century and a half record of distributing information to other institutions which will continue their solid century and a half record of preserving government information for future use while making sure it remains freely available over the internet. Since digital deposit is currently #2 on The Sunlight Foundation's Our Open Government List (OOGL) of top ideas for the President's open government initiative, we can only assume that the public -- or at least those that are most interested in government transparency -- agrees that a geographically dispersed system is a key ingredient in government transparency.
We also believe it is important in discussions of transparency to plan for preservation of and long-term access to information. If, in concentrating on short-term access and on information-as-service, we fail to consider long-term access and instantiation of information for long-term preservation, we will inevitably lose information -- and that would be bad for transparency.
Incomplete Access
We commend and support GPO for building APIs into FDSys. It is heartening and encouraging to see that GPO is publicly and officially proclaiming that "access" means more than providing a web site. But APIs and a web site are only two of the three parts of a complete access system. GPO has yet to acknowledge or even mention the third part of access: the provision of unfiltered bulk data access to government information.
A GPO web site can provide a human-friendly interface for the public and APIs can provide a computer-program-friendly way of querying, fetching, and using information. But, even taken together, these two access points provide only the government-approved, government-designed, government-hosted view of government information.
The problem with these government-only views of government information is that they are limited. No single provider (government or non-government) can provide unlimited access points or views or interfaces.
APIs are not magic. Each is a design for access and the product of choices made by the designer. Each has its own constraints built in. For example, an API might be tied to a particular agency or department, which would limit cross-agency utility. Or an API might be generalized to work across agencies or departments and thus lose rich access to agency-specific information content or structure.
One way to overcome these limitations is for the government to provide bulk data access. This means allowing the public to download raw content in bulk. Where web sites provide one "page" at a time and APIs can provide one or many "facts" at a time, bulk data access provides the raw information so that users can build their own collections, interfaces, and APIs.
This could improve access in ways that GPO could never hope to do all by itself. Imagine, for example, an agricultural library building a digital collection that contains agricultural reports, data, and audio visual content from the The Department of Agriculture, the EPA, the SBA, and NOAA combined with reports, maps, and GIS data from state and local government agencies and other content from its own institutional repository or university press. Then imagine that specialized digital collection having its own state-specific, agriculture-specific API and web site and bulk data access. Then imagine that these repositories are part of the rapidly expanding cloud and you get a sense of a rich govt information ecology.
Such scenarios are possible, but only if GPO and other government agencies make raw content easily, freely available in bulk for use and re-use and re-purposing. Providing only government web sites and government APIs without bulk data downloads and the ability for others to build collections for specific or general purposes will provide only a tiny fraction of open usability and transparency that we could have. There is nothing standing in the way of this happening today except the will of government agencies to make it happen.
Incomplete Preservation
The Public Printer's letter glosses over the problems of long-term access and preservation.
Let's be as clear as we can: we cannot and should not rely solely on GPO for long-term preservation and free access. The shift to digital does not change the methodology for long-term preservation and access. On the contrary, the tenuousness of digital information means that a distributed methodology is even more vital.
We cannot rely solely on GPO because the GPO Electronic Information Access Enhancement Act of 1993 does not even mention permanent access, nor does it guarantee that access will always be free. Indeed, the law specifically allows GPO to charge for access and even for use of its "directory" of information. The law also covers only "appropriate publications distributed by the Superintendent of Documents" -- effectively excluding huge bodies of born-digital information from the scope of what is GPO is allowed to handle. Regardless of GPO's intentions, there is no existing legislative mandate for GPO to provide free, permanent, public access to government information and we therefore cannot rely on it alone to do so.
We should not rely solely on GPO because no single digital archive or repository can ever be as secure and safe as multiple archives, libraries, and repositories. Even if GPO had a legislative mandate to provide permanent preservation and access (which it does not), and even if anyone could guarantee that GPO would always get adequate funding so that it never had to withdraw anything or charge for access for anything (which no one can), it would still be impossible to guarantee that GPO would never lose any information. The nature of digital information is that it can easily be corrupted, altered, lost, or destroyed. It can become unreadable or unusable without constant attention. Relying on any single entity is simply not as safe as relying on multiple organizations. It is more than a truism that Lots of Copies Keep Stuff Safe -- safer than backups and "mirror sites." But this is about more than redundant copies. It is also about relying on different organizations because they have different funding sources, different constituencies, different technologies, and different collections. No single digital collection can ever be as safe as multiple, reliable digital collections.
The good news
The good news is that there are existing organizations that can start working on this right away. There is nothing standing in the way of GPO and the existing FDLP libraries from implementing a digital depository system in which GPO enables FDLP libraries to download bulk data and build local digital collections.
There are existing technologies to facilitate this. The U.S. Government Documents Private LOCKSS Network is preserving "harvested" government information. [w:Peer-to-peer] (P2P) networks (like Napster and BitTorrent) have become increasingly popular because more and more people and some businesses have begun to realize that "distributed files" equals faster access and better preservation. (A geographically dispersed system of local, official digital repositories would be, for all intents and purposes, a P2P network.) Open source software for building digital repositories is widely available and increasingly easy to use.
Summary
APIs are good. They are a necessary part of adequate government information access. But digital distribution is also essential because only digital distribution will enable FDLP libraries and others to build new APIs, to de-ghettoize government information by better integrating it with non-government information, and to ensure long-term, free, public access and usability of government information.
Continue readingThe Intersection of Education, Technology, and Open Content
In a couple of recent posts, Lev Gonick, who is the CIO at Case Western Reserve University, has noted that we have "an educational economy that makes information abundant confronting an educational delivery system built for a time in which information was scarce."
- How Technology Will Reshape Academe After the Economic Crisis by Lev Gonick, Chronicle of Higher Education blog, "The Wired Campus" (February 24, 2009).
- A Small Proposal at the Intersection of Education, Technology, and Open Content, by Lev Gonick, Chronicle of Higher Education blog, "The Wired Campus" (February 26, 2009).
His description of the educational economy and the educational delivery system struck me as analogous to the situation we face with government information. We live in an environment where government information is abundant and gains value by being distributed and reusable. In this environment it is incredibly inexpensive to distribute information, yet governments too often treat it as if it were scarce and expensive to deliver.
It is ironic, for example, that GPO refuses to deposit ninety percent or more of government information in FDLP libraries because it is digital (SOD 301, Superintendent Of Documents Policy Statement, "Dissemination/Distribution Policy for the Federal Depository Library Program" Effective Date: June 1, 2006) and then wonders why libraries find it hard to justify being a depository library.
Imagine a system closer to what Gonick describes. Imagine a system that recognizes that digital information is different from paper and ink information: both more valuable (because it is more easily used and re-used) and less expensive to distribute. Imagine an approach that is a more modern, more appropriate response to digital information than what we have now. Go further and imagine what the depository system would look like if it adopted the vision that Carl Malamud proposes.
Carl says all government information should be available in three ways (all for free):
- as bulk data for downloading and repurposing;
- through an API for querying, retrieving, embedding in other web sites;
- as better official web sites aimed at end users.
(Carl outlines these in his interview with Timothy M. O'Brien February 24, 2009, and in his Rebooting the Federal Register document):
Lev Gonick expands on what truly open information could mean to communities. He contrasts "the largely proprietary learning economy that exists now" with the new environment of "more and more open educational resources." He sees these open resources as creating new opportunities that were not available when we could only rely on proprietary, closed, scarce information resources.
Goncik's specific ideas actually sound a lot like the kind of collaborative, civic-centered services that John Shuler has long advocated and is describing here. Specifically Goncik describes a "a university-led 'connected cities' project" in which "we could invite different communities within our cities (children, schools, professionals, unions, educators, artists, elected officials, and so forth) to communicate with others in this new connected Web." He continues:
They might share oral histories and multimedia presentations about their communities with one another. Or they might participate in formal educational and research exchanges. Scientists could discuss research on sustainability, for instance, in ways that connect to high-school students seeking to learn about ecology and the economics of recycling. We can and we should leverage our universities’ ability to create powerful networks of technology and learners to create binding partnerships that matter. The oceans that once separated us are now made smaller by the technology that we have helped invent and deploy. Deepening the linkages within and between our communities and across our cities is a 21st challenge worthy of great universities.
But this is not just about technology enabling sharing. It is also about having something to share. In order to do this, of course, we will need to guarantee free access to robust, preservable, re-usable collections of information. We could do that by hoping that GPO will always get the funding to do it for us and that it will do it right and meet all the needs of all communities equally well forever. We could hope that the government (GPO, OMB, Congress, etc.) will never privatize information or withdraw information, or alter information. Or, we could take on the task ourselves as depository libraries in the FDLP by demanding digital deposit. Then we could begin building digital collections for different communities-of-interest, world-wide. Libraries could then not only do interesting things with the information that they manage for their communities, but they could also facilitate others re-using the information.
Continue readingObama’s Inaugural Speech: visualized, video-searchable
President Obama's inaugural speech has generated some interesting examples of how technology can be applied to government information when the information is freely available for use and re-use and not locked into government databases or proprietary formats. It is a small piece of text with a lot of public interest and high visibility and, therefore, ripe for these kinds of demonstrations and experiments. Of course, to make use of the information, we have to actually have a copy of it. Imagine what would happen if all government information was actually distributed in open formats to libraries so that we could build collections that were index-able, search-able, visually browsable, and analyzable in interesting ways. Imagine freeing government information from its .gov silos and integrating it with non-government information in digital collections created for particular virtual communities of interest. Imagine the future of digital collections that are as easily re-usable as this small bit of text. Check out these examples!
- Inaugural Words: 1789 to the Present, New York Times. "A look at the language of presidential inaugural addresses. The most-used words in each address appear in [an] interactive chart..., sized by number of uses. Words highlighted in yellow were used significantly more in this inaugural address than average."
- Visual of the Inaugural Address, ProPublica. [Compare this to the NYT version. Stop words matter!]
- Search Inside Obama’s Inaugural Speech. Delve Networks. "We invite you to experience President Obama’s inaugural speech using our search inside technology. To do this, type what you’re looking for into the player searchbar above. A heatmap will show you where information related to your topic appears in the speech. You can move your mouse over the heatmap to see the matches. Click to jump to that place in the speech."
Affirmative Disclosure of Government Information
John Wonderlich, a Program Director of the Sunlight Foundation and a great friend of libraries, has posted some useful suggestions over at The Sunlight Foundation Blog:
I really like John's concept of "affirmative disclosure." I think we could go even further by explicitly addressing the problems of long-term preservation caused by the shift to e-government. I am starting from the assumption that society needs a reliable way to preserve an accurate, complete historical record. Unfortunately, the systems we have in place today makes it difficult, and in some cases impossible, to guarantee that we will preserve a record that is either complete or accurate. Consider, for example, the recent case where researchers at the University of Illinois discovered that the White House removed original documents from its web site, altered them, and replaced them with backdated modifications that appear to be originals but are not. Also consider the project of the Library of Congress, the California Digital Library, the University of North Texas Libraries, the Internet Archive and the U.S. Government Printing Office to try to capture web pages of the current administration by performing a "comprehensive crawl of the.gov domain." These examples illustrate the problem of preserving the historical record. The first shows how the historical record can easily be lost and altered (intentionally or unintentionally -- it doesn't matter which) by lack of accurate metadata (dates, versioning). The second shows the sad state of current preservation: the best record we will have of the government web will be a single, incomplete snapshot of the end of an eight year administration. (Harvesting is imperfect and incomplete: links can break, embedded content can be lost, databases can prohibit or inhibit crawls of their content, and crawls can only save a snapshot of dynamic sites.) In essence, the government has made a major change in information policy by changing the technology of information dissemination and has done so without really examining the implications of the change or even acknowledging that a policy has changed. What was the policy change? In the old policy, the role of government was to collect and assemble and edit and create information and then instantiate it in publications and distribute those instantiations to the public. At that point the role of preservation was in the hands of libraries (mostly FDLP libraries) and archives. But, in the new policy, the government does not actively distribute, but "posts" information on web sites where it is subject to alteration and removal without ever being instantiated anywhere. It is up to the public, consumer groups, individuals, libraries, and special projects to identify when information is posted or changed and then attempt to preserve that information. While that may succeed sometime, the approach has two fatal flaws. First, it is ad-hoc and therefore will almost certainly be incomplete at best. Second, it puts the responsibility of instantiation in the wrong hands: not those who create the information (the government) but those who "discover" the information. The government essentially is renouncing its responsibility to actively, affirmatively create a preseveable instance of the information it creates. While some agencies (e.g. GPO, EIA) are saying that it is now their role to preserve information, other agencies (e.g., NARA) are actually narrowing their role in long-term preservation (notice that NARA is not participating in the ".gov crawl" and says explicitly that "most web records do not warrant permanent retention"). So, let's explicitly expand the idea of "affirmative disclosure" to include "active deposit." By that I mean that the government should be required to actively inform and distribute to the public notifications (metadata) and documents (data) every time a "document" is created or modified or superseded. "Deposit" could be accomplished with technology (e.g., RSS, APIs, OAI and OAI-PMH, etc.) and should be required to include dates and version information. This is the right way to do this because it recognizes the appropriate roles for the different participants in the life cycle of information: government agencies create information products that are preservable and libraries and others preserve those products outside the .gov domain. Continue readingDigital Deposit: Lack of storage space is no excuse
This past weekend I was at my local Costco and not one, but two brands of 1 Terabyte (1000 GB) drives selling at around $300. I also saw a 500 GB (1/2 T) drive for $130. All of the drives were USB friendly meaning you could take one off the shelf and plug it into a USB port and have all that memory available to you.
What can you store in a Terabyte? According to an FBI article on digital forensics, plenty:
"a terabyte is equivalent to about 250 million pages of text, which would stack 10 miles high if printed on both sides of the page."
Surely that's enough space for even smaller libraries considering telling the Government Printing Office that they would like PDF ("access derivitives") delivered to them based on their profiles.
I admit, space isn't the only issue. But it's the objection I've heard most often and I honestly believe that technology has taken it away.
Continue reading
Latest Comments