Home » Articles posted by James A Jacobs (Page 368)
Author Archives: James A Jacobs
Don Waters on the responsibility of libraries
This, I think, is a must read for government information professionals:
- Managing Digital Assets in Higher Education: An Overview of Strategic Issues "[a work in progress]" by Donald J. Waters, October 28, 2005
Why? Because there are explicit parallels between the role libraries will play in the scholarly communications process (which Waters addresses) and the role libraries will play in access to and use of government information. In this paper, Waters (Program Officer of Scholarly Communication at The Andrew W. Mellon Foundation) presents another of his excellent overviews of the problems and opportunities for digital libraries. The paper was part of the ARL forum, Managing Digital Assets Strategic Issues for Research Libraries.
Waters minces no words and faces the issues honestly and creatively. p>
Here are a few excerpts, viewed from the perspective of government information:
- Waters notes the problem when libraries simply "use content stored on remote systems controlled by publishers."
- He notes that the libraries are responsible for "collecting, preserving, providing access to, and disseminating content" and that this mission differs from those who take on digital information projects for "business reasons" (e.g., the Google-Library project, Yahoo and Microsoft joining the Open Content Alliance, and now GPO with its intent to distribute electronic documents "on a cost recovery basis"). They do not do so for philanthropic purposes and these "for-profit competitors" bring their resources and set their sights "squarely on key parts of the higher education business."
- While a shift to electronic versions of government information can eliminate many library costs (ordering, receiving, processing, shelving, and circulating physical copies), many libraries see "covering the costs of preserving digital assets for the long term is a responsibility for someone else" and that this is a "jump-off-the-cliff shift" in responsibility.
- The advantages of digitization are not ends unto themselves. What makes digital information useful is that "the material becomes 'processable,' or subject to computational processing." In other words, if we have fully-functional copies of government digital information, then the information can be used and re-used, mixed and re-mixed.
- If we don't have copies of the information, libraries and citizens won't be able to do this for ourselves; our only avenue of use will be what we can purchase. Waters foresees that resulting in a bleak scenario in which "(1) libraries will not own the publications that form the scholarly record; (2) libraries will not own the archive of the scholarly record; and (3) publishers will charge whatever the market can bear for data-mining services because they control all the underlying resources." For "publishers" read "GPO" and "government agencies" and "private sector re-packagers."
- Waters sees that we need more than one useful database (e.g., FDsys) or one kind of indexing (e.g., Google). "The sheer volume of digitized material ... [will] require implementation of much more sophisticated indexing, searching, and filtering techniques, including broad application of computational linguistic and related statistical techniques as well as sophisticated techniques for filtering based on markup and thesauri, which would relate results to disciplinebased concepts and concerns. Above all, there will be growing demand for mechanisms to link search results flexibly across systems in ways that resemble but will be fundamentally different from metasearching across catalogs." And: "Solutions that the large search engines cannot supply will have to come from search applications developed within and for the academy, and finding these solutions should be a high priority for the academy, its libraries and publishers, to address."
A personal note: Several years ago, I was talking with a colleague in the University of California about the need for digital deposit of government information. She did not see the need for this and to support her point she said that we don't get copies of electronic journals so why should we get copies of government information? I tried to explain that we should get copies of electronic journals and that if we couldn't solve the same issue for non-copyrighted, legally-deposited government information how could we hope to solve the same problem for copyrighted journal articles? Now, we have come full circle and ARL and the Mellon foundation and others are saying that libraries must take responsibility and preserve scholarly journals. The Don Waters article complements the recent report that makes that case: Urgent Action Needed to Preserve Scholarly Electronic Journals . (See also New Mellon Report on Digital Deposit and ARL Endorses Call for Action to Preserve E-Journals.)
Continue readingGoogle begins to offer full-text scanned government documents
The Associated Press says that google print beta started offering the entire contents of books and government documents that aren't entangled in a copyright today.
- Google Offers Index of Public Domain Works, November 03, 2005.
I did a quick search for "government printing office" and got a result of 246 books "with 294000 pages." Of course, not all of those are government documents -- Google presumably returns any book with the phrase anywhere in the full text, including bibliographies.
Google provides a page that illustrates how to interpret your search results and shows the difference between how Google treats "Library books still in copyright" and "Public domain books" and "Books submitted by a publisher."
Interestingly, I found one books that appears to me to be a government publication: S. 1383, Children's Protection from Violent Programming Act of 1993, (Author(s): Science, and Transportation United States. Congress. Senate. Committee on Commerce, Publisher: For sale by the U.S. G.P.O., Supt. of Docs., Congressional Sales Office, Publication Date: 1994, Pages: 133, ISBN: 0160463270) that Google seems to treat as a "book still in copyright."
Continue readingHow do we guarantee authenticity in the digital age?
A new entry in the Issues section of FreeGovInfo addresses some of the issues of guaranteeing authenticity of digital government information:
As always we invite your comments. Continue readingAuthenticity
Who do you Trust? The Authentication Problem
How do we know when a digital document is "authentic"? While many in the library and academic communities hope that there will be a technological solution, the reality is that technology alone cannot solve the problem of authenticity. A report this week of research at a Chinese university illuminates one reason for this: technical tools are subject to failure, compromise, forgery, and hacking.
- U.S. mulls new digital-signature standard, By Anne Broache, and Declan McCullagh, CNET News.com, November 1, 2005.
The article reports a flaw in an official federal standard that was originally devised by the National Security Agency and is widely used to create and verify digital signatures in e-mail and on the Web. In fact, it is embedded in every modern Web browser and operating system. The CNET article notes that, while the flaw that Chinese scientists discovered in the "Secure Hash Algorithm" is "theoretical," it will eventually make it easier to forge electronic signatures.
But authenticity requires more than secure software. Even if we had a tool that could never be hacked and that would last forever, we would still only have part of a solution: the technical part. The other part of the solution is social: it is the issue of Trust.
Software provides the technical part of the solution
The technology of authentication provides a way to verify that a document is what it purports to be and determine if it has been altered or not. Document-creators can use software to create special files (called "hashes" or "signatures" or "keys") based on the original document. These special files are typically stored with a "trusted third party" -- neither the document creator nor the recipient. Document-users can then use software to check the authenticity of the document in hand against that "hash." The software is able to determine only if the document in hand is identical to the original. Even the smallest change (e.g., the insertion or removal of a blank space) will result in a report that the documents are not identical.
Trust is the social part of the solution
But this technological check does not solve the authentication problem by itself. The check against the hash is only as reliable as the trusted third party. The software just gives us a technical means of shifting who we trust -- instead of trusting the party that delivered the document to us, for example, we trust a third party that tells us that the hash is correct and authentic. If the hash isn't authentic and unchanged, the check against the hash is worthless.
This concept of a trusted third party is, therefore, an essential component of the authentication chain. That should lead us to an important question: who will we choose as our trusted third parties? This is important because the tools only work if we can trust the third party to do its job. In the case of government information essential to our democracy, this trust has to last forever.
Who do you trust?
Ask yourself who in society is the most trusted third party in delivering information? The government? The press? Publishers? Technology companies like Microsoft and Verizon?
What about libraries?
Now ask yourself what we will do if we think that technological-verification is all we need to ensure authentication and we find one day that the tools have failed as described in the CNET article.
A Social Solution built on Trusted Institutions and Legal Deposit
Trust is a social phenomenon, not a technical one. What if, instead of putting all our faith in potential technological "solution" for ensuring authenticity of government documents, we instead relied on the existing infrastructure of depository libraries to ensure authenticity through their collective possession of multiple copies of digital government publications, distributed by GPO at the time of their publication under the legal-mandate of 44 USC?
This solution promises to be a sound, sustainable one because it relies on libraries as the trusted repository of information. Libraries have a long, well-established social role of providing information; people trust libraries because of it. Libraries have a vested interest in ensuring that the information they provide is authentic and people trust them to do so because it is their primary mission -- not a byproduct of publishing or making money or the various missions of government agencies.
The trust people place in libraries in general can be increased in the digital environment by relying, not on one or two libraries, but on many libraries with different funding streams and missions. Any unforeseen compromise in one institution becomes a single error in a large system of information-provision. (See Article outlines bottom-up standards for digital preservation systems.) Even in the paper and ink world, forgeries are possible -- though more difficult than in the digital world -- and one important way we determine authenticity is by comparing multiple copies.
A different approach
This approach is subtly different from the approach of hoping for a technological solution to authenticity. It recognizes that the social issue of trust (along with the existence of multiple copies controlled by different parties) is paramount and the role of technology is secondary. The role of technology is simply to provide tools to help implement that trust. Indeed, if we used this social-trust legal-digital-deposit approach, libraries would still use technical tools (e.g., LOCKSS, PKI, state of the art hash technologies) to validate the integrity of digital files. Combine these tools with trusted institutions, legal deposit, and multiple copies under multiple jurisdictions and you have fail-safe a recipe for ensuring authenticity.
Summary
The problem with hoping for a technological solution was clearly articulated back in 2000 by Abby Smith, Director of Programs at the Council on Library and Information Resources.
Interestingly, the scholar-participants suggested that technological solutions to the problem [of establishing the authenticity of a digital object] will probably emerge that would obviate the need for trusted third parties. Such solutions may include, for example, embedding texts, documents, images, and the like with various warrants (e.g., time stamps, encryption, digital signatures, and watermarks). The technologists replied with skepticism, saying that there is no technological solution that does not itself involve the transfer of trust to a third party. Encryption -- for example, public key infrastructure (PKI) -- and digital signatures are simply means of transferring risk to a trusted third party. Those technological solutions are as weak or as strong as the trusted third party. To devise technical solutions to what is, in their view, essentially a social challenge is to engender an "arms race" among hackers and their police.
-- Digital Authenticity in Perspective in "Authenticity in a Digital Environment," Council on Library and Information Resources, Publication 92. (May 2000).
James A. Jacobs, November 3, 2005
Continue readingDo agencies exclude search engines?
Introduction
Recently, a speaker at a conference said that Thomas, the Library of Congress site with extensive information about federal bills and legislation, excluded all search engines from indexing the site. Why, the speaker asked, would a government site wish to block search engines? While government agencies might, like most web sites, exclude search engines from parts of the site for benign reasons (such as blocking temporary pages, out-of-date pages, etc.) It seems counter-intuitive that anyone would want to block access to search engines since they tend to drive most traffic at most web sites.
This prompted me to examine if and when and how government agencies are blocking search engines. I discovered that Thomas does not block all search engines after all -- but it only allows Google to fully index the site. That raises more questions than it answers. What follows is some background and findings and conclusions from a selective examination of what some agencies are doing.
Background
Search engines like Google, Yahoo, MSN search, etc. build indexes of web pages by using software (called "robots" or "spiders") that browse (or "crawl") the web, finding web pages and the links on those pages, following the links they find to find more pages, and so forth. When the spidering software reaches a site, it is supposed to do a couple of things before proceeding. It is supposed to look for a special file named "robots.txt", read the rules in that file that tell it where it may go and where it may not go, and then obey those rules. The rules (called The Robots Exclusion Standard) allow the site to exclude or allow robots to visit particular parts of the site. This is a web convention, not a mandate, but all "well behaved" robots are supposed to follow these rules. The rules can exclude robots from a directory or a particular file or even exclude a particular type of file. They can also specify that one robot can get to one part of the site, while another cannot. (This is useful for site administrators that want to allow their own indexing software to index parts of the site that they don't want to be visible to the world, for example.) There are other ways to ask a search engine to refrain from indexing a page, but I don't cover those here.
Some explanations of how the robots.txt files work are:
I examined the robots.txt files at some dot-gov (.gov) sites. The following is a report on what I have found so far. I used google to identify files (google search for file robots.txt at site .gov) and examined a selection of the federal sites from that result. I invite others to let me know of other interesting exclusions.
Findings
It is not always easy to know the effects of a robots.txt file on a complex site. When the file is quite long and lists many directories that are themselves links to other parts of the site, for instance, it is possible that these are old links and spiders will find the information through another route. The EPA site is an example of this; it has 975 exclusions and, without following every one and navigating the site looking for the information in the exclusions, it is difficult to determine if the information is still findable by spiders.
That said, here are the few sites that I examined that had some interesting features.
- Customs and Border Protection (visited 10/30/2005)
Excludes all spiders except Google and the "gsa-crawler", which it allows everywhere. - Department of Justice (visited 10/30/2005)
Exclude indexing of its archive - Department of Labor (visited 10/30/2005)
Evidently excludes spiders from older press releases (e.g., /dol/media/press/archives.htm), decisions of the Employees' Compensation Appeals Board and the Benefits Review Board. - Environmental Protection Agency (visited 10/30/2005)
Has a lengthy robots.txt file that excludes all robots from parts of the site such as the state Binary Base Maps of the Exposure Assessment Models, a page on the Columbia Space Shuttle Accident, the section on Environmental Monitoring for Public Access and Community Tracking, the entire "Human Exposure Database System," parts of the site that require a password (e.g., "mercury-archive"), and much more. - Thomas (visited 10/30/2005)
Allows Google's robot complete access. Allows the "Ultraseek" robot any plain file, but excludes it from following any dynamic links (which is most of the site). Denies all other robots any access. - USDA Forest Service Southern Research Station (visited 10/30/2005)
Whole sections of publications in pdf format are excluded (e.g., "/pubs/rp/*.pdf" [recent publications]) - Whitehouse (visited 10/30/2005)
An extensive list of over 2000 exclusions. Most of these seem to be simply aimed at excluding robots from the "text" versions of web pages that are probably duplicates of other pages; if so, these have essentially no effect on indexing but are probably used to reduce the number of times spiders hit the site and save bandwidth. Interestingly, though, most of the other exclusions are for directories named "iraq" that do not exist. Excluding non-existing directories is neither necessary nor functional. Perhaps these are in place to save bandwidth by preventing spiders from hunting for information about Iraq rather than following actual links. Examples: /infocus/healthcare/iraq /infocus/internationaltrade/iraq
Conclusions
It appears that most agencies are either not using robots.txt files to excluded search engine robots, or are using the exclusion rules sparingly and appropriately. At least two government web sites (whitehouse.gov and epa.gov), and perhaps others, warrant further study because they have such lengthy robots.txt files or because the exclusions otherwise obscure their overall effect.
Two other federal government web sites (customs.gov and thomas.loc.gov) have troubling exclusions of most search engines except Google.
Continue reading
Latest Comments