Authenticity

Who do you Trust? The Authentication Problem

How do we know when a digital document is "authentic"? While many in the library and academic communities hope that there will be a technological solution, the reality is that technology alone cannot solve the problem of authenticity. A report this week of research at a Chinese university illuminates one reason for this: technical tools are subject to failure, compromise, forgery, and hacking.

The article reports a flaw in an official federal standard that was originally devised by the National Security Agency and is widely used to create and verify digital signatures in e-mail and on the Web. In fact, it is embedded in every modern Web browser and operating system. The CNET article notes that, while the flaw that Chinese scientists discovered in the "Secure Hash Algorithm" is "theoretical," it will eventually make it easier to forge electronic signatures.

But authenticity requires more than secure software. Even if we had a tool that could never be hacked and that would last forever, we would still only have part of a solution: the technical part. The other part of the solution is social: it is the issue of Trust.

Software provides the technical part of the solution

The technology of authentication provides a way to verify that a document is what it purports to be and determine if it has been altered or not. Document-creators can use software to create special files (called "hashes" or "signatures" or "keys") based on the original document. These special files are typically stored with a "trusted third party" -- neither the document creator nor the recipient. Document-users can then use software to check the authenticity of the document in hand against that "hash." The software is able to determine only if the document in hand is identical to the original. Even the smallest change (e.g., the insertion or removal of a blank space) will result in a report that the documents are not identical.

Trust is the social part of the solution

But this technological check does not solve the authentication problem by itself. The check against the hash is only as reliable as the trusted third party. The software just gives us a technical means of shifting who we trust -- instead of trusting the party that delivered the document to us, for example, we trust a third party that tells us that the hash is correct and authentic. If the hash isn't authentic and unchanged, the check against the hash is worthless.

This concept of a trusted third party is, therefore, an essential component of the authentication chain. That should lead us to an important question: who will we choose as our trusted third parties? This is important because the tools only work if we can trust the third party to do its job. In the case of government information essential to our democracy, this trust has to last forever.

Who do you trust?

Ask yourself who in society is the most trusted third party in delivering information? The government? The press? Publishers? Technology companies like Microsoft and Verizon?

What about libraries?

Now ask yourself what we will do if we think that technological-verification is all we need to ensure authentication and we find one day that the tools have failed as described in the CNET article.

A Social Solution built on Trusted Institutions and Legal Deposit

Trust is a social phenomenon, not a technical one. What if, instead of putting all our faith in potential technological "solution" for ensuring authenticity of government documents, we instead relied on the existing infrastructure of depository libraries to ensure authenticity through their collective possession of multiple copies of digital government publications, distributed by GPO at the time of their publication under the legal-mandate of 44 USC?

This solution promises to be a sound, sustainable one because it relies on libraries as the trusted repository of information. Libraries have a long, well-established social role of providing information; people trust libraries because of it. Libraries have a vested interest in ensuring that the information they provide is authentic and people trust them to do so because it is their primary mission -- not a byproduct of publishing or making money or the various missions of government agencies.

The trust people place in libraries in general can be increased in the digital environment by relying, not on one or two libraries, but on many libraries with different funding streams and missions. Any unforeseen compromise in one institution becomes a single error in a large system of information-provision. (See Article outlines bottom-up standards for digital preservation systems.) Even in the paper and ink world, forgeries are possible -- though more difficult than in the digital world -- and one important way we determine authenticity is by comparing multiple copies.

A different approach

This approach is subtly different from the approach of hoping for a technological solution to authenticity. It recognizes that the social issue of trust (along with the existence of multiple copies controlled by different parties) is paramount and the role of technology is secondary. The role of technology is simply to provide tools to help implement that trust. Indeed, if we used this social-trust legal-digital-deposit approach, libraries would still use technical tools (e.g., LOCKSS, PKI, state of the art hash technologies) to validate the integrity of digital files. Combine these tools with trusted institutions, legal deposit, and multiple copies under multiple jurisdictions and you have fail-safe a recipe for ensuring authenticity.

Summary

The problem with hoping for a technological solution was clearly articulated back in 2000 by Abby Smith, Director of Programs at the Council on Library and Information Resources.

Interestingly, the scholar-participants suggested that technological solutions to the problem [of establishing the authenticity of a digital object] will probably emerge that would obviate the need for trusted third parties. Such solutions may include, for example, embedding texts, documents, images, and the like with various warrants (e.g., time stamps, encryption, digital signatures, and watermarks). The technologists replied with skepticism, saying that there is no technological solution that does not itself involve the transfer of trust to a third party. Encryption -- for example, public key infrastructure (PKI) -- and digital signatures are simply means of transferring risk to a trusted third party. Those technological solutions are as weak or as strong as the trusted third party. To devise technical solutions to what is, in their view, essentially a social challenge is to engender an "arms race" among hackers and their police.
-- Digital Authenticity in Perspective in "Authenticity in a Digital Environment," Council on Library and Information Resources, Publication 92. (May 2000).

James A. Jacobs, November 3, 2005

Continue reading

Continue Reading →

Nevada Library Assn presentation: Privacy

FGI volunteers Daniel Cornwall, James R. Jacobs and Shinjoung Yeo were invited to give a panel at the Nevada Library Association's 2005 annual conference in Reno, NV, on October 21, 2005. Below is the text of James' part of the presentation about privacy and government information in the digital age. We'd love to hear your comments, corrections, suggestions or ideas.


NLA presentation, Friday October 21, 2005, 10:30 - 12:00 Government Documents Interest Group (GODIG) Who's Government Information? Our Government Information! Presenters: Cornwall, Yeo, Jacobs --Introduction: Welcome and thanks for coming. My name is James R. Jacobs. I'm a government information librarian at University of California San Diego. These are my colleagues Daniel Cornwall, from the AK State Library, and Shinjoung Yeo, also at UC San Diego. The three of us -- along with Jim Jacobs (yes there are actually 2 of us, both in the same department at UC San Diego!! and James Staub from the Tennessee State Library -- started the advocacy organization Free Government Information almost a year ago as a way to reach a broader audience and create dialogue among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, citizens etc.) who have a stake in the preservation of and perpetual free access to government information. Today, we'd like to give some background of the debate as we see it currently working its way through the docs community, the libraries in the Federal Depository Library Program, the members of the Depository Library Council, and the Government Printing Office (GPO). We'll discuss the move, since about 1985, toward digital-only government information and how this move to digital is affecting and will continue to affect, how libraries do their jobs. If you've visited the website at http://freegovinfo.info, you know that we are concerned that the move to digital, although exciting and full of possibilities for increased library services, contains some major obstacles to overcome in order to provide free, fully-functional digital information while protecting users' privacy. We have discussed in various forums the need to continue the FDLP tradition of document deposit and will hopefully make a strong case here for digital deposit of government information as the primary means of preserving government information, giving widespread access to that information and protecting users' privacy in what they read and access online. --Background and history of government information: I'd like to start with some background and history of government information. First off, what exactly IS government information? I found this poignant quote from Anne Morris Boyd in 1949 that sums up what I mean when I talk about government information.
"Government publications, commonly called public documents, are among the oldest records, and if measured by their influence on civilization, are probably the most important of all written records. They are the sources of political, economic, and social history of peoples of all times; they contain the authentic accounts of the world's great explorations, discoveries, and inventions in every field of human endeavor; they reveal and explain the phenomenal scientific and technological developments of modern times; they open up great treasuries wherein man has attempted to give expression to his artistic impulses. They contain the history of civilization itself in all its aspects."
Perhaps a bit hyperbolic but you get her drift! Government information is the information collected, compiled and created by governments in their official capacity, funded by our tax dollars, and used by us for our common understanding and our common good. Government information belongs to citizens who have a right to know their government's activities in order to participate fully in the democratic process. Since 1860, there has been a system in place to insure public access to government information through a partnership between the Government Printing Office (GPO) and the hundreds of libraries in the Federal Depository Library Program (FDLP). Similar systems have been set up in nearly every state for state-produced information. The California State Library has been collecting documents from CA state legislative, executive, and judicial branches, commissions, etc. since 1850, before even the FDLP was around. CA has always been on the cutting edge! You may have seen Daniel's previous presentation which shows that now Alaska is on the cutting edge in terms of digital govt information!! He'll talk more about that later. In essence, the FDLP has provided for centralized processing for a deliberate, distributed collection able to have been well-preserved over time. For almost 150 years, this system has been largely successful in providing for one of the core tenets of a democracy, an informed citizenry, not to mention being a veritable goldmine for researchers across the academic disciplines. So now I've set the stage for Shinjoung who will talk about the current conditions we find ourselves in and ACCESS issues regarding the move to digital government information. I will then be back to talk about PRIVACY and then Daniel will wrap up with a discussion about PRESERVATION. PRIVACY: --Philosophy of privacy: Libraries have traditionally protected patron privacy as one of the core tenets of librarianship. Patron privacy is an integral part of the practice of intellectual freedom inherent in the First Amendment of the Bill of Rights to the US Constitution. Like the Hippocratic oath, ALA's Code of Ethics includes principles that guide our work as librarians. Among them are:
--We uphold the principles of intellectual freedom and resist all efforts to censor library resources. --We protect each library user's right to privacy and confidentiality with respect to information sought or received and resources consulted, borrowed, acquired or transmitted. ALA Code of Ethics
This professional code, along with individual libraries' strong privacy protecting policies, has led ALA and librarians to work tirelessly to protect patron privacy. In fact the American Bar Association, in an article in their ABA Journal in August of this year, recognized the growing lobbying clout of ALA, calling them, "one of the most active players in legal fights over technology, copyright, national security, censorship and privacy law." Librarians have been at the forefront in the battle against such legislation as the Children's Internet Protection Act (CIPA) and against certain provisions of the USA PATRIOT (Uniting and Strengthening America by Providing Appropriate Tools Required to Intercept and Obstruct Terrorism) Act (or USAPA in government documents parlance!) -- primarily section 215, the "library" section. Section 215 allows the FBI to access business records (including library circulation records showing what patrons are reading) with no probable cause and places a non-disclosure gag order on those librarians or libraries that are being asked for library records. USAPA precludes state statutes protecting private information. The strong stance against USAPA has led many libraries, in an effort to circumvent USAPA, to proactively delete circulation records and Web server logs and has led many to evaluate what personal information they keep and for how long. This strong historical pillar of privacy protection is perhaps most important where govt information is concerned. What you read about your government (or about anything else for that matter) needs to be protected and kept strictly confidential. In the analog world, this was easier to assure since govt information was distributed to libraries that then gave access and had the power and organizational policies to protect privacy. Privacy in the digital world is on much more tenuous footing. In the digital world the servers of the information decide how much privacy they'll let you have, and this decision is not necessarily based on Codes of ethics, or explicit policies protecting user privacy, but on business plans and economic realities. Downloading a digital document from a Federal Agency web site, or from a server at the Government Printing Office will allow those organizations to collect the following information:
  • The IP address of your computer and Internet provider.
  • The date and time you accessed their site.
  • Information about the document or piece of data downloaded.
  • The Internet address of the web site that referred you to their site.
  • Tracking information via cookies.
In other words, with the personal information that is collected as a matter of course in the digital world, the government will know what you're reading and from where you're reading it! --The Technology of privacy: The Government Printing Office (GPO), as everyone here is well aware of, are in the midst of creating their Future Digital System (FDSys) which, they maintain, will collect, describe, preserve and make accessible all past, present and future government information. On the face of things, this sounds like a good thing. However, this does not obviate libraries and librarians from doing what they have done so well for so long: selecting, acquiring, organizing, and preserving information; providing services for and access to that information; protecting the privacy of readers and users of that information; providing information without fees or stipulations. We have asked GPO about privacy concerns inherent in some of the technologies being used in the Future Digital System (Digital Object Identifiers (DOI) and Public Key Infrasructure (PKI)) and have asked for clarification on the use of Digital Rights Management (DRM) technologies within FDSys. We have been assured that GPO follows Office of Management and Budget (OMB) recommendations and has a long-standing tradition of protecting the privacy of its customers and users of digital information content. However, GPO has yet to release a statement on whether or not they will use or reject DRM technologies. This could be a fatal flaw for FDsys and could undermine libraries' abilities to serve our users. Despite these assurances, the technologies being implemented in the Future Digital System have the possibility of being used to abuse the privacy of users. Let me just say a few words about DRM, DOI and PKI. I promise to keep this short and not get into technological alphabet soup! Digital Rights Management (DRM) is an umbrella term comprising several technologies used to control or restrict the use of digital content. Two common technologies are the Digital Object Identifier (DOI) and Public Key Infrastructure (PKI). There are many others used for specific file formats (like CSS encryption for DVDs). Two things are certain: all of these technologies can be circumvented by other technologies; All of these technologies are based on systems having access to the user's private information (IP address, name, address, credit card number...). The Digital Object Identifier (DOI) -- which you may have seen in links to electronic journal articles in some of the larger article databases like this one -- is similar to a persistent URL (purl). But instead of simply being a persistent pointer to a digital object, the DOI is designed to verify the authenticity of a digital document, and also to check a user's authority to access a document. Thus the DOI is designed to protect copyright, and prevent "piracy." David Sidman, wrote a very interesting article about DOI in 2001 entitled "The Digital Object Identifier (DOI): The Keystone for Digital Rights Management (DRM)." Public Key Infrastructure (PKI) is a crytographic program that provides for third-party vetting of, and vouching for, user identities and will be used by GPO to certify electronic content as authentic and official. Think of PKI as a kind of digital water mark that is verified by a third party. The drawback of PKI is that it shifts the verification of trust in a document to a third party (usually a private company) and again relies on private information in order to work. Both the creator and user of information must have access to the public key in order to unencrypt the information. This means that the holder of the public key for a document can not only withdraw permission to access but can easily collect information on who is reading what. These tools by themselves can be inoccuous, harmless and even quite useful in the right situations. However, the private information used to make them run can also be saved, harvested and mined for other purposes. --Proposed Solutions: Does this mean we cannot have Internet access to government information without Uncle Sam looking over our shoulder? The jury is still out on this. In GPO's Future Digital System, the government will not need the USA PATRIOT Act because their content management system will be able to collect and mine personal data quickly and easily from their own servers. GPO takes great pains to note that the FDsys will be "policy-neutral." As my colleague Jim Jacobs has pointed out, policy-neutral does not mean "neutral policies." The FDsys is actually designed to accommodate changes to government information policy, including changes for how they deal with users' private information. One solution to this privacy conundrum would be to deposit electronic copies of government information with libraries and let the libraries serve the information on their own servers. That way, electronic government documents would be accessed from privacy minded libraries. The issue of trust would be shifted from government agencies and third party corporations to libraries, which have a long, well-established social role of providing information and protecting privacy. Even if the government used its new powers under the PATRIOT act, it would have to make literally thousands of requests to find out who had handled a given document. History tells us that the Government and private companies do not play well with others' personal information and often misuse their power to access personal information for their own goals. From 1956 - 1971, the FBI's counterintelligence program (nicknamed COINTELPRO), attempted to neutralize political dissidents by, among other things, targeting library records to find out which users were consulting sensitive -- although unclassified -- technical information in libraries. More recently, In 2001 the Bureau of Indian Affairs had its internet and email privileges taken away by a federal judge because of their gross mismanagement of and failure to protect Indian land trust records (Court appointed experts had actually hacked into and gained access to supposedly protected electronic files!). 4 years later, the BIA's Web site STILL says "temporarily unavailable." An August, 2005 report from the Government Accountability Office (GAO), titled, "Data Mining: Agencies Have Taken Key Steps to Protect Privacy in Selected Efforts, but Significant Compliance Issues Remain," (PDF) examined programs at various agencies like the Small Business Administration, the Agriculture Department's Risk Management Agency, the IRS, the State Department and the FBI. The report found that the agencies' implementation of privacy and security measures was haphazard. The General Services Administration (GSA), which provides a database to the State Department, actually claimed that the Privacy Act did not apply to its system. On the corporate side, there is even more evidence of mismanagement of private information. Here are a couple of examples of corporations selling personal information or otherwise not protecting private information: --In 1999, RealNetworks - the company that popularized streaming audio broadcasts, slipped some code into its RealJukebox music player that surreptitiously transmitted information about the individual listening habits of its 13.5 million users to the company's servers. The program secretly searched the users' hard drives for every type of music file, and, even worse, it made a note of every music disc they played in their CD-ROM drives. RealNetworks made no mention of this little intrusion in the privacy statement posted on the company's Web site. This clandestine snooping went on for months until Richard Smith, an Internet consultant, discovered it and blew the whistle on the company. -- In September, 2002, JetBlue sold the itinerary information of over 1.5 million passengers, including passenger names, addresses, and phone numbers to another private company called Torch Concepts, under contract with the US Army to do research on data mining. --Just last month, Yahoo provided private information to Chinese State Security authorities that helped to convict a Chinese dissident journalist named Shi Tao. --What presentation would be complete without a mention of google! Google's gmail service indexes private email in order to target advertising. While this is technically not illegal, it is, to my mind, unethical. Google has also amassed a huge amount of personal information, search history, images, email etc. meaning there is a great likeliness for abuse. Google records everything they can: For all searches they record the cookie ID, your Internet IP address, the time and date, your search terms, and your browser configuration. Increasingly, Google is customizing results based on your IP number. This is referred to in the industry as "IP delivery based on geolocation." Google retains all data indefinitely and has no data retention policies. There is evidence that they are able to easily access all the user information they collect and save. Google won't say why they need this data. Inquiries to Google about their privacy policies are ignored. When the New York Times (2002-11-28) asked Sergey Brin about whether Google ever gets subpoenaed for this information, he had no comment. But enough of that. I'd like to conclude with a question for this audience: Who do you want protecting your privacy and right to read in the digital age? Organizations (i.e., LIBRARIES) with a long history and a primary duty of protecting privacy, backed up by a strong code of ethics and institutional policies? Or organizations (government or private) who's tertiary duty MAY be to protect privacy, but who may have other reasons, nefarious or economic, for NOT protecting privacy? Now I'd like to turn things over to Daniel who will discuss preservation of government information. Continue reading

Continue Reading →

Nevada Library Assn presentation: Access

Below is the text of my part of FGI presentation at the Nevada Library Assn. Annual Conference on October 21, 2005. ShinJoung Yeo's part of the panel was regarding access to government information in the digital age. Please feel free to give us feedback.


I am only talking about the issue of access to government information here. But, it is important to think about the issue of access within a larger context since it is tightly intertwined with long term preservation, local control, privacy etc. We have to remember that Information has a cycle -- creation, collection, distribution, access and preservation. They are not mutually exclusive. With that said: 1. Why access? We are living in a society where economic forces are at the front of many social and political decisions. However I believe there are certain things in our society that need to be free or have to free from a purely economic motivation such as water, air, education, health care etc. I hope you all agree with me that government information falls into this category. If you haven't thought about government information in this way, I hope our talk today will convince you of its inherent importance. Imagine that you aren't able to access information about local environmental conditions (water and air quality…), or about current legislation pending in Congress, or find out about government research into cancer cures /for your loved one. When you think about government information in this way, access to government information is an inherent right of citizens. 2. Current Conditions From the beginning, the U.S. government has recognized the importance of government information. Title 44 was written to codify its importance in our legal system. Under Title 44, GPO has historically had primary responsibility for the printing, distribution, and sale of government publications. Thus, government publications passed through GPO and GPO distributed the publications to depository libraries. The geographically dispersed system of the FDLP libraries was then responsible for providing free, local access to government information. However, the recent development of the Internet and its associated technologies has brought a shift from paper to purely digital information, bypassing the depository library system. Judy Russell (who oversees the FDLP), the Superintendent of Documents, estimated that only 14 percent of federal government documents is deposited in the FDLP libraries. The other 86% is available only through the Internet and only from government-controlled Web servers. Ms Russell has stated that by 2007 fully 95% of all government information will be digital-only. Because much of government information is now being produced digitally and not in paper, GPO is doing much less printing and distributing less print materials to depository libraries. In addition, it is becoming increasingly routine for government agencies to produce their own documents digitally and make them available directly to the public through the Internet. So without visiting a physical library building / now people are able to access government information anywhere there is an Internet connection. Sounds great right? However, behind all this seemingly quick and easy access, there are far-reaching consequences of bypassing FDLP libraries. I am not saying this because I am a librarian. Rather, I am talking as a citizen here as well. If we don't take this issue seriously and critically now, then we might completely lose access to government information. 3. Who controls access to information? In the print world, the government collected, created bibliographic information, printed and distributed documents to FDLP libraries. After the government information was distributed in the FDLP libraries, the role of government was ended. However, in a digital world, it becomes up to government agencies and the GPO what information is accessible and how information is accessed. So Basically, the responsibility for access has shifted from libraries to the government. That does not mean that this responsibility cannot or will not shift back to libraries. I hope to persuade you that it MUST return to libraries. In response to the shift from print to digital, GPO is proposing the creation of a centralized digital content management system (called their Future Digital System or "FDSys") to provide access to all government information. GPO's proposal implicates that GPO will be responsible for the collection, description, access and preservation and will also bear the full cost of these responsibilities. In other words, libraries will relinquish their traditional responsibilities –collecting, organizing, providing free access along with services -- so they will be merely service points for helping patrons with search engines. In this scenario that I've just painted -- which is closer to reality than you might think -- what could affect, change, limit access to government information? I would like to talk about 4 areas that are related to access in the event that libraries no longer have collections: Economics, Technologies, Politics of government information, and the Digital Divide. Economics Let's say that GPO's funding is fine now, but there is no guarantee that the government will fund GPO at the level needed to continue to provide no fee access to digital government information. We are already seeing GPO needing to fight for funding and under constant pressure from the Office of Management and Budget (OMB). In the midst of a budget crisis, can we assume that GPO's funding will remain a government priority? In this hypothetical situation, where GPO is the sole information provider, if GPO's budget line fails then there will be no access at all to government information. Another possibility is that GPO or some government agencies might want to sell their information to private corporations for profit, or create a fee-based system based on cost-recovery, or even privatize popular or marketable documents or serials. GPO's strategic plan in November of last year states that they will provide free access AND distribute information on a cost recovery basis. Actually GPO tried to do this with GPO access about 10 years ago and failed due in part to the fact that FDLP libraries had the same information available for free. So, it would seem obvious that in order to make a profit or at least to recover their costs, GPO or agencies will need to create information that is somehow limited or less-than-fully-functional in order to be able to charge for fully-functional information. We're not saying that GPO WILL do this, but it seems to us that GPO's contradictory mission statement of free access and cost recovery will lead to reduction or limitations on free and fully-functional access. Technologies Technologies that GPO and other agencies implement could easily facilitate this fee based system -- restrict user access, and/or render digital documents unusable or barely usable. For instance, Digital Rights Management tools, which are designed to authenticate users to prevent piracy or copying copyrighted materials, can easily restrict access by identifying users based on whether or not they have paid a fee or subscription to access information in the FDsys. Even the Depository Library Council's Vision paper has recognized the problem of DRM and has stated that, "GPO should work with agencies to ensure that the standard for web-publishing is fully-enabled digital files." Content management systems like FDsys can be set up to reduce the functionality of information products -- prohibiting downloading, printing or transferring of text to other programs. This is not simply my paranoia. Systems like I describe are already in place. For instance, at the National Academies Press Website, users may view a document one page at a time for free, but a fee is charged for downloading or printing the document. Another issue that I'd like to bring up is the limiting of access based on software. GPO and other government agencies have decided in many instances to use only specific software for whatever reason. The decision to adopt a specific piece of software can restrict or limit users' ability to access needed information. For instance, FEMA's online application for assistance after Hurricane Katrina could only be used with Internet Explorer browser. Another example is GPO's own annual report. In order to view the fully functional annual report one would need to download the Vizio, a document viewing software and register with Vizio to use the software. As you can see it is not that difficult to manipulate technologies to limit or restrict information according to economic, political and social motivation. So do we really want to rely on the government to provide no fee and full-functional government information? Or do we want to rely on libraries have diverse technologies and backgrounds and are committed to free public access to government information? Politics of government information As you know the issue of access to government information is highly political. Depending on the political climate at any given time, what is able to be accessed might change regardless of public interest or public's right to know. In the current FDLP system, after the government deposits their documents, libraries have control over their collections. GPO can recall documents, but, in this system, it is difficult and cumbersome for the government to remove or alter or restrict those deposited documents located in FDLP libraries. However, in the digital realm, documents can be removed or altered or restricted. We have already seen this on many occasions. For instance, in May of this year, the Overseas Base Closing Commission report was released on the commission's web site, but the report was pulled off of the site a few days later by order of the Secretary of Defense who didn't like some of the information in the report. In March, 2002, the EPA announced that it would no longer allow direct access to its Envirofacts databases. EPA stated that "As part of our continuing efforts to respond to Homeland Security issues. I could go on and on and I bet you have a story like this as well. But you get the picture. In the digital realm, information can be more easily and strictly controlled to the detriment of those who need access to the information: citizens, students, researchers, mothers, etc. So the question is do we really want to put our trust to the government and expect them to provide us with free and fully functional government information? Or trust our 1300 libraries who have been fighting and advocating for people's access to information for 150 years? Digital Divide Another issue that has barely been on the radar in the docs community is the issue of the digital divide. Daniel raised this issue recently on Govdoc-l in a response to the DLC Vision Statement, but we have not heard others discuss this. This issue has largely been ignored because we all think of the Internet as ubiquitous. We frequently forget that those with lower incomes or in rural areas do not enjoy the privilege of internet technologies, or do not have access to high speed internet connections necessary to download and view large PDFs, audio or video files. According to the recent Pew report on the Digital Divide in the United States: 68% of adults use the Internet, 32% do not. 73% of adults live in a household with an Internet connection and 27% do not. 22% of adults have never used the Internet and do not have access in their homes. 38% of adults living with disabilities have access to the Internet. 22% of adults over 70 have Internet access whereas 53% of adults between 60 and 69 have access. 11% of Internet non-users say that getting access is too difficult, frustrating or expensive. We have to find ways to make sure that EVERYONE has access to government information -- not just those who are privileged -- and need to take into account those on the other side of the digital divide when making our decisions about future systems of government information or shifting our primary role as to merely being service center. 4. Solutions So what are the possible solutions for the problems that I've just described above. Go back to paper? Print out every digital government document? That will be highly unlikely. I think the white elephant in the living room that has been in front of our eyes this whole time is revitalizing FDLP instead of relinquishing its responsibilities to the government. I often hear that the roles of FDLP are changing because of digital technologies. Remember only the formats have changed. The primary roles of FDLP libraries (and libraries in general!) -- collecting, organizing, distributing, providing free access to government information -- have not changed and I don't think this should be changed. The only change that libraries need to make is accepting digital documents instead of paper documents. GPO needs to deposit digital documents instead of paper documents. It's simply a matter of a format change. Libraries dealt with microcards, microfiche, CDROMs, DVDs…, they'll and can deal with digital as well. How does digital deposit solve these problems? Let me count the ways! We believe that digital deposit and the creation of a digital FDLP system will solve the economic problem by dispersing and sharing the cost of digital access among many libraries instead of one chronically cash-strapped agency. Libraries will guarantee no fee access to information like they have been doing for the last 150 years. Digital deposit will allow libraries to save digital documents on their local servers, have local control over those collections, and share them with other libraries in a collaborative manner. By doing this, we, libraries, can assure the provision of no fee and fully-functional access and better and more expanded services to government information. By having local digital collections we will be able to use and reuse information and won't need to worry about possible restrictions, information alternation or removal. Additionally, fugitive documents will be greatly reduced. Digital deposit can also go a long way toward alleviating the digital divide by for instance, facilitating print on demand in libraries and/or allowing local libraries to create and maintain their own digital collections that could be used without having to have an expensive T-1 line, or burned to CDs or other media for off-line use. Some might argue that not every library can afford to create their own digital repository, do not have the technological know-how, or might not have a need for or have the staff to process local collections. These same arguments were brought up not so long ago in regards to other formats, and at the advent of the Internet age. We dealt with those changes and we can deal with the current set. 5. Conclusion I'd like to end with a story. In 1999, following World Bank advice and a condition for the country's development loans, the Bolivian government granted a 40 year privatization lease to a subsidiary of the Bechtel Corporation, giving the company control over the water on which Bolivian citizens needed to survive. Immediately, the company doubled and then tripled the water rates for some of South America's poorest families. How do you think the people of a small town in Bolivia -- a large majority of whom are poor peasants -- responded to this? How could these peasants even imagine challenging Bechtel, one of biggest multinational corporations in the world? The people believed that access to water was a sacred right, not a commodity to be bought and sold, so they fought with their lives for access to water. Eventually, because of the fervor of the protests, Bechtel's contract was rescinded and Bolivian citizens took back their water right. There are only a few countries in the world where the right to access to government information is given to citizens. I believe this is your privilege and it's your sacred right. I hope that the library community will fight to assure citizens' access to freely available, fully functional digital government information. Continue reading

Continue Reading →

The politics of printing: 9/11 commission report

An interesting article entitled, " UNDOCUMENTED EVIDENCE The Politics (and Profits) of Information: The 9/11 Commission One Year Later" has just come out from the Washington Spectator. The article, written by Max Holland, touches on various government information issues: privatization, the role of GPO and FDLP, access, government secrecy etc. As a consequencs of this high level commission's decision to give Norton the rights to publish their report instead of going through the GPO, much of the commission's work was not included in the commission report, was only available online (if at all) or through private publishers for substantial cost. This highlights the weakness of the govt information system. That is, Title 44 has no teeth to make commissions, agencies, etc comply with the law and go through GPO for their publications. If GPO had had a role in the publication of the commission's work, depository libraries would have had the commission's final report, as well as staff monographs and other supplemental volumes. Instead, only the 567 page final report -- with no index! -- has been deposited. Washington Spectator has graciously allowed free access to this article. If you have a problem with the link, please download the PDF of the article attached to this post below. We will also add it to the FGI library. Continue reading

Continue Reading →

Do agencies exclude search engines?

Introduction

Recently, a speaker at a conference said that Thomas, the Library of Congress site with extensive information about federal bills and legislation, excluded all search engines from indexing the site. Why, the speaker asked, would a government site wish to block search engines? While government agencies might, like most web sites, exclude search engines from parts of the site for benign reasons (such as blocking temporary pages, out-of-date pages, etc.) It seems counter-intuitive that anyone would want to block access to search engines since they tend to drive most traffic at most web sites.

This prompted me to examine if and when and how government agencies are blocking search engines. I discovered that Thomas does not block all search engines after all -- but it only allows Google to fully index the site. That raises more questions than it answers. What follows is some background and findings and conclusions from a selective examination of what some agencies are doing.

Background

Search engines like Google, Yahoo, MSN search, etc. build indexes of web pages by using software (called "robots" or "spiders") that browse (or "crawl") the web, finding web pages and the links on those pages, following the links they find to find more pages, and so forth. When the spidering software reaches a site, it is supposed to do a couple of things before proceeding. It is supposed to look for a special file named "robots.txt", read the rules in that file that tell it where it may go and where it may not go, and then obey those rules. The rules (called The Robots Exclusion Standard) allow the site to exclude or allow robots to visit particular parts of the site. This is a web convention, not a mandate, but all "well behaved" robots are supposed to follow these rules. The rules can exclude robots from a directory or a particular file or even exclude a particular type of file. They can also specify that one robot can get to one part of the site, while another cannot. (This is useful for site administrators that want to allow their own indexing software to index parts of the site that they don't want to be visible to the world, for example.) There are other ways to ask a search engine to refrain from indexing a page, but I don't cover those here.

Some explanations of how the robots.txt files work are:

I examined the robots.txt files at some dot-gov (.gov) sites. The following is a report on what I have found so far. I used google to identify files (google search for file robots.txt at site .gov) and examined a selection of the federal sites from that result. I invite others to let me know of other interesting exclusions.

Findings

It is not always easy to know the effects of a robots.txt file on a complex site. When the file is quite long and lists many directories that are themselves links to other parts of the site, for instance, it is possible that these are old links and spiders will find the information through another route. The EPA site is an example of this; it has 975 exclusions and, without following every one and navigating the site looking for the information in the exclusions, it is difficult to determine if the information is still findable by spiders.

That said, here are the few sites that I examined that had some interesting features.

  • Customs and Border Protection (visited 10/30/2005)
    Excludes all spiders except Google and the "gsa-crawler", which it allows everywhere.
  • Department of Justice (visited 10/30/2005)
    Exclude indexing of its archive
  • Department of Labor (visited 10/30/2005)
    Evidently excludes spiders from older press releases (e.g., /dol/media/press/archives.htm), decisions of the Employees' Compensation Appeals Board and the Benefits Review Board.
  • Environmental Protection Agency (visited 10/30/2005)
    Has a lengthy robots.txt file that excludes all robots from parts of the site such as the state Binary Base Maps of the Exposure Assessment Models, a page on the Columbia Space Shuttle Accident, the section on Environmental Monitoring for Public Access and Community Tracking, the entire "Human Exposure Database System," parts of the site that require a password (e.g., "mercury-archive"), and much more.
  • Thomas (visited 10/30/2005)
    Allows Google's robot complete access. Allows the "Ultraseek" robot any plain file, but excludes it from following any dynamic links (which is most of the site). Denies all other robots any access.
  • USDA Forest Service Southern Research Station (visited 10/30/2005)
    Whole sections of publications in pdf format are excluded (e.g., "/pubs/rp/*.pdf" [recent publications])
  • Whitehouse (visited 10/30/2005)
    An extensive list of over 2000 exclusions. Most of these seem to be simply aimed at excluding robots from the "text" versions of web pages that are probably duplicates of other pages; if so, these have essentially no effect on indexing but are probably used to reduce the number of times spiders hit the site and save bandwidth. Interestingly, though, most of the other exclusions are for directories named "iraq" that do not exist. Excluding non-existing directories is neither necessary nor functional. Perhaps these are in place to save bandwidth by preventing spiders from hunting for information about Iraq rather than following actual links. Examples: /infocus/healthcare/iraq /infocus/internationaltrade/iraq

Conclusions

It appears that most agencies are either not using robots.txt files to excluded search engine robots, or are using the exclusion rules sparingly and appropriately. At least two government web sites (whitehouse.gov and epa.gov), and perhaps others, warrant further study because they have such lengthy robots.txt files or because the exclusions otherwise obscure their overall effect.

Two other federal government web sites (customs.gov and thomas.loc.gov) have troubling exclusions of most search engines except Google.

Continue reading

Continue Reading →

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM