Home » Posts tagged 'digitization' (Page 6)

Tag Archives: digitization

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

The Elusive Prerequisite: Text Extraction

As we look to relying increasingly on digital texts for discovery, access, and use of government information, it is worth understanding the issues (and difficulties) in accurately extracting text from a variety of sources. Here is a thirteen page paper that outlines the issues.

  • Herceg, Paul M., and Catherine N. Ball, Reliable Electronic Text: The Elusive Prerequisite for a Host of Human Language Technologies, Mitre Technical Report (Mitre, 30 September 2010). Electronic text is a prerequisite for text-processing applications such as indexing and search, named entity recognition (NER), and machine translation (MT). For example, text that is printed on paper must first be converted to electronic form to make it searchable--typically, by scanning the pages and using Optical Character Recognition (OCR) to create electronic text. Not all electronic text is created equal, however. Electronic text comes in a wide variety of containers, from Microsoft Office documents to Oracle database records. Electronic text comes in a wide variety of character set encodings, such as Chinese "Big 5" or Unicode UTF-8. Electronic text may be written in any of the world's languages, including the special sublanguages of blogs, chat and forums. And from the end user's perspective, even formatting aspects of the electronic text may be important: for example, if the use case is translating a document for a customer, and the customer requires the translated document to be formatted to look like the original. From the application perspective, all aspects of electronic text potentially make a difference. Almost all products are language-specific, and will produce useful results only on supported languages. In terms of file types, some products accept Microsoft Word documents as input, while others may accept only plain text files. Some products may work best if all text is "normalized" to a single character set encoding such as UTF-8. The purpose of this paper is to explore factors to be considered when trying to match up text-containing files (in a variety of file formats, character set encodings, languages, etc.) with text-processing applications and use cases.
Hat tip to Info Docket! Continue reading

Continue Reading →

book scan wizard + internet archive = DIY public domain digital book repository

At the last couple of depository library council meetings, I've heard comments from documents librarians -- especially from librarians at smaller institutions -- that they'd love to participate in the digitization process of historic government documents, but for various reasons (lack of $$, staffing, time, technical infrastructure etc) could not undertake large scale digitization projects. Now there's a way for lots of libraries to chip in on the greater goal of increased access to historic government documents with very little $$ or infrastructure. We've mentioned before about BookLiberator and DIYbookscanner, two projects working on low cost hardware solutions for digitizing books using off the shelf digital cameras and free opensource software called Book Scan Wizard. But there were still 2 pieces missing to make the whole workflow run smoothly for libraries and government documents collections of all sizes. The third piece to the puzzle just became a reality with yesterday's announcement that Book Scan Wizard had teamed up with the Internet Archive to provide automatic uploads of scans to the Internet Archive (directions and more information here). Hardware: check. Software: check. Digital infrastructure: check.

With the new version of Book Scan Wizard, or even through just uploading directly to the Internet Archive, any PDF composed of images of book pages or organized zip file filled with images of book pages will be automatically processed. The Internet Archive’s servers will then automatically perform optical character recognition (OCR) on the book and make a pdf, epub, kindle (mobi), daisy, djvu, and text file copy of the entire book available for download by anyone, anywhere. You can see a sample book from this process to get a better idea. All this happens within a few hours of the book being uploaded and then anyone can download it. This is free OCR for anyone in the world.
Now there's one last piece needed: Scan on demand. This idea has already been put into practice by the Internet Archive's Open Library and their partnership with the Boston Public Library. What we need is to open up the Catalog of Government Publications (CGP) -- which will soon include over 1 million records from GPO's historic shelflist spanning 1870s - 1992 -- similar to the way the BPL's scan on demand project (now retired it seems) allowed users to request a scan of a public domain book directly from the Open Library catalog. GPO could manage this scan on demand process -- or allow libraries to pick and choose documents from the CGP -- connect the bibliographic metadata from the historic shelflist, and upload to both the Internet Archive and FDsys. The circle is complete. Am I missing anything? Would love to hear readers' thoughts.
Continue reading

Continue Reading →

GPO and LoC to collaborate on two projects to enhance digital access

Here's some good news on this stormy day (at least in NorCal). GPO and the Library of Congress are set to work together on better digital access for the historic [w:United States Statutes at Large] and the [w:United States Constitution]. Anyone want to add this to the the [w:Conan the Librarian] wikipedia page?

The U.S. Government Printing Office (GPO) and the Library of Congress (LOC) recently received approval from the Joint Committee on Printing (JCP) to proceed on two collaborative efforts. One project involves the digitization of some of our nation's most important legal and legislative documents and the other involves enhanced public online access to the Constitution of the United States: Analysis and Interpretation (CONAN). The digitization project will include the public and private laws, and proposed constitutional amendments passed by Congress as published in the official Statutes at Large from 1951-2002. GPO and LOC will also work on digitizing official debates of Congress from the permanent volumes of the Congressional Record from 1873-1998. These laws and documents will be authenticated and available to the public on GPO’s Federal Digital System (FDsys) and the Library of Congress’s THOMAS legislative information system. The other project will provide enhanced public online access to the Constitution of the United States: Analysis and Interpretation (CONAN), a Senate Document that analyzes Supreme Court cases relevant to the Constitution. The project involves creating an enhanced version of CONAN, where updates to the publication will be made available on FDsys as soon as they are prepared. In addition to more timely access to these updates, new online features will also be added, including greater ease of searching and authentication. GPO authenticates the documents on FDsys by digital signature and these authenticated documents are also available on the Library’s THOMAS system. This signature assures the public that the document has not been changed or altered since receipt by GPO. This digital signature, viewed through the GPO Seal of Authenticity, verifies the document’s integrity and authenticity.
Continue reading

Continue Reading →

Smithsonian digitization strategic plan

The Smithsonian has just released their digitization strategic plan for fiscal years 2010 - 2015 called "Creating a Digital Smithsonian" -- executive summary and full report. I'm in 2 minds about this as well as similar digitization plans. On the one hand, the digitization of Smithsonian collections -- books, research reports, data, music, film and other sounds (like frog vocalizations!) -- will mean potentially a boon to online access to some really amazing materials. On the other hand, this quote from the executive summary worries me:

To preserve our collections, the Smithsonian constantly battles the destructive forces of time and environment. Despite our best efforts, plastics discolor, wax cylinder recordings distort, and botanical specimens become brittle. Digitization offers a way to make objects — and the valuable information they contain — available without jeopardizing their integrity by handling or by exposure to the elements.
While they mention a "life cycle-management approach to digitization," there doesn't seem to be a serious amount of thought given to the fact that digital objects degrade faster than physical objects, and that digital preservation is an ongoing and potentially more expensive effort. I worry that SI.edu will broker the same kind of disastrous deal that GAO did with Thomson-West whereby a whole swath of public domain information was privatized. I would call on SI.edu and ALL .gov agencies to insert a clause into ANY digitization contract that ALL digital files and metadata will be accessible via free and open sites. That means where applicable, copies of all digital content would be ingested into GPO's FDsys, Library of Congress, NARA and/or publicly accessible non-profit sites (eg. UNT digital library or Internet Archive). Please help us get this message across to your friends in the .gov sector. Public information should remain public! Continue reading

Continue Reading →

Open Library redesign and proposal for collaborative digitizing of documents

The Open Library announced yesterday that their redesigned site is now available with lots of new features and functionalities. As I suggested in my tweet a few minutes ago, wouldn't it be great if lots of depository libraries bought cheap book scanners like the Decapod (A Mellon funded project), digitized government documents and uploaded them to the Open Library? There are tons of records for government documents just waiting for the attachment of a digital file. And GPO could help by sharing their records from the Catalog of Government Publications (CGP) with the Open Library where librarians and others could enhance to make more robust metadata (which could be fed back in to the CGP!). Lots of libraries with Decapods make light work! (Full disclosure: I'm on the board of QuestionCopyright, a 501(c)(3) non-profit which has its own book scanning hardware/software project called Book Liberator. BL developers are in close contact with Decapod folks. But I get no economic benefit from either Book Liberator or Decapod.) Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM