Home » Posts tagged 'Internet archive' (Page 6)

Tag Archives: Internet archive

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

Status of the Wayback Machine

Roy updates us on the status of the Wayback machine with an example from the White House:

  • Back to the Wayback Machine, Roy Tennant, Library Journal (May 18th, 2011). But that means that any claims to be "archiving the web" should be taken with a grain of salt. Maybe say "archiving the parts of the web that matter" or "ignoring what doesn’t matter so much".
And, don't forget Archive-It, the web archiving service from Internet Archive.
Through a user-friendly web interface, Archive-It partners can catalog, manage, and browse their archived collections using web archiving tools developed at the Internet Archive. Collections are hosted at the Internet Archive data center and are accessible to the public, including full-text search.
Continue reading →

Continue Reading →

book scan wizard + internet archive = DIY public domain digital book repository

At the last couple of depository library council meetings, I've heard comments from documents librarians -- especially from librarians at smaller institutions -- that they'd love to participate in the digitization process of historic government documents, but for various reasons (lack of $$, staffing, time, technical infrastructure etc) could not undertake large scale digitization projects. Now there's a way for lots of libraries to chip in on the greater goal of increased access to historic government documents with very little $$ or infrastructure. We've mentioned before about BookLiberator and DIYbookscanner, two projects working on low cost hardware solutions for digitizing books using off the shelf digital cameras and free opensource software called Book Scan Wizard. But there were still 2 pieces missing to make the whole workflow run smoothly for libraries and government documents collections of all sizes. The third piece to the puzzle just became a reality with yesterday's announcement that Book Scan Wizard had teamed up with the Internet Archive to provide automatic uploads of scans to the Internet Archive (directions and more information here). Hardware: check. Software: check. Digital infrastructure: check.

With the new version of Book Scan Wizard, or even through just uploading directly to the Internet Archive, any PDF composed of images of book pages or organized zip file filled with images of book pages will be automatically processed. The Internet Archive’s servers will then automatically perform optical character recognition (OCR) on the book and make a pdf, epub, kindle (mobi), daisy, djvu, and text file copy of the entire book available for download by anyone, anywhere. You can see a sample book from this process to get a better idea. All this happens within a few hours of the book being uploaded and then anyone can download it. This is free OCR for anyone in the world.
Now there's one last piece needed: Scan on demand. This idea has already been put into practice by the Internet Archive's Open Library and their partnership with the Boston Public Library. What we need is to open up the Catalog of Government Publications (CGP) -- which will soon include over 1 million records from GPO's historic shelflist spanning 1870s - 1992 -- similar to the way the BPL's scan on demand project (now retired it seems) allowed users to request a scan of a public domain book directly from the Open Library catalog. GPO could manage this scan on demand process -- or allow libraries to pick and choose documents from the CGP -- connect the bibliographic metadata from the historic shelflist, and upload to both the Internet Archive and FDsys. The circle is complete. Am I missing anything? Would love to hear readers' thoughts.
Continue reading →

Continue Reading →

Holiday gift idea: a piece of the public domain

Carl Malamud's FedFlix project is a joint venture with the National Technical Information Service (NTIS) whereby he takes NTIS videos, digitizes them and uploads them to the Internet Archive. Well now he's expanding FedFlix to include public domain videos from the National Archives. He's released 41 videos into the public domain in this way, but has put together an Amazon Wish List in order to expand public access to public domain video content from the National Archives. If you see anything you'd like to buy the public domain, they'll take your DVD and upload the video to YouTube, the Internet Archive, and to public.resource.org's own rsync/ftp public domain stock footage library. So why not add a gift of the public domain to your favorite person's/people's stockings this year? We'll all be glad we did! UPDATE 12/25/09: The wish list has been fulfilled. You can watch all of the donated NARA videos on YouTube, Internet Archive, or public.resource.org's bulk server. Thanks Carl! [HT BoingBoing!] Continue reading →

Continue Reading →

Tools of Change: BookServer

This is a presentation that is definitely worth seeing!

  • Web of Books, Peter Brantley, Presentation at Tools of Change Frankfurt introducing the BookServer architecture. Brief description of history, motivations, and technical outline. "Creating a new architecture using common, open standards that permits people to ?nd, buy, acquire, and read books from any source, on any device, using many different ebook applications."
Continue reading →

Continue Reading →

Internet Archive proposal for mass digitization

I had known that the Internet Archive had submitted a response to the GPO's RFP for mass digitization. A friend just sent me the link to the proposal submitted to GPO (embedded below and here's the link to the proposal and supporting documents). As you can probably guess, we've been pulling for the Archive to get the bid, not least of which because the Archive is a 501(c)(3) non-profit library and we've stated on more than one occasion that privatization of public domain government information is a very bad idea. But also, we've been heartened by the quality of the Archive's scans to date, their openness and willingness to be collaborative in their processes and data access and sharing. Those qualities certainly come through in their proposal for mass digitization -- not to mention the fact that they've actually made their proposal public! While the award has not been officially announced, we really hope that the Archive wins the award. Perhaps GPO will name them as an official depository library and work with them not only on the "legacy" collection (there needs to be a better description of the deep and rich collections of depository libraries than the somewhat pejorative "legacy" :-| ) but on digital deposit of government documents going forward. --that is all.

Continue reading →

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM