Home » Posts tagged 'Government websites'

Tag Archives: Government websites

Our mission

Free Government Information (FGI) is a place for initiating dialogue and building consensus among the various players (libraries, government agencies, non-profit organizations, researchers, journalists, etc.) who have a stake in the preservation of and perpetual free access to government information. FGI promotes free government information through collaboration, education, advocacy and research.

Sunlight’s Web Integrity Project shutting down. A sad day for govt oversight and transparency

This is bad news for government transparency advocates. The Sunlight Foundation shut down its Web Integrity Project (WIP) yesterday. WIP was created almost two years ago and was tracking “tens of thousands of federal government webpages each week, has reported on and sourced hundreds of stories about federal websites, has provided materials for congressional oversight, […]

Continue Reading →

FDLP.gov hacked — cute but a little NSFW

I happened to surf over to the FDLP site today around 4:30pm and found the FDLP.gov had been hacked and taken over by SoWa BeZ OkA — which translates from Polish into “Owl without an eye.” The group seems to be a band of some kind, but I can’t tell. But it wasn’t just a […]

Continue Reading →

Government Shutdown: Status of federal websites

Update #2 10pm PST 10/2/13 : Our friends over at the Sunlight Foundation have an interesting post, "What Happens to .gov in a Shutdown?" They explained the .gov shutdown matrix:

...drawn on an agency-by-agency basis, and the specific determination is based on the importance of the function and how illegal ceasing to do it might be. But aside from some obvious ones--national parks would be closed; the CO2 scrubber on the International Space Station would stay plugged in--it'll be agency leadership that makes the determinations.
(and love the unix joke!) UPDATE #1 3pm PST 10/2/13: Arstechnica, checked 56 .gov sites and found 10 that went dark. See "Shutdown of US government websites appears bafflingly arbitrary." A bunch of federal websites will shut down with the government, By Andrea Peterson, Washington Post, Published: September 30 at 5:28 pm. Continue reading

Continue Reading →

Healthcare.gov: The Innovative Development of a .gov Website

In October, the healthcare.gov website will be the site millions of Americans use to choose their health insurance. The new site has been built in public for months, iteratively created on Github using cutting edge open-source technologies. Healthcare.gov is the rarest of birds: a next-generation website that also happens to be a .gov. It will use Jekyll, which allows developers to build a static website from dynamic components. This will make the website faster and more efficient. A fascinating story! Continue reading

Continue Reading →

US Executive Branch Closure Crawl

The State of the Federal Web Report issued in late 2011 noted that Federal agencies planned to eliminate or merge several hundred domains, as part of the President's Campaign to Cut Waste. The goal was to reduce outdated, redundant, and inactive domains. As part of this work, the .gov Task Force overseeing the process asked members of the National Digital Stewardship Alliance (NDSA) to archive and preserve all .gov Executive branch domains slated to be decommissioned or merged. NDSA members immediately agreed that an important step in this process was to preserve the content of these sites as part of our national digital heritage - instead of simply eliminating them. Rather than start a separate, standalone project, we chose to launch a collaborative crawl under the auspices of the End of Term Web Archive project (EOT). Although the EOT project has primarily focused on transitions occurring at the end of administrative terms, part of the goal of the project is to document changes in all online presences of the US Federal government during key periods of transition, regardless of when or under what circumstances they occur. So, a comprehensive harvest, using a targeted list of domains supplied by the .gov Task Force and a general list of all Executive branch domains downloaded from data.gov, began on Saturday, October 8, 2011. The crawl concluded on November 5, 2011 and encompassed 46,278,384 captures and ~13TBs of data compressed. Here's a general outline of the sequence of events of the Fall 2011 crawl:

  • Agencies identified recommended actions for domains in their Interim Progress Reports and Web Inventory
  • The .gov Task Force collected a list of outgoing .gov domains and shared those with the NDSA
  • Internet Archive crawled outgoing sites and the full suite of Executive branch domains (note: for some resources it took several weeks to crawl sites in their entirety)
  • GSA eliminated domains after they were archived
The End of Term Web Archive project, including the archival capture of Executive Branch domains last Fall, is not meant in any way to satisfy agency records management obligations. The domains are archived solely for the purpose of preservation and posterity. Agencies separately discuss records management obligations and handle those processes independently. However, we do make every effort to replicate resources in their entirety – at least what can be supported by available tools, techniques and best practices. Some portion of every web site is housed server-side and that subset of content and/or user experience cannot be archived and replicated using traditional web crawler/capture software that is dependent on files being downloaded to the client. The biggest challenge of this project, however, was not Web 2.0/Web 3.0 server side rendering or content serving. The biggest limiting factor was time. When we archive resources, there is a big difference between visiting and sampling a web resource using a set of scoping rules and guidelines versus going out and attempting to “drain” a site, i.e. replicate it soup to nuts as fast as the server can respond to your requests. Some of these resources house thousands to tens of thousands of PDF files, videos &/or other network intensive resources. And, most servers are programmed to meter how fast they respond to requests from the same IP address or an IP address range, so we have to wait appropriate intervals between requests in order to avoid being ignored or blacklisted by an automated process. There are ways to parallelize capture, but without dedicated funding, few institutions are able to marshal those kinds of resources on a volunteer basis. The End of Term project is built on the collaborative best efforts of a network of partners who share a passion for preservation of online government. For more information about the streamlining of agency website management, please visit www.usa.gov/WebReform.shtml. This effort is now part of the larger Digital Government Strategy. For more information on the End of Term Web Archive project, please visit http://eotarchive.cdlib.org, and follow us @eotarchive. Kris Carpenter Negulescu Director Web Group Internet Archive Continue reading

Continue Reading →

Latest Posts

Latest Comments

Blogroll

Archives

Meta

Archives

Powered by WordPress / Academica WordPress Theme by WPZOOM