New CRS Report on Data Mining
Thanks to Docuticker for pointing out this new Congressional Research Report on Federal data mining efforts:
Data Mining and Homeland Security: An Overview (PDF; 231 KB)
RL31798
Source: Congressional Research Service (via Federation of American Scientists)
Aside from cataloging currently known datamining efforts by the federal government, the report identifies four areas of concern:
As with other aspects of data mining, while technological capabilities are important, there are other implementation and oversight issues that can influence the success of a project’s outcome. One issue is data quality, which refers to the accuracy and completeness of the data being analyzed. A second issue is the interoperability of the data mining software and databases being used by different agencies. A third issue is mission creep, or the use of data for purposes other than for which the data were originally collected. A fourth issue is privacy. Questions that may be considered include the degree to which government agencies should use and mix commercial data with government data, whether data sources are being used for purposes other than those for which they were originally designed, and possible application of the Privacy Act to these initiatives.
I've heard people say that data mining by the government is no big deal since advertisers and other corporate interests do it all the time in efforts to focus marketing and improve profits. If we don't have privacy from corporate types, why should the government worry us? Because ratty data used by a marketer might result in a bald man getting shampoo ads, but when the government relies on ratty data for law enforcement, innocent people can get jailed or harrassed.
Hopefully Congress will be more vigilant on this issue.
Continue reading
Happy 75th to UCLA!
In a govdoc-l posting on December 21, 2007, Jan Goldsmith informed the depository community that the Federal Depository LIbrary at UCLA celebrated it's 75th anniversary. Her post states how they celebrated:
The UCLA Library celebrates its 75th anniversary as a member of the Federal Depository Library Program this year. We were designated a depository in 1932 by Senator Samuel Shortridge. Over the course of these past years, we have provided free access to federal government information for the University community and beyond. In honor of our anniversary, Kris Kasianovitz and Jan Goldsmith, with the professional assistance of Ellen Watanabe and Dawn Setzer, mounted an exhibit in the Young Research Library lobby. Gary Strong signed letters notifying all legislators whose district we serve, including our Senators and Congress representatives. For more information regarding 75 years of collection and services of the Federal Government Documents at the UCLA Library visit our Anniversary web page, at: http://www.library.ucla.edu/libraries/yrl/depository_exhibit.html
Let's hear it for UCLA for providing three quarters of a century of public service to the people of Los Angeles!
Continue readingGovernmentDocs.org provides enhanced access to FOIA documents
Back in November, Citizens for Responsibility and Ethics in Washington, along with Project on Government Oversight, Public Citizen, Electronic Frontier Foundation and the Sunlight Foundation launched a new website, GovernmentDocs.org, which will house governmetn documents obtained through Freedom of Information Act (FOIA) requests. From the press release:
The database will house Freedom of Information Act (FOIA) responses, and other government documents, from a number of organizations, that can be browsed, searched and reviewed. It is the only one of its kind.
Traditionally, government watchdog groups have either posted FOIA documents on their websites as unsearchable PDFs, or statically highlighted several pages within a document to bolster their findings. This has historically limited the public's access to FOIA documents, and minimizes the opportunities for use by researchers, journalists and citizen reviewers for further research and disclosures. Governmentdocs.org changes that:
- Each and every document goes through an optical character recognition (OCR) process, so that the text of each document is entirely searchable.
- A powerful search engine provides full-text searches and hit highlighting.
- Citizen reviewers can add information to each document page and highlight important findings, allowing for more robust and targeted searches.
- Every page of every document has its own unique URL so that documents can be linked, shared, or posted onto websites.
- The database is a coalition effort, so all of the organizations' documents will be housed on governmentdocs.org and searches will work across collections.
Continue reading
How many government websites are there?
Still No Directory of Federal Websites, E-Gov Act Ignored. By Coby Logen, .gov Watch. November 05. 2007
How many government websites are there? How many HHS or DOJ sites are there? You and I have no way to know. American taxpayers cannot even know how many public websites their government is funding. By law, we should—but the system is broken.
The E-Government Act of 2002 set a deadline of two years to develop a "public domain directory of public Federal Government websites" (Section 207(f)(3)). But this directory still does not exist 5 years later.
Continue reading
Most fed data is un-Googleable
As we've noted here before (Is your search engine finding the government information you need?), the problem of relying on commercial search engines to find government information is that a lot of government information on the web goes un-indexed by those search engines.
- Most fed data is un-Googleable By Jason Miller FCW (December 17, 2007). "After five years, a major E-Gov Act provision goes unmet because of search problems."
Sen. Joseph Lieberman (I-Conn.), chairman of the Homeland Security and Governmental Affairs Committee, says "There are more than 2,000 federal government Web sites not included in commercial search engine results. Is it accidental, or is there a policy, or it is just laziness? I would like to know why" and "Agencies do not let commercial search engines index their sites."
I wonder if that is true? I wonder if there is any document librarian who can answer that question or point to which sites are not indexed?
It is probably closer to the truth to say, along with John Needham, Google's manager for public-sector content partnerships, that government "databases" are being missed by web crawlers and that "Agencies are concerned more about how information is presented than if users are finding it." In other words, agencies would probably like to have their information indexed, but haven't figured out how to do so, or don't have the budgets to do what is necessary. It probably isn't "laziness" but lack of funds and other resources; it probably is sometimes "accidental" in that some may not know what to do. It is probably sometimes even "policy" -- but probably less often.
But, one big problem is that we don't really know the scope of the problem or the cause. FDLP librarians should be pushing GPO, researchers, and library schools to research these issues so we have answers.
Continue reading