Home » Posts tagged 'Data' (Page 6)
Tag Archives: Data
Govistics and the need for library data microservices
Several of us here at Stanford library who deal with data and/or govt information have recently received emails asking if we'd be interested in a free trial of the Pro level of subscription to the Govistics Government Spending Database built by the Center for Governmental Research (CGR). I'm a sucker for free trials, so took them up on their offer. Here's what I found -- and please take it with an FGI grain of salt ;-) The interface is easy for quick results and high-level comparisons, but I found it lacking for any kind of in-depth scholarly pursuits -- the researchers and students I work with would most likely be interested in historic data for all counties or all municipalities in a state or region or ALL states; and they'd probably want the data exportable so they could do further analysis with a statistical package (SPSS etc) or GIS software. I also didn't find the maps or charts particularly compelling. $50/year for an individual subscription (I didn't ask about an institutional subscription) seems like too steep a price to pay when there are other *free* tools out there -- my personal favorite is Many Eyes (also check out their new project Many Bills visual bill explorer!). Many Eyes allows a person to upload datasets, share them, run a variety of visualizations (charts, graphs, maps, clouds etc), and most importantly embed those visualizations in other Web pages. Govistics doesn't do any of that. And what about the underlying data you say? Govistics is basically US census of govts which is available for free on factfinder.census.gov (although only in PDF with no data export :-|). Many of the same variables are also available via the Census' County and City Data Book (again only PDF :-|). Govistics only offers data export with the pro version and the data only goes back to 2007. I don't begrudge the govistics folks trying to make a quick buck on public domain data that's already available online for free (well maybe a little). Perhaps for the casual user, this service will work well. But what I'd love to see is libraries creating interfaces like this *for free*. There needs to be free tools that include access + visualization + preservation. UVA has done gotten a great start with their historical county and city data books 1944-2000(!). This is especially cool because it not only gives access to historical data back to 1944 (no visualization yet, but users can use Many Eyes!) and allows for export of data for reuse, but it provides a preservation model as well. And THAT'S why I'd love to see more libraries doing this sort of thing. This is an increasingly data driven world and it would behoove libraries to combine these kind of access/visualization services with libraries' traditional strength in long-term preservation. --that is all. Continue reading
New maps, charts, tables from BLS’s Quarterly Census of Employment and Wages
A very nice new online application from the Bureau of Labor Statistics:
- Introducing the QCEW State and County Map Application The Bureau of Labor Statistics (BLS) has developed an interactive state and county map application available at http://beta.bls.gov/. The application displays geographic economic data through maps, charts, and tables, allowing users to explore employment and wage data of private industry at the National, State, and county level. Throughout this application, URLs are specific to the data displayed, so links can be bookmarked, reused, and shared. The application includes maps, charts, tables, and a link to standard BLS data tables and graphs.
- QCEW State and County Map
Changes at Data.gov
If you haven't looked at data.gov lately, you should. It was launched one year ago and has had a bit of a makeover recently and has added lots of new data. OMB Watch has a quick overview and comment about the current state of data.gov (Data.gov Celebrates First Birthday with a Makeover, by Roger Strother, OMB Watch. 05/24/10). Check out these highlights:
- Apps where developers are creating a wide variety of applications, mashups, and visualizations. From crime statistics by neighborhood to the best towns to find a job to seeing the environmental health of your community...
- Semantic Web where they highlight a set of data.gov resources reformatted into Resource Description Framework (RDF) format. These allow new kinds of rich interaction with the data. See, for example, the White House Visitor Search. Also see the Thetherless World Weblog from the Rensselaer Polytechnic Institute where some of this work is being done.
Workshop: Providing Social Science Data Services: Strategies for Design and Operation
Announcement of Workshop: Providing Social Science Data Services: Strategies for Design and Operation http://www.icpsr.umich.edu/icpsrweb/sumprog/courses/0041 August 9-13, 2010 Ann Arbor Michigan Instructors: Chuck Humphrey, Head of the Data Library, University of Alberta Jim Jacobs, Data Services Librarian Emeritus, University of California San Diego This five-day workshop is being offered for individuals who manage or provide local support services for ICPSR and other numeric data for quantitative research. Providing access to data has taken on greater prominence over this past decade with the emergence of several significant developments, including, e-Science infrastructure funding, the open data movement, national and institutional digital preservation strategies & services, data enclaves for confidential data, lifecycle data management planning, and data mash-up technologies on the Internet. Given these major environmental changes, how does one plan and design appropriate levels of data service in her or his local institution? This workshop is structured around a five-stage data lifecycle model that focuses on data production, data dissemination, data repositories, data discovery and data repurposing. A day is dedicated to each stage in this model during which discussions address issues for local data services and computer exercises demonstrate service activities. In this context, fundamental data topics are covered, including understanding the data reference interview, working with variables, interpreting data documentation, coping with various dissemination formats, accessing different online services (e.g., SDA and Nesstar), searching for social science data, subsetting data using Web-based tools, selecting and downloading ICPSR data, and options for local data delivery. Throughout the workshop, an emphasis will be placed on social science concepts and terminology, as well as on practical solutions to service delivery. Who Should Attend: Anyone who is new to providing services for numeric social science data or is seeking to revitalize an existing service. This is not a course in statistics and attendees are not expected to know how to analyze data. Online Registration: http://www.icpsr.umich.edu/icpsrweb/sumprog/2010/index.jsp Workshop will remain open only until the Summer Program office has received 20 paid applications. Questions? If you have questions about registration, fees, travel, housing, or other courses at the ICPSR Summer program, please get in touch with ICPSR directly: http://www.icpsr.umich.edu/icpsrweb/sumprog/contact.jsp If you have any questions about the workshop content, please feel free to send email to Chuck or Jim: Chuck: humphrey at datalib.library.ualberta.ca Jim: jajacobs at ucsd.edu Dates: August 9-13, 2010 Location: University of Michigan, Ann Arbor MI. Fees (Participants from ICPSR member institutions): $1,500 Fees (Participants from institutions that are not members of ICPSR): $3,000 http://www.icpsr.umich.edu/icpsrweb/sumprog/2010/application.jsp List of ICPSR member institutions and Official Representatives: http://www.icpsr.umich.edu/icpsrweb/ICPSR/membership/ors.jsp Information about transportation and housing: http://www.icpsr.umich.edu/icpsrweb/sumprog/visiting.jsp This workshop is part of the ICPSR Summer Program in Quantitative Methods of Social Research http://www.icpsr.umich.edu/icpsrweb/sumprog/index.jsp --- James A. Jacobs jajacobs at ucsd.edu Continue reading
What do we mean by “effective” access to data ? (Part II)
In my last post, I described the possibility of a systematic approach to data validation. A key feature of such an approach must be it’s availability to all who are responsible for data – and of special importance, its capacity to support efficient and timely use by creators or managers of data. Bill Michener (UNM), leader of one of the currently funded DataNet projects has published a chart describing the problem of “information entropy” [SEE: WK Michener “Meta-information concepts for ecological data management,” Ecological Informatics 1 (2006): 4 ] Within recent memory, I have heard an ecologist say that were it not possible to generate minimally necessary metadata “in 8 minutes,” he would not do it. Leaving aside -- for now -- the possibility of applying sticks and/or carrots (i.e. law and regulations, norms and incentives), it seems clear that a goal of applications development should be simplicity and ease of use. [ Within the realm of ecology, a good set of guidelines to making data effectively available was recently published – these guidelines are well worth reviewing and make specific reference to the importance of using "scripted" statistical applications (i.e. applications that generate records of the full sequence of transformations performed on any given data) this recommendation complements the broader notion -- mentioned in my last post -- of using work flow mechanisms like Kepler to document the full process and context of a scientific investigation. SEE “Emerging Technologies: Some Simple Guidelines for Effective Data Management” Bulletin of the Ecological Society of America, April 2009, 205-214. http://www.nceas.ucsb.edu/files/computing/EffectiveDataMgmt.pdf ] As a sidebar, it is worth noting that virtually all data are “dynamic” in the sense that they may be and are extended, revised, reduced etc. For purposes of publication – or for purposes of consistent citation and coherent argument in public discourse – it is essential that the referent instance of data or “version” of a data set be exactly specified and preserved. (This is analogous to the practice of "time-stamping" the citation of a Wikipedia article...) Lest we be distracted by the brightest lights of technology, we should acknowledge that we now have available to us, on our desktops, powerful visualization tools. The development of Geographic Information Systems (GIS) has made it possible to present any and all forms of geo-referenced data as maps. Digital imaging and animation tools give us tremendous expressive power – which can greatly increase the persuasive, polemical effects of any data. (For just two instances among many possible, have a look at presentations at the TED meetings [SEE: http://www.ted.com/ ] or have a look Many Eyes [SEE: http://manyeyes.alphaworks.ibm.com/manyeyes/ ] .) But, these tools notwithstanding, there is always a fundamental obligation to provide for full , rigorous and public validation of data. That is, data must be fit for confident use. +++++++++++++++ Unanticipated uses of resources are one of the most interesting aspects of resource sharing on the Web. (At the American Museum of Natural History, we made a major investment in developing a comprehensive presentation of the American Museum Congo Expedition (1909-1915) – our site included 3-D presentation of stereopticon slides and one of the first documented uses of the site was by a teacher in Amarillo, Texas who was teaching Joseph Conrad – we received a picture of her entire class wearing our 3-D glasses.) It seems highly unlikely to me that we can anticipate or even should try to anticipate all such uses. In the early 1980’s, I taught Boolean searching to students at the University of Washington and I routinely advised against attempts to be overly precise in search formulation – my advice was – and is – to allow the user to be the last term in the search argument. An important corollary to this concept is the notion that metadata creation is a process not an event – and by “process” I mean an iterative, learning process. Clearly some minimally adequate set of descriptive metadata is essential for discovery of data but our applications must also support continuing development of metadata. Social, collaborative tools are ideal for this purpose. (I will not pursue this point here but I believe that a combination of open social tagging and tagging by “qualified” users -- perhaps using applications that can invoke well-formed ontologies – holds pour best hope for comprehensive metadata development.) Continue reading
Latest Comments