Home » Posts tagged 'Data' (Page 7)
Tag Archives: Data
What do we mean by “effective” access to data?
As previously discussed, “free” and “open” dissemination of data are primary values, are fundamental premises for democracy. Data buried behind money walls, or impeded or denied to users by any of a variety of obstacles or “modalities of constraint” (Lawrence Lessig’s phrase) cannot be “effective”. But even when freely and/or openly available data can be essentially useless. So what do we mean by “effective”? One possible definition of “statistics” is: “technology for extracting meaning from data in the context of uncertainty”. In the scientific context – and I have been arguing that all data are or should be treated as “scientific” – if data are to be considered valid, they must be subject to a series of tests respecting the means by which meaning is extracted... By my estimation, these tests in logical order are: Are the data well defined and logically valid within some reasoned context (for example, a scientific investigation – or as evidentiary support for some proposition)? -- Is the methodology for collecting the data well formed (this may include selection of appropriate, equipment, apparatus, recording devices, software)? -- Is the prescribed methodology competently executed? Are the captured data integral and is their integrity well specified? -- To what transformations have primary data been subject? -- Can each stage of transformation be justified in terms of logic, method, competence and integrity? -- Can the lineages and provenances of original data be traced back from a data set in hand? The Science Commons [SEE: “Protocol for Implementing Open Access Data” http://www.sciencecommons.org/projects/publishing/open-access-data-protocol/] envisions a time when “in 20 years, a complex semantic query across tens of thousands of data records across the web might return a result which itself populates a new database” and, later in the protocol, imagines a compilation involving 40,000 data sets. Just the prospect of proper citation for the future “meta-analyst” researcher suggests an overwhelming burden. So, of course, even assuming that individual data sets can be validated in terms of the tests I mention above, how are we to manage this problem of confidence/ assurance of validity in this prospectively super-data-rich environment? (Before proceeding to this question let’s parenthetically ask how these test are being performed today? I believe that they are accomplished through a less than completely rigorous series of “certifications” – most basically, various aspects of the peer review process assure that the suggested tests are satisfied. Within most scientific contexts, research groups or teams of scientists develop research directions and focus on promising problems. The logic of investigation, methodology and competence are scrutinized by team members, academic committees, institutional colleagues (hiring, promotion, and tenure processes), by panels of reviewers – grant review groups, independent review boards, editorial boards -- and ultimately by the scientific community at large after publication. Reviews and citation are the ultimate validations of scientific research. In government, data are to some extent or other "certified by the body of agency responsible.) If we assume a future in which tens of thousands of data sets are available for review and use, how can any scientists proceed with confidence? (My best assumption, at this point, is that such work will proceed with a presumption of confidence – perhaps little else?) Jumping ahead, even in a world where confidence in the validity data can be assured, how can we best assure that valid data are effectively useful? A year ago in Science a group of bio-medical researchers raised the problem of adequate contextualization of data [SEE: I Sim, et al. “Keeping Raw Data in Context”[letter] Science v 323 6 Feb 2009, p713] Specifically, they suggested: “a logical model of clinical study characteristics in which all the data elements are standardized to controlled vocabularies and common ontologies to facilitate cross-study comparison and synthesis.“ While their focus was on clinical studies in the bio-medical realm, the logic of their argument extends to all data. We already have tools available to us that can specify scientific work flows to a very precise degree. [SEE for example: https://kepler-project.org/ ] It seems entirely possible to me that such tools can be used – in combination with well-formed ontologies built by consensus within disciplinary communities to systematize the descriptions of scientific investigation and data transformation. – and moreover – by the combinations with socially collaborative applications -- to support a systematic process of peer review and evaluation of such work flows. OK -- so WHAT ABOUT GOVERNMENT INFORMATION??? We’re just government document librarians or just plain citizens trying to make well-informed decisions about policy? Stay tuned… Continue reading
What is NOT “science”? Why we have a right to “data” as “evidence”…
Most of us accept a priori the institutionalized distinction between the sciences and the humanities. If asked, we can tick off the names of “disciplines” that are “scientific” and those that constitute “the humanities”… (The “social sciences” are somehow less centrally – more vaguely? -- “scientific” -- but what do we mean essentially by these distinctions?) [It's worth noting that novelist CP Snow famously posited this distinction in his Cambridge lecture and subsequent book "The Two Cultures" -- ca. 1959.] We might say that science is “empirical” meaning that it is based upon real, physical evidence? Or perhaps that it’s “inductive” – its theories or “laws” flowing from observations of facts… Or perhaps that it is “quantitative” or "technical" – its conclusions determined by the use of sometimes very complex mathematical logic or by complex apparatus. We might also say that it employs a rigorous methodology that includes exact logical provisions for “falsifiability” [SEE: Karl Popper, The Logic of Scientific Investigation – and elsewhere], for open peer review – including test by replication – and for validation by demonstration of predictive power… Science also is systematically accretive and depends on careful citation and documentation, building upon itself like a coral reef… But, it strikes me that any humanist should feel uncomfortable at the assumption that the humanities do not – or are incapable of – meeting these standards at least most of them in most cases? (I'll leave it to the reader to assess what is most essentially “humanist” – but I often have the uncomfortable sense that the humanities may too often depend for their esoteric authority upon the incoherence of their evidentiary base or upon the imprecision of language or between languages…?) I attribute "beauty" as a primary motive/value to “the arts”… (The American poet Randall Jarrell once said: “Criticism is the poetry of the prosaic.”) And I heard, anecdotally, a few years ago that the performance artist, Laurie Anderson, was invited to a discussion about “the arts” and “the sciences” and before too long was asking “What are we doing here?” I understood this to be an intuitive recognition that the arts and the sciences are on very similar tracks… I believe that artists are able to operate more spontaneously, intuitively and imaginatively -- perhaps more "aesthetically"? but less "systematically" ? Scientists often operate on that same frontier but with the requirement that they test their intuitions using the scientific method and then publicly disclose their “tests”. "Belief" is ultimately the subjective preserve of the individual -- and the institutional preserve of religion. Maintaining the distinction between "belief" and reason (or logic) is a fundamental value of the Enlightenment -- particularly in public discourse. OK so what am I getting at here? And why? Ultimately all policy -- whether "scientific" or not -- and all human decisions should be based on logical analysis and on evidence. Both evidence and analysis are susceptible to testing, to evaluation and thus to reasoned discussion. Our civil discourse will always be improved by clear specification of analytical logic and by free, open and effective disclosure of empirical evidence or DATA. Respecting data there are a series of fundamental criteria that must be satisfied to validate it’s “authenticity” and its probative value (its effectiveness as evidence). As citizens, we have the right to demand that public policy and public decisions be based on well-formed logic and on valid evidence… Discussion that occurs in our public fora should always distinguish between matters of logic and fact and matters of belief. We’ll pursue these notions – in the context of free, open effective access to data and in the context of science literacy – in future posts… Continue reading
Data.gov.uk Launches Soon!
Looks like the UK version of data.gov, developed by Sir Tim Berners-Lee, is going to be released soon. It is "language-based" where "linkages are based on human language, rather than hard-coded hyperlinks", a.k.a. the Semantic Web concept that Berners-Lee has been touting for years. I like the way Nancy Scola of Personal Democracy Forum describes the Semantic Web:
[Berners-Lee] vision is of a web that understands the connections between disparate bits of information in a way similar to how the human mind might effortlessly connect an address on London's Whitehall with the events of World War II that Winston Churchill directed from an underground bunker there. Data woven through with more human ways of interpretation might, just might, make the gap between making government information public and making it useful a little smaller.The BBC reports that "Data.gov.uk is built with semantic web technology, which will enable the data it offers to be drawn together into links and threads as the user searches...we will also be able to look for patterns...visitors to data.gov.uk will want to make their own mash-ups from the information available." Yes, and we should be making mashups from our country's data.gov for our library patrons too! Let's get to it! I'll be working on mine and will show you how it can be done. Continue reading
DataTO.org
Check out DataTO.org, which is similar to data.gov, but users request data sets from the Toronto municipal government. The first phase of the website will allow one to:
...publish a request for data to the community, where members can comment and rate the request. In future iterations of this site, publishers and others will be able to post details of known and existing data sources so that community members can rate them for prioritization. Users will then be able to find data sources that have been published.Kudos to O'Reilly Radar. Continue reading
FEC makes data available in multiple formts
Disclosure Data Catalog, Federal Election Commission
"Each of the files listed here can be downloaded in either csv or xml formats. Each also has a metadata page that describes the information included and the structure of the file itself. There is a pdf version of each file if you need to print the information. You can also subscribe to RSS feeds for each of the files so you're notified whenever new data is available or a change is made."Also see the Commission's Disclosure Data Blog where the FEC will post information about the files and its future plans. And: they say that "you can get help with any questions about the data we're providing here." Continue reading
Latest Comments