Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Monday, January 23, 2012

"Open Data" Needs to Die

Amongst all the UK GovCamp 2012 buzz, point #18 from Tom Sprints' write-up caught me as being one of the more curious:
18. A lot of “open data” sessions just seemed to me to be variations on a theme, and didn’t sell themselves to me at all. I am therefore worried that some of those discussions are either very esoteric, or insufficiently informed by people who understand the issues rather than the tech.

Where have we come from?

As a data geek (I like the word "mechanic" myself), it's been intriguing to see the conversation around "open data" change over successive GovCamps. A few years back, the question was heartily "How can we get hold of data?" - Tim Berners-Lee was starting out on his comeback tour, and mySociety were beginning to show that data could be made useful with some clever tools.

As I remember it (likely in a fairly biased narrative way), the conversation then switched fairly rapidly into "What's the best way to open up data?" - in terms of what data and what platforms were most useful to developers. Suddenly data stores had (experimental) APIs, and the public realm had massive amounts of spending data. There was some loose rhetoric about transparency and accountability, while developers picked things apart with fine Excel toothcombs.

Then things got more interesting, as it turned out everything that had happened so far didn't automagically lead to Amazing Stuff Happening. The question became a necessary "So what?" - as if transparency and accountability weren't enough by themselves! The topic turned to users and reasons and (more often) to interesting examples. Surely, somebody was clamouring for this stuff after all this?

I'm kind of hoping this explains something about why "open data" sessions are a bit fumbly-jumbly now.

Open data got complicated, quickly. Because data is complicated. Jump to the present, and conversations rapidly flit between all of the above either because everybody is involved at the same time, or the people who should be involved, aren't.

"Open Data" is harmful

Or both. The paradox is that it's become difficult to talk about open data firstly because those who were talking about it from one point of view are now talking about it from many points of view. And secondly because those who weren't talking about it before aren't talking about it now. Data silos still exist. Most people still use Excel. Statisticians still output reports.

The term "open data" is meaningless now. Not just meaningless - actively harmful. If you're used to talking about it, then the conversation has begun to fragment and coalesce around more subtle outcrops. And if you're not used to talking about it, then you're put off because nobody can explain what it means - and more importantly, what it means to you. So you carry on as normal.

My session at GovCamp on Data Engagement was, in retrospect, an attempt to get back to the previous question of "So what?". What I really want to do is fence the conversation off from the technical, economic and political aspects of data (although I'm still into all these things) and focus on the why. I desperately tried not to use the term "open data" because I think it would have distracted the discussion. (To be honest, I wanted to find something better than "data engagement" too, hence the phrase "Everyday data".)

And I'm really glad that some of the idea got taken up on day 2 by Tim Davies and others. A "Charter" for engaging with data really starts to delve into how we think about how to make data useful.

I admit I'm a little afraid that the term "Open Data Engagement" just makes the discussion even more vague. What does that mean to you if you have no idea what it is, or what Open Data is supposed to be? Is it all at risk of becoming another buzzword? What about "Data Usability", or "Public Data Engagement"? I'm still aware just how much I hate the terms "Public Understanding of Science" and "Public Engagement with Science". Are we going round in circles?

Should we call a Stats Spade a Stats Spade?

Many people with useful, everyday data and databases really don't think in terms of data. Because the data is about stuff they know, they think of it as "information". Maybe even a "resource". But ask them what "data" they have and they'll probably give you a back-up of their website.

One of the interesting points coming out of the Data Engagement session was that people deal with data all the time - think football, Formula 1, house prices, etc. But do people even refer to this as "data"? Or - more likely - do they call them "stats"? Mention "stats" and people think of tables, averages, and counts.

In a way, "stats" makes sense where "data" doesn't. "Information" makes sense where "data" doesn't. "Data" is tricky because it's all of this and more. It's figures, it's formats, it's visualisations. No wonder even those who understand this get confused when talking to each other. The more you try to take "Data" into the real world, the less the term applies.

Should the "open data" moniker be scrapped instead of more "useful" terms like these? Would this make talking about implementing it more difficult, or easier? After all, any conversation on how to make data useful quickly turns away from talk of even databases and on to other issues (standards, protocols, best practice, comprehension).

Maybe if we talk about our bus times as "public information", and spending figures as "spending figures" then people will be interested in it, and we can stop trying to work out what "open" means.

Maybe.

Sunday, January 23, 2011

#UKGC11: Hooking into DIY Data

Following on from data relativity, the second thing that grabbed me at #ukgc11 (UKGovCamp 2011) happened during Will Perrin’s introduction to “Making a Difference with Data”, and followed up in conversation with Helen Jeffrey (@imhelenj) on community-led data. Will talked of comparing his local authorities’ data on lamp-post repair time to his own count for how long a lamp-post had been broken for (encouraging crime). Helen talked about data that a group of volunteers had collected themselves, and then turned into a report to feed into government decision-making.

In both cases, what struck me was that people don’t have such an aversion or fear of data that is often assumed - if they start by generating it themselves.

To go back to data relativity, the most confusing and scary part of data is figuring out the thought processes and assumptions that have gone into a dataset, as well as figuring out what the hell’s important (often around 0.01%-0.5% of the data) and what’s “noise” - to an individual. Self-generated data doesn’t suffer from this, because all the scary background and assumption bits are part of the citizen’s/volunteer’s mindset and experience. Voila, understanding data comes from experience. And as such, successful engagement with data is about creation as well as consumption.

OK, it’s clearly a little more complicated than that, but it’s a principle that’s often far too implicit for some datasets (crowdsourced maps, etc), and far too often forgotten about for others (most centrally-gathered stats). There are also a whole bunch of stereotypes about what “people” “want” from “data”, and often these stereotypes do little except re-establish the status quo. When data comes up against real-world users (yes, even geeks), and the magic “fails to happen”, we’re left wondering if natural engagement is such a given after all.

It’s difficult to get excited about thousands of datasets when you have no idea where to start or where they can be relevant to you. It’s much, much easier to get excited about data that is relevant to to you, that you understand, and that you can see how it will benefit you. I think that’s why I love the idea of Mappiness or Christian Nold’s maps - both involve getting people to create data they find interesting, and through this “personal” data, the need and relevance for other data suddenly becomes instant and appreciated. (For example, if I’m feeling most happy on street X, what other properties does that street have, and how do they relate to me - house prices? Pollution? Streetworks?)

This seems to be an issue that is bubbling along - everyone knows that organisations collect data, for instance, so an open data system worth its salt will take that into account. But there are still assumptions about the scale and complexity of that data, I would argue - whereas really, data can be as simple as counting, and everyone can count.

Thursday, November 25, 2010

Structured data: Accessible magic?

Magic Numbers (source)

Have you ever dreamt in SQL? Few of us have, and we only ever speak of it in hushed, yet secretly astounded tones. Database development is a weird way of looking at the world. Those who venture into it too deeply may never come back to being "normal".

But at the same time, it's also like any language - reaching a state of fluidity involves a fundamental shift in thinking. To think in French is to adopt a new philosophy, and the same is true of database logic; linking and manipulating rows of data is a world away from editing each row by hand.

To explain the importance of structured data, is it important to first get across this conceptual paradigm shift? Is the ultimate draw of structured data tied inherently to a new way of seeing the world? One in which we, as data/content hackers don't see the data at all, but merely instruct a computer to do stuff with it.

Maybe this explains the "format divide" between those publishing in "closed" formats, and those publishing in open ones. If you do not know how to automate the data-munging process, then you do stuff by hand, you take a long time to do it, and you have absolutely no need for "structures" other than those in your head.

This happens everywhere, all the time: half the world lives in Excel during office hours. At some point, computers became popular as difference engines, but not necessarily good at being them. Human operators became part of the machines, rather than directors of them. In a way, this mirrored the huge factory production lines, and the endless supermarket checkouts, so most humans simply accepted this as the new way of life. Any sufficiently different technology is a form of magic. Processing as a manual task.

This happens today. Never forget that. Seeing data is more important than defining a structure for it, because structure is *hard*. Datasets have peculiarities, errors, and specifics that resist simple structuring. And changing these structures is effort - effort that involves communicating these changes to other. In short, it's easier to stick with "sloppy" data if you're not using the right tools. It's even easier if the other people using the data don't care about the tools either. Content is King.

So how do we bridge this divide between "manual data labour" and "magic"? On the up side, I believe it must happen, as data - and talk of data - becomes a public matter. Those not structuring their data will need to structure it, or face a new kind of exclusion - call it "un-APIness" perhaps.

But this doesn't help us to move into a culture of automation, of magic, which I think is important because it determines what you *believe* you can do with data. Understanding structured data is essential in coming up with new services, new applications, and new answers.

Working with people to build answers will help too. It's not enough to just want "raw data now" - to build bridges, we need to build real things based on data. We, as geeks, need to find out what people actually want. We need to show that questions can be answered with "magic", but also be open enough to demonstrate that structuring data has a direct impact on what can be done, and how quickly.

Change the tools. Rethink processes. It's time to end the conveyor-belt, factory line approach to data.

Wednesday, July 01, 2009

Transparency should not be for Blame

I was going to write a small blog post, but Peter Kawalek's discourse on what isn't said around MP expenses says what I was going to say, and far, far more elegantly.

I'm encouraged by the flurry of interest and activity surrounding the release of expenses data, but at the same time I can't help but question whether it really matters after all that.

Did people really care when this all hit the headlines? I, for one, got the impression that the weeks of blathering waffle on the radio and in the papers was being drummed up and forced on to stage by either politicians wanting to embarrass other politicians, politicians wanting to un-embarrass themselves, or media outlets looking to embarrass politicians - which, incidentally, is like taking sweets from a baby in a sweet shop.

For everyone else, talk of expenses was dull, dull, dull, and generally a good excuse to flick channel, turn off the radio, or go and do something interesting like make pasta shakers.

The point of this rant is this then: Is it important to spend time and energy releasing the kind of data that, while ideas of transparency might be in the public eye, doesn't actually either a) contribute to our understanding of the state of things, or b) offer a positive solution?

After all, the main reason for releasing expenses data is to find people to point fingers at, rather than to actually applaud MPs for not spending money. (Personally, I'm thinking of sending my MP some better coffee than the Kenco stuff he orders...)

Data can be good. Transparency can be good. But shouldn't we be careful that we're not just opening up an attitude of blame culture? Can we avoid a society transparency and monitoring are no better than CCTV or a nanny state - a culture of wrist-slapping people for their mistakes, rather than encouraging and rewarding valued behaviour?