Showing posts with label opendata. Show all posts
Showing posts with label opendata. Show all posts

Wednesday, July 13, 2011

Open data funding - experiments and ecosystems

Paying for the Open

The funding aspect of open data development came up at Open Data Brighton & Hove (#odbh) last night - who should (or shouldn't) pay for it?

While one camp says that there are lots of people who will build on top of open data for free and for passion, the camp at the other end of the hall wants to see return on investment for work paid for. The latter works both ways - people want to be paid to develop, and people want to pay for development. If the payback is enough, of course.

In a sense, both camps are "right" - the model you believe in depends on your daily interests, daily funding models, and where else you get money from. So it's easy to see that some people are fine building free side-projects, while for others it's a day job. Sometimes one person may have a foot on both sides, depending on what's going on that particular day/week/whatever.

This will always be the case. So it's really really important to understand that there is no "correct" model. Any open data ecosystem needs to fundamentally take this into account. Making data available is great - some people will run and play with it. But working out funding and collaboration is also great. Both are essential, even in the context of open-source, cutbacks, austerity and liberal progressiveness etc etc.

Bountiful

The more I think about #odbh, the more I notice how much I'm influenced by the openness of the Bitcoin community. Other open communities exist, of course, and do similar things, I'm sure, but Bitcoin is the one I'm closest to at the moment.

(Background sidenote: Ignore what Bitcoin is, and whether it's a good idea or not. The relevant and important point is how people are organising around it.)

One funding model that seems to work is the "Bounties" model - a kind of funding pledge, but one based on identifying desired functionality rather than, say, group activity or a band's next output. This list of bounties isn't complete, but it illustrates how it works and the kind of work people want done.

Could this work for open data development? If people are serious about wanting an idea turned into reality, shouldn't they put their wallet where their mouth is? Does it offer a "third way" to both working for free or having to "prove" your idea in advance?

I suppose what I'd envision is a bit like the Ideas section on data.gov.uk, but with more ... oomph, more "I really want this" instead of "This'd be nice".

So...

To wrap up, what this says to me is that open data is more than just about getting data out there, and even more than just about how we weave data into our everyday lives. It's about how we commission progress, how we organise collaboration, and how we identify needs.

All stuff we've been doing for ages, really. But here's a chance to try new ways of approaching old problems, and to bring all of that experimentation together. To create a very real open ecosystem.

Monday, February 14, 2011

Hackers, transparency, and the zen of failure


This week I picked up the Hacker Ethic from my library (remember them?) by Pekka Himanen, with an intro by Linus Torvalds and an outro by Manuel Castells. It was hard to resist it with names like that splashed across it.

The book was written in 2001, back when the browser wars were in full swing and streamed video was still a bit of a novelty (so nothing's changed that much). But the theme addressed by Himanen is anything but dated - and contains some key threads which I want to think about and blog about more.

A lot of the text deals with the idea of what makes a "hacker" tick -and not just geek hackers, but anyone with a passion for what they do, rather than a bitter feeling that they work because they have to. Central to this hackerness is a passionate creativity, and a desire to share knowledge - including the results of that knowledge, such as code.

The truth and untruth of progress

This act of sharing knowledge is partially a form of status, true. But when you read articles like this about incorrect data being published, you start to notice what else open knowledge (including data) is about - social learning.

Hackers and open government are both (now) keen on sharing data - knowledge, code, ideas. But the real difference is in how they learn - for the hacker, openness brings about learning and improvement through public failure - there is an assumption that what you create can be improved, and an attitude that anyone else is welcome to improve it.

Now compare this to CLG's response to the LGC article above - a response filled with defensive language and finger-pointing. There is something rather scientific - or, rather, legal - about this discourse: claims are made by one party and refuted by another. Slowly the "truth" is "sculpted" from what is left.

But for the hacker, the truth is only what is created - not what is undisputed. Hackers fork code, create new communities, start new websites, run unconferences. If "truth" exists, then it is what emerges, not what is discovered, or what remains.

What do hackers sit on?

Can a highly hierarchical structure such as our democracy adapt to be creative rather than competitive? The open data movement is driven by both of these - data for transparency can be thought of as "evidence" in a legal bid for the justification of an organisation's existence. Data for new apps, on the other hand, only needs a use, and to be useful in a creative context.

This split in attitude is key when considering efforts like the recent consultation on local data transparency, which clearly puts "open data" into the evidential context:

The Government wants to place more power into people’s hands to increase transparency by seeing how their money is spent.

Transparent screen 1
img by AMagill
My fear is that this inherently makes "open data" unuseful to the hacker crowd - an essential crowd to interest when the data is being released in CSV files, or other formats that require some parsing. If hacker's can't create something with the data, they won't do anything with it. The idea of an "army of armchair auditors" becomes a functional paradox, as the people the Government has in mind for the data apparently sit in armchairs, while the hackers sit in cafes, meet in pubs, and generally find comfy chairs far too comfy to code in.

To return to this post's title, what role will failure and learning play in this paradox? Looking at the draft code, we can see a desire to use the "many eyes" approach to fixing data:

18. Data should be as accurate as possible at first publication. While errors may occur the publication of information should not be unduly delayed to rectify mistakes. Instead, publication and use of the data should be used to help address any imperfections and deficiencies.

The hacker approach agrees with this - fix things as we go along. But does this fit with the idea of "armchair auditors"? As we saw in the LGC article, how can an auditor tell the difference between what is incorrect, and what is wholely disagreeable? And if they can't, why should they trust any of it?

(Maybe we need a "stable release" system like that of open source projects? Maybe, like the Linux kernel or desktop distributions, data can be released with an "unstable/testing" tag, then marked up as "stable/trustable" after enough testing has been done on it.)

Transparency is a lovely thing - but everyone has different uses for it. If it's used for creativity, then there is, perhaps, an implicit assumption that things can and will change. If, on the other hand it's to be used for accountability, then there needs to be trust in it.

Sunday, January 23, 2011

#UKGC11: Hooking into DIY Data

Following on from data relativity, the second thing that grabbed me at #ukgc11 (UKGovCamp 2011) happened during Will Perrin’s introduction to “Making a Difference with Data”, and followed up in conversation with Helen Jeffrey (@imhelenj) on community-led data. Will talked of comparing his local authorities’ data on lamp-post repair time to his own count for how long a lamp-post had been broken for (encouraging crime). Helen talked about data that a group of volunteers had collected themselves, and then turned into a report to feed into government decision-making.

In both cases, what struck me was that people don’t have such an aversion or fear of data that is often assumed - if they start by generating it themselves.

To go back to data relativity, the most confusing and scary part of data is figuring out the thought processes and assumptions that have gone into a dataset, as well as figuring out what the hell’s important (often around 0.01%-0.5% of the data) and what’s “noise” - to an individual. Self-generated data doesn’t suffer from this, because all the scary background and assumption bits are part of the citizen’s/volunteer’s mindset and experience. Voila, understanding data comes from experience. And as such, successful engagement with data is about creation as well as consumption.

OK, it’s clearly a little more complicated than that, but it’s a principle that’s often far too implicit for some datasets (crowdsourced maps, etc), and far too often forgotten about for others (most centrally-gathered stats). There are also a whole bunch of stereotypes about what “people” “want” from “data”, and often these stereotypes do little except re-establish the status quo. When data comes up against real-world users (yes, even geeks), and the magic “fails to happen”, we’re left wondering if natural engagement is such a given after all.

It’s difficult to get excited about thousands of datasets when you have no idea where to start or where they can be relevant to you. It’s much, much easier to get excited about data that is relevant to to you, that you understand, and that you can see how it will benefit you. I think that’s why I love the idea of Mappiness or Christian Nold’s maps - both involve getting people to create data they find interesting, and through this “personal” data, the need and relevance for other data suddenly becomes instant and appreciated. (For example, if I’m feeling most happy on street X, what other properties does that street have, and how do they relate to me - house prices? Pollution? Streetworks?)

This seems to be an issue that is bubbling along - everyone knows that organisations collect data, for instance, so an open data system worth its salt will take that into account. But there are still assumptions about the scale and complexity of that data, I would argue - whereas really, data can be as simple as counting, and everyone can count.