This scoping study is an explorative study of how public bodies currently make decision about data privacy in relation to open data publishing, it will be conducted through semi- structured interviews (see Section 3.4.6). The intention was that the interviews could be used to guide the study and establish how LAs currently deal with open data publishing and also, test the waters to see whether there would be sufficient interest in collaborating with the researcher in testing any prototype tool created as part of this work.
LA practitioners face a lot of constraints on their time and therefore, getting practitioners to agree to participate in research is not an easy task. Six of the LAs who responded to the FOI requests in Section 4.2, were contacted again by telephone and/or email to ask if they would be willing to participate in a semi-structured interviews (see Section 3.5.2). Two of these agreed to take part which, in view of the restrictions on practitioners time, was considered a successful outcome. Further, because these interviews were intended to help inform and guide the study, a small sample of interviews was considered sufficient to provide useful insight (de Ruyter and Scholl 1998).
Three interviews were conducted at the LAs offices with five practitioners participating from across the two LAs, two managers and three practitioners who worked with open data. The interviews were taped and transcribed and loaded into NVivo for analysis. The participants consisted of two members of staff from each LA. These interviewees worked in the following capacities; three were IT professionals who work with data and open data; one was a General Manager with responsibility, among other things, for data and FOI requests; and one was the Manager of an Open Data Department. The interviews took place at the offices of the LAs in three sessions. The first session was attended by two interviewees (an IT professional and the General Manager), lasting lasted 1.5 hours. The other two sessions were attended by the other two interviewees separately, lasting
approximately 1 hour each.
As part of the conversation, it was explained that the outcome of these interviews would be used to gain a better understanding of their data publishing practices and what barriers to publication they may have encountered. The interview guide for these interviews can be found in Appendix A, Section A.1.
Prior to the interviews, the participants were sent brief details of the research being conducted, a list of questions and a participant agreement form (see Appendix A, Sections A.2 and A.3 for copies of these).
4.4.1 Coding
The interview transcripts were coded based on Corbin and Strauss (2008) grounded theory model. The first stage of coding in grounded theory is the open coding phase where snippets from the transcripts are grouped and broken down into named codes through examination, conceptualisation and comparing (Corbin and Strauss 2008). The coding was led by the information within the transcripts rather than by the questions asked. This was done deliberately so as not to miss any points made that might have arisen from the conversation flow rather than from a specific question. This exercise resulted in 77 codes being created.
In the second stage, the Axial coding stage (Corbin and Strauss (2008), see Section 3.4.3), these codes were amalgamated into 13 overarching categories by grouping them into related themes. The 13 themes were:
1. Aspirations - this category relates to the aspirations LA has for open data in future and what they would like to achieve;
2. Benefits - this category relates to what the practitioner(s) see as the perceived benefits of open data publishing for the LA;
3. Current processes - relates to how whether open data is currently published and, where applicable, the processes in place for publication;
4. Data management and manipulation - relates to how data is currently managed and/or published;
5. Departments - relates to the department and the team(s) involved in open data publication;
6. History - information relating to the history of open data publishing from the perspec- tive of the LA;
8. Other LA Platforms - information about what other LAs are doing in the open data publishing field;
9. Policy - relates to any formal policies in place;
10. Privacy - this category includes privacy concerns, risks and current approach to privacy;
11. Publication Scheme - relates to releasing information under the FOI publication scheme;
12. Roadmap - what plans are in place for the future of open data publishing within the LA;
13. Transparency - any discussion about transparency and how this relates to open data publishing.
No further coding took place of this data because the coding was not intended to be used for generating theory in the traditional sense of grounded theory, where the third phase would be to carry out selective coding to generate theory (Corbin and Strauss 2008). Rather, the coding exercise was done to gain a better understanding of the underlying messages and make sense of the interviews.
Relating these categories to the questions asked the following findings became evident:
Q1 - Team/Department
One LA stated they originally had 40 people scattered around the authority who were working with data and, as part of the open data initiative within the authority, ”that bringing them together in one space would be good” (P1). However, it also transpired that, of those, only one person worked full time on open data publication with a second assisting on an ad-hoc basis. The other LA did not have a team in place rather they stated that their IT Information Governance Officer was the only person who did any work in this area, and that involved mainly ensuring that the LA’s obligation under the FOI publication scheme were met.
Q2 - Policy
Neither LA had any formal policies in place for publishing data in open format, one practitioner stated: ”it is easy enough to say, yes, lets adopt a policy but then you find nothing has happened” (P3). The reason for this may be that getting agreement on what should be included and creating formal policies in this area might be considered too difficult, as eluded to by another practitioner who stated they decided to: ”just get going on releasing data but not get bogged down and do anything that was too hard” (P1).
Q3 - Process
When asked about processes one practitioner stated: ”we rely on the process maps...and they are quite specific for creating/updating particular types of datasets” (P2). Thus, through the interviews it became quite apparent that few formal processes were in place for what data to publish or indeed how to publish open data, one practitioner stated: ”it has been more driven by trying to reduce FOI request numbers rather than trying to be proactive with the data....we sort of free wheeled” (P4). This is perhaps not surprising given that no policies are in place for this area either. When asked why they thought there were so few formal processes in place, one practitioner stated it would be good to ”put it into a good standard, keep it updated and take ownership responsibility for it. That’s the difficult bit” (P2).
Q4 - Data
From the interviews it became evident that one of the LAs practice was not to publish until pressure dictates otherwise: ”why do we have to do anything when we can get away with the bare minimum?” (P3). In contrast, the other LA published open data regularly with performance indicators (processed statistical information only) representing around 80 percent of output, ”we are looking at a massive number of indicators, 1000’s and 1000’s of indicators” (P2) . The publishing of these indicators was predominantly automated. When asked how many raw datasets have been published, the practitioner stated: ”outside of the performance stuff, I think if you got up to 100, you would be doing fantastically well...its 80 percent of our time that is spent on that 20 percent of the data” (P5). Therefore, although the LA publishes data in open format, most of this is performance indicators produced as part of reporting processes for different areas of the business. When asked what the reason for this low number of raw datasets being published when compared to performance indicators, the reply was: ”It’s economy of scale” (P1).
Q5 - Problems/obstacles
There were many barriers raised for why data could not be published in open format. These included:
• Lack of resource - ”the issue has always been the cost of releasing that data and the difficulty in doing it on a repetitive basis” (P1);
• Culture - ”When you are trying to get service areas to try and publish data that they would much rather be kept secret, just because it is culture” (P4);
• Technical challenges - ”Data is spread across a lot of different systems...it is difficult in some cases to get that data out of the system” (P5), ”It’s like an algorithm that
fights back at you” (P2);
• Lack of collaboration - ”Not everyone wants to collaborate” (P3);
• Management buy in - ”what we have got to do is get buy in and the attitude is, well, what’s in it for us?” (P3);
• Meeting user expectations - ”I don’t think we could ever really meet the ambition of the market and what they would want us to put out” (P4).
Q6 - Standard
Although the practitioners interviewed all agreed that having a standard for open data publication would be useful, they felt that the technical barriers such as legacy IT systems and the difficulties in getting data output from these in consistent formats would prevent this from being a realistic goal for now. However, they would all be happy to participate in developing such a standard if a group was set up to achieve this. For the moment, the best hope, according to one practitioner, would be to create standardised schemas: ”there are attempts towards standardising the actual schemas which will be absolutely brilliant. It is never going to work 100 percent but even if it works 50 percent it will still be brilliant” (P3).
4.4.2 Other findings from interviews
When asked about privacy, practitioner opinions ranged from cautious: ”some information, particularly datasets containing sensitive personal data, will clearly present a need for caution, and the anonymisation issues may be complex for large datasets containing a wide range of personal data” (P4)” to unconcerned: ”I am not too concerned about the privacy angle because we would never put out personalised data.” (P1), and very concerned: ”I am almost convinced that if I went back through our data that we have published over the last 4-5 years, I would find something that we’d missed” (P2). This may explain why only 6 of the LAs contacted as part of the FOIR currently have an open data portal.
Where public open data is published, innovative uses have been made of the data. For example, OpenStreetMap used public open data to provide users with free access to maps of their local area (Open Data Institute 2015). However, while the LAs interviewed appreciated and appeared to endorse government’s eagerness to promote open data publication as a means of increasing and promoting transparency (see Section 2.3), e.g. ”by opening up information to people you can foster growth” (P1), most do not practice what they preach.
These findings suggest that there is great disparity in how each of these LAs approach open data publishing. One of the LAs confirm the finding from the two previous studies
into LAs open data publishing approaches that, unless the information is requested, it is unlikely to be made available (see Sections 4.2 and 4.3), the other LA interviewed however, do publish open data proactively. However, looking at the findings from the Boswarva dataset, only six of the LAs have published in excess of 50 datasets on the central data.gov.uk site (see Figure 4.3), suggesting this LA is in the minority.