skip to main |
skip to sidebar
We publish a monthly e-newsletter for which I am responsible (content, creation, distribution, data management, subscriber satisfaction / evaluation). I have done quite a bit to improve on each of the first four aspects and now need to turn my attention to the last. Apparantly, we conducted a survey almost two years ago now so I think that another one is about due. On top of that, I’m interested to know what subscribers think of the changes or whether they have even noticed them. Some of the changes they are not likely to have noticed (such as the changes to the template which ensure that every issue conforms to XHTML, WCAG and CSS guidelines - all links open in new window), some they won’t know about (such as the new software we use to manage subscriber data and e-newsletter distribution) and others they almost certainly will have noticed (such as the new look our e-newsletter has received).
I haven't done a whole lot of thinking on this one yet but I can immediately see three ways forward:
- The first is to put together a few questions in an email to which we ask subscribers to reply. This is probably the cheapest option but:
- it’s probably the least professional and slick looking to the subscribers
- it’s more onerous on the subscribers
- and certainly creates a fair old bit of work for me in compiling the responses for analysis.
- The second is to outsource the entire thing to a company specialising in this sort of work. This is likely to yield the most professional-looking result for the subscribers and would be the least onerous for me in terms of data compilation and analysis but is almost certainly going to be the most expensive option.
- The third option is to use something like surveymonkey.com (opens in new window) which would enable us to create our own online surveys and does the data complilation and some basic analysis for you. It’s the least flexible of the options (you are working within the confines of an OTS product) but it probably offers the best combination of professional appearance, simplicity for subscribers, simplicity for me, and cost.
I have had a go at putting something together use surveymonkey.com (there is a free version available but I think that we would want to use one of the paid subscription versions as the benefits of doing so are in our favour – e.g. limitations on the number of forms submitted are raised or removed depending on which package you opt for) and am happy enough with it.
Has anyone out there done a survey like this before using a tool like surveymonkey.com? What software did you use? Any ‘top tips’ I might use?
The Business Support team (of which I am a member) has been working on the last few aspects of a project to introduce a workflow management system that we refer to as the Research Management System (RMS). When I started looking at how we might apply a metadata schema to our publications and how we might store that metadata (database vs. embedded), the RMS seemed like the most sensible way of capturing the metadata and certainly presented a way of storing it as well (though this isn’t my preferred option – more on that later).
Creating a schema
I looked around at other organisations and tried to find a schema that we might adopt but they were either too tailored to their current applications or were too general. My conclusion was that it was best to create something that could be adopted by the industry as a whole but that would certainly meet R&D’s needs. I have already written a bit about this work but as a reminder, I basically assembled a series of elements from the Dublin Core and the e-Government Metadata Standard and then created a few that were specific either to the industry (e.g. Asset Type) or to R&D (e.g. Research Objective). For as many elements as possible, I have used existing, internationally-recognised encoding schemes (e.g. W3C’s Date-Time format) and for the R&D-specific ones, I have used schemes developed through a consultative process with our Heads of Research Section and a selection of Research Managers (e.g. Audience Group). I have created the Application Profile though I have yet to create the XML definitions and publish them on our website, though this is ultimately my objective.
Applying the metadata
The RMS is due to go live next week and the metadata schema, along with the controlled vocabularies that I have had to create to support some of the R&D- and industry-specific fields, will be put to the test. Without going into too much detail here, we have worked as many of the elements into the process flow as possible so that as our Research Managers work through a project and record it in the RMS, some of the data that they enter is held in metadata fields for later application to any publications that emerge from the research. Clearly, not all of the elements can be populated this way (title for example can only be completed once the publication has been completed) but many can (such as research topic).
At the end of the research process, there is a knowledge management stage where the method for publishing, promoting and evaluating the publication is captured and the remaining metadata elements are completed. Some of these are set to defaults that will almost certainly not need to change, such as the publisher (that is pretty much always going to be our organisation). It is all looking like it should work but of course there is really only one way to find out for sure.
"Subjectlessness"
My one disappointment in this project was the inability to sort out a controlled vocabulary for subject element in time for the launch of the RMS. The problem I encountered is that existing controlled vocabularies are either too granular (for example, SELCAT (http://www.levelcrossing.net/) have a highly detailed thesaurus on the topic of Level Crossings, just one of the many areas of research that we pursue) or insufficiently granular (the Integrated Public Sector Vocabulary, or IPSV, directs users to categorise anything having to do with the railways, from electrification to passenger crowding, under Rail transport). Developing one of our own is just too big a task to try to sort out in only a few weeks (at the same time as all of the other projects that I am working on) so for the time being, the field in the RMS will be populated with: [IPSV] Rail transport. Although this is largely useless to us, it helps us tie in our work with that of the Department for Transport and other government bodies applying the IPSV.
Towards completion and "subjectfulness"?
The next stage of this project for me will have three aspects:
- The first is to refine the schema and existing vocabularies (Do we have all of the elements that we need? Are there any that have been included that just aren’t necessary? Are the controlled vocabularies: sufficiently granular? too granular? incomplete?)
- The second will be to create a final version of the application profile and to publish the element and encoding scheme definitions that I have had to create to accommodate some of the metadata that is specific to our requirements.
- The third and final aspect is to resolve this issue of subject headings. I am attending a CILIP workshop tomorrow called “How to construct a thesaurus” which I am hoping will give me some ideas and strategies for solving this problem. I would like to use existing vocabularies (to build in as much interoperability as possible) where possible. Maybe the solution will be to use things like the SELCAT vocabulary and the rolling stock manufacturers’ parts vocabularies but to only go down to a particular level within them (not sure what the implications of doing so would be, yet).
Metadata in the database vs. metadata in the document
I have already written about this little bug-bear of mine but it is still an issue for me. At the moment, I am going along with a database-held metadata solution but this is largely due to the presence of this option and the distinct lack of any alternatives. I think that once the metadata schema is relatively set and the encoding schemes in use, that I will turn my attention to resolving the issue of how we embed the metadata into the documents themselves...
We have now created and agreed a working model (for lack of a better term) of a metadata schema. This schema has been integrated into our workflow software and it is this way that the metadata will be gathered. Clearly, as we use it, tweaks will need to be made and the workflow software is flexible enough for us to do so.
We are also trying to introduce XML content-level metadata and to date, we have managed to get a series of templates agreed. These templates are tightly controlled in terms of their structure so introducing XML from a structural perspective should be straight forward and we’re evaluating a couple of tools for ensuring that the XML tags are correctly applied.
Software Challenges
The biggest problem that we’re having with the software that will enable the move from templates in MS Word to XML encoded content is the fact that much of the writing is fulfilled by external suppliers and this raises licensing problems. They aren’t insurmountable but it will require a flexible software vendor, disciplined suppliers and fair bit of negotiation to agree something.
Document- and Content-level metadata: relationship
So if the XML metadata is focussed mainly on document structure (e.g. this content is the Introduction, this content is the Methodology, etc.) how does this relate to the document-level metadata which looks at subject, relational and bibliographic aspects of the document? The two are mutually exclusive to some extent but how relevant is the document-level metadata to the content? Should it be captured and accompany the content? I don’t really know the answers to this these questions and they’re the easier ones!
At the moment, I think that the best solution would be to include within the content-level metadata a reference to the document(s) of which it forms part. Someone could then move from content-level to document-level metadata if they wanted to see subject, relational, or bibliographic data. Of course, the minute you reuse some content, say a “Findings” paragraph in an “Introduction” paragraph, that structural element changes. So the structural element needs to exist within the context of a document of origin / reuse. But wait, because here is where it starts getting really tricky…
If we want to assign subject metadata at the content-level will we have to double our work? I can’t think of any other way...the subject of a particular paragraph will not be the same as that of the document as a whole.
Also, how do we manage bibliographic metadata at the content-level? For example, the first time a paragraph is written (probably as part of a larger document), the author associated with the paragraph and the document are one and the same and is probably pretty easy to establish. What do we do when the paragraph is combined with paragraphs that are also taken form other documents and ones that are new? I think that one can argue that the author of the individual paragraphs is clear but who is the author of the document?
Conclusion (resignedly)
The more I think about it, the more I think that we are just going to have to manage two levels of metadata. At the point of creation, the content- and document- level metadata are the same but as content is reused, two distinct and different levels of metadata emerge. To be honest, it would be a big step forward if we were to introduce structural metadata at the content-level and as this presents the fewest or simplest (not simple, mind you, simplest) challenges, I think we will pursue this objective and reassess where we go from there.
Currently, all of the publications that our Research and Development team publish carry a copyright statement that allows readers to reproduce the publication free of charge as long as it’s for research, private study or internal circulation within an organisation. It goes on to stipulate that the subject needs to be reproduced and referenced accurately and that it must acknowledge the company before finishing by directing enquiries for any other uses to the Head of R&D.
I’d quite like this copyright notice to be on the web as well and it is proving quite a struggle to get it hosted there. Within the metadata structure that we are implementing, there is a copyright element and I would like this to have to contain nothing more than a URL but there is resistance and I can’t get anyone to articulate why – I think that it just hasn’t been done before and the reason for doing so isn’t sufficiently obvious. One option for getting around this point is to go down the route of a Creative Commons (CC) licence . Doing so would enable us to mark our publications, record our CC licence in the metadata and be able to point to it on the Web. The question is, though, what is the difference between using a CC licence to protect our IP and the copyright statement that we currently use? Surely what I have described above would be defined as ‘some rights reserved’ rather than ‘all rights reserved’…? In fact, a quick bit of surfing leads me to believe that what we are after is a ‘Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 License’.
I think that a little more investigation is probably required and probably a quick call to our solicitors to verify things before we commit to making this change but to be honest, the threat of pushing for a CC licence might be so much change that ‘they’ concede on the hosting of the existing notice and that would at least get me where I want to be.
Having done some research, it looks like there isn’t yet a metadata standard that is either designed for use or in common use by the rail industry. Although it wasn’t a surprise, it was a bit of a set back as it meant having to assemble a proposed metadata schema for the description of materials within the rail industry, engage in consultation starting with the R&D senior managers, and testing to validate it.
Building the metadata standard: the birth of RIMS
I have chosen to draw from the Dublin Core Metadata Standard (DCMS), and the Dublin Core Terms (DCTerms) to flesh it out a little. I have also drawn from the e-Government Metadata Standard (eGMS) to ensure that any public sector aspects are also covered. Now, the eGMS is based on the DMS and so far, what I have described would pretty much describe the eGMS. The rail industry, and our R&D work in particular, requires a little more granularity than the eGMS currently provides, something the Cabinet Office acknowledges by encouraging enhancement and refinement to suit different contexts. So I have also added some additional elements that are specific to the rail industry (e.g. asset type) and to our organisation (research topic). The complete picture is what we will consider to be the Rail Industry Metadata Standard (RIMS).
Identifying, assembling and building controlled vocabularies
Once I had a proposed set of elements, I set about sorting out the required controlled vocabularies. In my, albeit relatively limited, experience, this is the most difficult part. For many of the fields drawn from established standards, it was pretty straight forward (e.g. date formats us the W3C-recomended date-time format). For some of the elements that I had to create (e.g. research topic), it was also pretty straight forward because such lists were specific to the company and, in many cases, already in current use. Others from both established standards and the new set, however, were much more difficult. One such example is the asset type element – how granular do you go? For most of us, the term ‘locomotive’ is sufficiently descriptive but for our engineers it’s just too broad. My approach to these controlled vocabularies has been to put together a starting point and seek input and comments. So far, I have only engaged the R&D team and the lists have been heavily refined and accepted by them.
Subject
My other challenge has been sorting out a subject matter controlled vocabulary and it is proving to be a somewhat daunting task. The Integrated Public Service Vocabulary (IPSV), recommended as the controlled vocabulary for DCMS Subject, treats everything to do with the rail industry as ‘Rail Transport’. Clearly, this isn’t going to be sufficient for our requirements. I started to have a go at this task in the same way as I approached sorting out some of the other controlled vocabularies but it has proven to be too big. At the moment, it’s on hold while I move the rest of the project forward with a space reserved for subject tags and start to look for other initiatives both here and around Europe that are working towards creating a controlled vocabulary of some sort for the rail industry.
Handling the metadata
There are basically two different ways of managing document metadata: you can hold the metadata in a table which includes the location of the document described and then use this table to search and retrieve documents or you can embed the metadata into the documents themselves and search that (In reality, the search software or engine will most likely create its own table of metadata as in the case of the first method but this is a temporary table that is understood to need regular updating so is not the source of the metadata). Each method has its strengths and weaknesses (e.g. the table is quicker and simpler to deliver while the embedded data means that when someone downloads the document to a local space, the metadata travels with it and isn’t lost).
At the moment, we are also in the process of introducing a business process management system (we are calling it the Research Management System or RMS). The RMS will allow us to store documents as well as manage their production and approval. As a result, it makes sense that we piggy back the metadata assignment on the RMS work meaning that we will be going down the table route. This isn’t my preferred option but it is the one that will mean that we get metadata gathered and stored sooner. Once that process in embedded, we can look at technologies that will enable us to embed that gathered metadata into the files so that users downloading them from our website take the metadata with them.
One challenge that remains, and for which we have a few options but haven’t decided on any one yet, is what we do with the legacy collection. It has been decided that past projects and their associated documents will not be uploaded into the RMS. So the RMS presents us with the solution for future publications but it doesn’t deal with the existing collection. It is most likely that we will upload the previous publications to a separated segment of the RMS which will store the metadata in the same way but there are a couple of alternatives solutions…more on this as it develops.
Where from here?
The next thing to do with this standard is to confirm that it works in practice which will be part of the embedding process for the RMS. We will then look to consult the rest of the organisation on the suitability of the metadata schema and its associated controlled vocabularies for wider use in the company. I guess you could think of our work as a bit of a pilot for the rest of the company.
I’d like to publish our schema and controlled vocabularies under creative commons and invite other organisations to comment on it or use it in their organisations.
Looking through my RSS (opens in new window) feeds this morning, I came across an article by someone with the same name as me. Upon closer inspection, it turned out to be an article by me! I had been told that the article that I had submitted sometime ago (September or so) would be published either as part of the December or January issue of Update Magazine (opens in new window), CILIP’s monthly magazine for members. Okay, so it isn’t the Guardian but I’m pretty pleased all the same.
The December issue is currently featured as part of the main Update page: http://www.cilip.org.uk/publications/updatemagazine (opens in new window)
As a more future-proof link to the article, though, here is its home in the archives:
http://www.cilip.org.uk/publications/updatemagazine/archive/archive2006/december/Bruce.htm (opens in new window)
As part of the Industry Schema metadata work that I am doing at the moment, we need to create a few organisation- and industry-specific controlled vocabularies. The last time that I did something like this (establish a metadata schema for an organisation), I don’t think that we did a great job of getting the controlled vocabularies sorted out. Obviously, it was pretty easy for the elements that used external ones (like the Integrated Public Service Vocabulary or IPSV) but for the bespoke ones, we just didn’t get our act together.
So, for this one, I am going to get my extremely knowledgeable colleagues to create a starting point at our next team meeting which I will then put in front of the section heads before getting input on it from key people around the rest of the organisation. Once we have the necessary lists sorted out for my department and the organisation as a whole, I will start to get in touch with other organisations in the industry to try to achieve some convergence on this issue.
Creating the lists and identifying the elements (and creating the necessary supporting documentation) isn’t proving to be the difficult thing. It’s getting everyone to agree on a set of terms…
One of the challenges that we are going to face in this metadata project is the mapping of the DITA schema to the document metadata schema that we choose to implement (let’s call it Industry Schema).
We have a Technical Writer who is eager to introduce a content reuse policy (something very much in line with my own objectives) and would like to use a metadata architecture called the Darwin Information Typing Architecture (DITA) to do it. It’s basically an XML-based architecture that helps organisations create technical publications without having to recreate content. If you want to more, there is a pretty good Wikipedia entry on DITA and you could check out the DITA section of the Oasis website, the organisation now responsible for its maintenance (it started out as an IBM architecture).
My problem with using only this architecture is that when it comes time to create a composite document, how you know which bits of metadata should be used to describe the document? For example, it is entirely possible that each fragment has a different Subject and Creator. When assigning these elements of metadata to the composite document, I wouldn’t use any of the Creators as the Creator for the document nor would I use all of them – they would now be considered Contributors (at least in the Dublin Core Metadata Standard) and the Creator would be (I suppose) the organisation. As for the Subject, clearly a combination of individual components about certain concepts, when pulled together, do not make up a document that is about all of those separate concepts. So we need a document level metadata schema – enter Industry Schema.
I’m thinking that by creating a document that maps the Industry Schema to the DITA Document Type Definition (DTD) that we use, we can populate the DITA architecture based on the metadata associated with the original document.
When it comes time to catalogue a composite document, we will have to do so from scratch. This isn’t the end of the world; we’d have to do so if we weren’t using DITA to create composite documents so it isn’t like we have to do more. I just can’t help but think that the DITA metadata could be used to inform the Industry Schema metadata that we choose to assign. Is there a tool that we could use?
There is no doubt an article in here somewhere on the use and application of metadata at different levels of content granularity (document versus segment in this case). I would really like to speak to someone who has done this before…
One of the objectives that I have agreed with my line manager is to identify, or if necessary, to develop a metadata schema that our department, organisation and possibly even our industry, could use to organise our publications and documents.
Having had a look around, there doesn't seem to be any industry-wide schema in use and having spoken to a few key people in the organisation, there doesn't seem to be a understanding of what I'm on about let along something in use. So...
Today I have completed a draft version of our metadata schema and sent it around my team for their comments / questions / suggestions. The schema is based on the Dublin Core, DC Terms and the e-Government Metadata Standard. I'm quite pleased with it, though it still needs some work. For starters, there are at least two controlled vocabularies that I need to sort out. I might even have to create them. Hmm...although that's going to delay things a little, I suppose that, along with the final standard, it will provide me with more amunition for my portfolio!