Showing posts with label controlled vocabularies. Show all posts
Showing posts with label controlled vocabularies. Show all posts

Monday, July 02, 2007

Thesaurus vs taxonomy vs subject headings

A few weeks ago, I attended a course on how to build a thesaurus (opens in new window). Having thought about it a fair bit since the course, I’ve come to the conclusion that what we probably want to do here, at least in the short term, is not to create a thesaurus but to put together a list of subject terms. Here’s my thinking…

A thesaurus is a comprehensive listing of all the possible terms that someone might use to describe content within our subject (not an official definition but it’ll do for my purposes here). Although this means that whatever term someone might use, the thesaurus should make sure that they use the same one as the cataloguer, the problem is in the time required to create one. One of the things that I picked up from my course is that building a thesaurus from scratch is a very large task. Having discovered that there isn’t anything quite right out there already, I think I’d be looking at quite a bit of work. Also, although it is more flexible, I need to use more terms to describe the subject. Where a single subject heading will capture it, I might need to use several terms from the thesaurus.

A taxonomy has the same description limitation as the thesaurus (potentially many terms required) but is much quicker to produce. Of course, a user needs to guess the right term where a thesaurus will direct them to it, providing they’ve had a reasonably good guess in the first place.

Subject headings are a little less quick to produce than a taxonomy but still a quicker process than a thesaurus and they can describe the subject a little more effectively (this is because they provide context which an isolated word does not possess). Subject headings have their own pitfalls, of course, particularly when it comes to consistency of use over time.

So here is how I sort of see things looking:

(I don't know what's going on with this table...scroll down...it's there...I will look into it later as I must get going now!)










































 

Taxonomy

Thesaurus

Subject headings

Time to create

Fast

Slow

Medium

Versatility

High

High

Low

Ease of use

Low

High

Medium

Ease of Maintenance

High

Low

Low

Structure

Low

High

Medium



The time to create is an initial outlay of resource and is not ongoing. As a result, from a long-term perspective, the fact that this might be Slow isn’t too critical. The bit that makes a thesaurus less attractive as an option is the effort required to maintain it (of course there is software available to assist). The appeal to me of the subject headings is their time to create and their relative ease of use. I think that I would start the subject heading creation process by building a taxonomy and use that to develop the subjects. After creating and introducing subject headings initially, I think that I should use that same taxonomy (and the finished subject headings) to develop a thesaurus and aim to use that in the longer-term.

What do you think? Are there flaws in my thinking here? Are there other aspects that I should consider? This hasn’t exactly been a scientific process so there is every chance that I have missed something…

So, was the course a waste of time/money? Not at all. I couldn’t have arrived at a conclusion around the best way forward for my organisation without having attended it, even though the conclusion was to not build a thesaurus. Also, if/when we do get to the point where we want to introduce a thesaurus, I will have an idea as to where to start.



Thursday, June 14, 2007

How to construct a thesaurus

This is the first CILIP course that I've attended though it isn't the first time I've been here (while studying, I carried out a feasibility study for what was then the Library Association).

There are a couple of courses that were on yesterday and the coffee and registration is in one room for both courses. This is a great idea from a networking perspective but the room is full of tables and chairs meaning that delegates aren't as encouraged to mingle as much as they might be if there were only a few tables. The result, though, is that if people are talking they're only talking to the one or two people in their immediate vicinity. (Hmmm...got here first and sat at a table...room began to fill up but no one is sitting at my table :( Could it be the shaved head, goatee and swastika tattooed on my forehead? Kidding.)

So, at 9:30, we were called up to our meeting room to begin the training session. The presenter, Keith Trickey, was very knowledgeable and a good speaker and he dove straight into things. I have a bit of an academic hidden inside me who really enjoys talking/thinking about things like the ways in which nouns/verbs define adjectives/adverbs as much as the adjectives/adverbs describe the nouns/verbs. For example: heavy suitcase (a suitcase that is heavy) vs. heavy smoker (a person who smokes a lot). This relationship is important because in a thesaurus, words are usually separated from each other meaning that the information that travels in the context is lost. The same is true when verbs are turned into nouns – the verb carries some additional information in its conjugation and context (e.g. manages --> management) that it loses when it becomes a noun.

Anyway, having talked about these problems with dissecting words and compiling them into a thesaurus, it was straight into some of the rules and methods. (It is not my intention to record all of my notes here – I have assembled a mind map to supplement the course handouts for my future reference). Then it was in to some exercises. The exercises were fun and helped to illustrate some of the challenges and to demonstrate some of the rules in practice.

Finally, we spent a bit of time talking about the process of creating a thesaurus. This included ‘top tips’ and a look at some of the software that can be used for creating thesauri. My only criticism of the course, which I thought was otherwise really good, was that the only piece of software that we really talked about in any detail and had a look at was Multites. Quite a lot of time was given over to playing around with it – I started to get bored here...I don’t need to see each of the ways you can add a term in this particular programme.

On the whole, my feedback form contained positive comments/scores (with the exception of this one point about the time spent on Multites).

So, what did I learn? Well, the rules for assembling a list (like whether terms should be singular or plural) are good practical things that I can take away and put into use when it comes to creating a subject thesaurus. I think that I also have a better understanding about the different strengths and weaknesses of the different options (subject headings, taxonomies, thesauri, etc). I also feel more confident about evaluating existing thesauri for application to my contexts.

Having consolidated my thoughts from yesterday, the next set of actions on this matter that I have are around identifying potential existing thesauri for our use and then looking for anyone else who is looking to introduce or has recently introduced, a rail industry-specific thesaurus. Having seen what is involved in creating one from scratch, doing so is definitely my last choice!


Monday, June 11, 2007

Metadata and the RMS

The Business Support team (of which I am a member) has been working on the last few aspects of a project to introduce a workflow management system that we refer to as the Research Management System (RMS). When I started looking at how we might apply a metadata schema to our publications and how we might store that metadata (database vs. embedded), the RMS seemed like the most sensible way of capturing the metadata and certainly presented a way of storing it as well (though this isn’t my preferred option – more on that later).

Creating a schema

I looked around at other organisations and tried to find a schema that we might adopt but they were either too tailored to their current applications or were too general. My conclusion was that it was best to create something that could be adopted by the industry as a whole but that would certainly meet R&D’s needs. I have already written a bit about this work but as a reminder, I basically assembled a series of elements from the Dublin Core and the e-Government Metadata Standard and then created a few that were specific either to the industry (e.g. Asset Type) or to R&D (e.g. Research Objective). For as many elements as possible, I have used existing, internationally-recognised encoding schemes (e.g. W3C’s Date-Time format) and for the R&D-specific ones, I have used schemes developed through a consultative process with our Heads of Research Section and a selection of Research Managers (e.g. Audience Group). I have created the Application Profile though I have yet to create the XML definitions and publish them on our website, though this is ultimately my objective.

Applying the metadata

The RMS is due to go live next week and the metadata schema, along with the controlled vocabularies that I have had to create to support some of the R&D- and industry-specific fields, will be put to the test. Without going into too much detail here, we have worked as many of the elements into the process flow as possible so that as our Research Managers work through a project and record it in the RMS, some of the data that they enter is held in metadata fields for later application to any publications that emerge from the research. Clearly, not all of the elements can be populated this way (title for example can only be completed once the publication has been completed) but many can (such as research topic).

At the end of the research process, there is a knowledge management stage where the method for publishing, promoting and evaluating the publication is captured and the remaining metadata elements are completed. Some of these are set to defaults that will almost certainly not need to change, such as the publisher (that is pretty much always going to be our organisation). It is all looking like it should work but of course there is really only one way to find out for sure.

"Subjectlessness"

My one disappointment in this project was the inability to sort out a controlled vocabulary for subject element in time for the launch of the RMS. The problem I encountered is that existing controlled vocabularies are either too granular (for example, SELCAT (http://www.levelcrossing.net/) have a highly detailed thesaurus on the topic of Level Crossings, just one of the many areas of research that we pursue) or insufficiently granular (the Integrated Public Sector Vocabulary, or IPSV, directs users to categorise anything having to do with the railways, from electrification to passenger crowding, under Rail transport). Developing one of our own is just too big a task to try to sort out in only a few weeks (at the same time as all of the other projects that I am working on) so for the time being, the field in the RMS will be populated with: [IPSV] Rail transport. Although this is largely useless to us, it helps us tie in our work with that of the Department for Transport and other government bodies applying the IPSV.

Towards completion and "subjectfulness"?

The next stage of this project for me will have three aspects:
  1. The first is to refine the schema and existing vocabularies (Do we have all of the elements that we need? Are there any that have been included that just aren’t necessary? Are the controlled vocabularies: sufficiently granular? too granular? incomplete?)
  2. The second will be to create a final version of the application profile and to publish the element and encoding scheme definitions that I have had to create to accommodate some of the metadata that is specific to our requirements.
  3. The third and final aspect is to resolve this issue of subject headings. I am attending a CILIP workshop tomorrow called “How to construct a thesaurus” which I am hoping will give me some ideas and strategies for solving this problem. I would like to use existing vocabularies (to build in as much interoperability as possible) where possible. Maybe the solution will be to use things like the SELCAT vocabulary and the rolling stock manufacturers’ parts vocabularies but to only go down to a particular level within them (not sure what the implications of doing so would be, yet).
Metadata in the database vs. metadata in the document

I have already written about this little bug-bear of mine but it is still an issue for me. At the moment, I am going along with a database-held metadata solution but this is largely due to the presence of this option and the distinct lack of any alternatives. I think that once the metadata schema is relatively set and the encoding schemes in use, that I will turn my attention to resolving the issue of how we embed the metadata into the documents themselves...


Monday, April 02, 2007

Folksonomies

This is one of those things that I am quite interested in but know relatively little about. It seems to me that given my area of work tends to focus a lot on metadata, I should have a greater knowledge about folksonomies; we are, after all, talking about metadata (generally subject metadata) assigned by members of the public. As a result, I decided to do a bit of research on the web to learn more. One of the many articles that I came across was a literature review (opens in new window) carried out by an assistant librarian at the Royal College of Music. In her review, Edith Speller outlined some of the themes and provided a long and helpful list of articles for further reading (it was a lit review afterall!). I have picked a few of those articles out and had a look and below are what I have learnt and some of my thoughts:
  • I wonder how easy it is to get people to add their own metadata…if you look at the metadata assigned to music files, the data is partial at best and this is data that can be freely and automatically downloaded.

  • With potentially an infinite number of “taggers” how can we hope to achieve consistency and accuracy? For example, do we use singular or plural nouns and what about the use of capitalisation?

  • A major plus to this concept (over the dictated thesaurus or taxonomy) is that the terms that are used to describe or classify are chosen by the end users themselves and so, in volume, are going to be the right ones.

  • The fact that there are so many different people potentially tagging items, the basic level of variation comes into effect. This concept refers to the level of detail into which an individual will go to describe the subject or format of the item (e.g. personal music player vs. iPod nano)
    • The tags can be perfectly accurate and yet totally useless if they’re not at the right level for the individual’s needs

  • Folksonomies are flexible and evolve with time to match current terms and concepts unlike fixed systems like the DDC into which librarians are continuously cramming new concepts and terms.
    • The flip side of this flexibility, or the price of it, is that unless previously tagged items are reclassified, the items become lost and inconsistency creeps in making these items difficult or even impossible to find
    • It could be argued that it something can’t be found using current terminology, it is likely that the contents of that document have been superseded…

  • There is no active management of synonyms in folksonomies leading to reduced recall
    • To mitigate this problem, some sights (e.g. Del.icio.us) make all tags used visible

  • As for synonyms, there is no active management of homonyms (Apple the company vs apple the fruit) which leads to reduced search precision
    • By searching using more than one term (Apple and music vs apple and pastry) limited context can be achieved and search precision improves

  • My final thought on folksonomies is that users will need to ‘learn’ a slightly different vocabulary for each new database, at least initially. Over time and use, the differences could be ironed out and convergence could occur for those sights that appeal to a broad base of users.
I think that there could be a great deal of value in using a "folksonomy" approach to tagging content on the corporate Intranet. Rereading the points above, the benefits are ones that would be welcomed in that environment (e.g. terms chosen by ‘the people’) and the draw-backs are somewhat mitigated by the close relationship between the users (e.g. the use of synonyms and homonyms).

We are looking at our overarching IT, IS, IM, and KM strategy at the moment and attention will eventually focus on the very poor Intranet system that we are using. At that time, I think that I will suggest it!


Wednesday, February 28, 2007

Rail Industry Metadata Standard (RIMS)

Having done some research, it looks like there isn’t yet a metadata standard that is either designed for use or in common use by the rail industry. Although it wasn’t a surprise, it was a bit of a set back as it meant having to assemble a proposed metadata schema for the description of materials within the rail industry, engage in consultation starting with the R&D senior managers, and testing to validate it.

Building the metadata standard: the birth of RIMS
I have chosen to draw from the Dublin Core Metadata Standard (DCMS), and the Dublin Core Terms (DCTerms) to flesh it out a little. I have also drawn from the e-Government Metadata Standard (eGMS) to ensure that any public sector aspects are also covered. Now, the eGMS is based on the DMS and so far, what I have described would pretty much describe the eGMS. The rail industry, and our R&D work in particular, requires a little more granularity than the eGMS currently provides, something the Cabinet Office acknowledges by encouraging enhancement and refinement to suit different contexts. So I have also added some additional elements that are specific to the rail industry (e.g. asset type) and to our organisation (research topic). The complete picture is what we will consider to be the Rail Industry Metadata Standard (RIMS).

Identifying, assembling and building controlled vocabularies
Once I had a proposed set of elements, I set about sorting out the required controlled vocabularies. In my, albeit relatively limited, experience, this is the most difficult part. For many of the fields drawn from established standards, it was pretty straight forward (e.g. date formats us the W3C-recomended date-time format). For some of the elements that I had to create (e.g. research topic), it was also pretty straight forward because such lists were specific to the company and, in many cases, already in current use. Others from both established standards and the new set, however, were much more difficult. One such example is the asset type element – how granular do you go? For most of us, the term ‘locomotive’ is sufficiently descriptive but for our engineers it’s just too broad. My approach to these controlled vocabularies has been to put together a starting point and seek input and comments. So far, I have only engaged the R&D team and the lists have been heavily refined and accepted by them.

Subject
My other challenge has been sorting out a subject matter controlled vocabulary and it is proving to be a somewhat daunting task. The Integrated Public Service Vocabulary (IPSV), recommended as the controlled vocabulary for DCMS Subject, treats everything to do with the rail industry as ‘Rail Transport’. Clearly, this isn’t going to be sufficient for our requirements. I started to have a go at this task in the same way as I approached sorting out some of the other controlled vocabularies but it has proven to be too big. At the moment, it’s on hold while I move the rest of the project forward with a space reserved for subject tags and start to look for other initiatives both here and around Europe that are working towards creating a controlled vocabulary of some sort for the rail industry.

Handling the metadata
There are basically two different ways of managing document metadata: you can hold the metadata in a table which includes the location of the document described and then use this table to search and retrieve documents or you can embed the metadata into the documents themselves and search that (In reality, the search software or engine will most likely create its own table of metadata as in the case of the first method but this is a temporary table that is understood to need regular updating so is not the source of the metadata). Each method has its strengths and weaknesses (e.g. the table is quicker and simpler to deliver while the embedded data means that when someone downloads the document to a local space, the metadata travels with it and isn’t lost).

At the moment, we are also in the process of introducing a business process management system (we are calling it the Research Management System or RMS). The RMS will allow us to store documents as well as manage their production and approval. As a result, it makes sense that we piggy back the metadata assignment on the RMS work meaning that we will be going down the table route. This isn’t my preferred option but it is the one that will mean that we get metadata gathered and stored sooner. Once that process in embedded, we can look at technologies that will enable us to embed that gathered metadata into the files so that users downloading them from our website take the metadata with them.

One challenge that remains, and for which we have a few options but haven’t decided on any one yet, is what we do with the legacy collection. It has been decided that past projects and their associated documents will not be uploaded into the RMS. So the RMS presents us with the solution for future publications but it doesn’t deal with the existing collection. It is most likely that we will upload the previous publications to a separated segment of the RMS which will store the metadata in the same way but there are a couple of alternatives solutions…more on this as it develops.

Where from here?
The next thing to do with this standard is to confirm that it works in practice which will be part of the embedding process for the RMS. We will then look to consult the rest of the organisation on the suitability of the metadata schema and its associated controlled vocabularies for wider use in the company. I guess you could think of our work as a bit of a pilot for the rest of the company.

I’d like to publish our schema and controlled vocabularies under creative commons and invite other organisations to comment on it or use it in their organisations.


Friday, October 13, 2006

Creating Controlled Vocabularies

As part of the Industry Schema metadata work that I am doing at the moment, we need to create a few organisation- and industry-specific controlled vocabularies. The last time that I did something like this (establish a metadata schema for an organisation), I don’t think that we did a great job of getting the controlled vocabularies sorted out. Obviously, it was pretty easy for the elements that used external ones (like the Integrated Public Service Vocabulary or IPSV) but for the bespoke ones, we just didn’t get our act together.

So, for this one, I am going to get my extremely knowledgeable colleagues to create a starting point at our next team meeting which I will then put in front of the section heads before getting input on it from key people around the rest of the organisation. Once we have the necessary lists sorted out for my department and the organisation as a whole, I will start to get in touch with other organisations in the industry to try to achieve some convergence on this issue.

Creating the lists and identifying the elements (and creating the necessary supporting documentation) isn’t proving to be the difficult thing. It’s getting everyone to agree on a set of terms…