Concepts and their measurement
TYPES OF VARIABLE
One of the most important features of an understanding of statistical operations is an appreciation of when it is permissible to employ particular tests. Central to this appreciation is an ability to recognize the different forms that variables take, because statistical tests presume certain kinds of variable, a point that will be returned to again and again in later chapters.
The majority of writers on statistics draw upon a distinction developed by Stevens (1946) between nominal, ordinal and interval/ratio scales or levels of measurement. First, nominal (sometimes called categorical) scales entail the classification of individuals in terms of a concept. In the Job-Survey data, the
58 Concepts and their measurement
variable ethnicgp, which classifies respondents in terms of five categories— white, Asian, West Indian, African and other—is an example of a nominal variable. Individuals can be allocated to each category, but the measure does no more than this and there is not a great deal more that we can say about it as a measure. We cannot order the categories in any way, for example.
This inability contrasts with ordinal variables, in which individuals are categorized but the categories can be ordered in terms of ‘more’ and ‘less’ of the concept in question. In the Job-Survey data, skill, prody and qual are all ordinal variables. If we take the first of these, skill, we can see that people are not merely categorized into each of four categories—highly skilled, fairly skilled, semi-skilled and unskilled—since we can see that someone who is fairly skilled is at a higher point on the scale than someone who is semi-skilled. We cannot make the same inference with ethnicgp since we cannot order the categories that it comprises. Although we can order the categories comprising skill, we are still limited in the things that we can say about it. For example, we cannot say that the skill difference between being highly skilled and fairly skilled is the same as the skill difference between being fairly skilled and semi-skilled. All we can say is that those rated as highly skilled have more skill than those rated as fairly skilled, who in turn have greater skill than the semi-skilled, and so on. Moreover, in coding semi-skilled as 2 and highly skilled as 4, we cannot say that people rated as highly skilled are twice as skilled as those rated as semi-skilled. In other words, care should be taken in attributing to the categories of an ordinal scale an arithmetic quality that the scoring seems to imply.
With interval/ratio variables, we can say quite a lot more about the arithmetic qualities. In fact, this category subsumes two types of variable— interval and ratio. Both types exhibit the quality that differences between categories are identical. For example, someone aged 20 is one year older than someone aged 19, and someone aged 50 is one year older than someone aged 49. In each case, the difference between the categories is identical—one year. A scale is called an interval scale because the intervals between categories are identical. Ratio measures have a fixed zero point. Thus age, absence and income have logical zero points. This quality means that one can say that somebody who is aged 40 is twice as old as someone aged 20. Similarly, someone who has been absent from work six times in a year has been absent three times as often as someone who has been absent twice. However, the distinction between interval and ratio scales is often not examined by writers because, in the social sciences, true interval variables frequently are also ratio variables (e.g. income, age). In this book, the term interval variable will sometimes be employed to embrace ratio variables as well.
Interval/ratio variables are recognized to be the highest level of measurement because there is more that can be said about them than with the other two types. Moreover, a wider variety of statistical tests and procedures are available to interval/ratio variables. It should be noted that if an interval/ratio variable like age is grouped into categories—such as 20–29, 30–39, 40–49, 50–59, and so
on—it becomes an ordinal variable. We cannot really say that the difference between someone in the 40–49 group and someone in the 50–59 group is the same as the difference between someone in the 20–29 group and someone in the 30–39 group, since we no longer know the points within the groupings at which people are located. On the other hand, such groupings of individuals are sometimes useful for the presentation and easy assimilation of information. It should be noted too, that the position of dichotomous variables within the threefold classification of types of variable is somewhat ambiguous. With such variables, there are only two categories, such as male and female for the variable gender. A dichotomy is usually thought of as a nominal variable, but sometimes it can be considered an ordinal variable. For example, when there is an inherent ordering to the dichotomy, such as passing and failing, the characteristics of an ordinal variable seem to be present.
Strictly speaking, measures like satis, autonom and routine, which derive from multiple-item scales, are ordinal variables. For example, we do not know whether the difference between a score of 20 on the satis scale and a score of 18 is the same as the difference between 10 and 8. This poses a problem for researchers since the inability to treat such variables as interval means that methods of analysis like correlation and regression (see Chapter 8), which are both powerful and popular, could not be used in their connection since these techniques presume the employment of interval variables. On the other hand, most of the multiple-item measures created by researchers are treated by them as though they are interval variables because these measures permit a large number of categories to be stipulated. When a variable allows only a small number of ordered categories, as in the case of commit, prody, skill and qual in the Job-Survey data, each of which comprises only either four or five categories, it would be unreasonable in most analysts’ eyes to treat them as interval variables. When the number of categories is considerably greater, as in the case of satis, autonom and routine, each of which can assume sixteen categories from 5 to 20, the case for treating them as interval variables is more compelling.
Certainly, there seems to be a trend in the direction of this more liberal treat- ment of multiple-item scales as having the qualities of interval variables. On the other hand, many purists would demur from this position. Moreover, there does not appear to be a rule of thumb which allows the analyst to specify when a variable is definitely ordinal and when interval. None the less, in this book it is proposed to reflect much of current practice and to treat multiple-item measures such as satis, autonom and routine as though they were interval scales. Labovitz (1970) goes further in suggesting that almost all ordinal vari- ables can and should be treated as interval variables. He argues that the amount of error that can occur is minimal, especially in relation to the considerable advantages that can accrue to the analyst as a result of using techniques of analysis like correlation and regression which are both powerful and relatively easy to interpret. However, this view is controversial (Labovitz 1971) and
60 Concepts and their measurement
whereas many researchers would accept the treatment of variables like satis as interval, they would cavil about variables like commit, skill, prody and qual. Table 4.1 summarizes the main characteristics of the types of scale discussed in this section, along with examples from the Job-Survey data.
In order to help with the identification of whether variables should be classified as nominal, ordinal, dichotomous, or interval/ratio, the steps articulated in Figure 4.1 can be followed. We can take some of the Job-Survey variables to illustrate how this table can be used. First, we can take skill. This variable has more than two categories; the distances between the categories are not equal; the categories can be rank ordered; therefore the variable is ordinal. Now income. This variable has more than two categories; the distances between them are equal; therefore the variable is interval/ratio. Now gender. This variable does not have more than two categories; therefore it is dichotomous. Finally, we can take ethnicgp. This variable has more than two categories; the distances between the categories are not equal; the categories cannot be rank ordered; therefore, the variable is nominal.