• No results found

Measures of Variability

In document Six Sigma Green Belt Manual (Page 35-38)

Mode The most frequently appearing number(s) in a set of data. Useful when data displays wide variation, perhaps due to mixed processes.

For the data set:

1,2,2,3,3,3,3,4,4,5,5,6,7 three is the mode

Measures of Variability

Measure Description & Use How to Calculate

Range The difference between the largest and smallest values in a data set.

R = x

max

x

min

Variance The sum of the squared differences of the data from the mean, divided by the number of data less one.4 Forms the basis for the standard deviation.

s

The square root of the variance. The Standard Deviation can be thought of as a

“distance” measure - showing how far the data are away from the mean value.

s= s2

4This is the sample standard deviation. If the entire population is known, there is no need to subtract one from n, the number of data.

2.3 LINE GRAPHS & RUN CHARTS

Purpose

Line graphs are basically graphs of your performance indicator/CTQ taken over time. They help you see where the “center” of the data tends to be, the variability in performance, trends, cycles and other patterns. Line graphs are very simple to construct.

One of the most important factors to keep in mind for line graphs is that the data must be plotted in the order in which it occurs. Losing this order will prevent you from seeing patterns that are time dependent.

Application

Virtually any data can be placed on a line graph (as long as you’ve kept it in order of occurrence).

Some typical line graph applications include:

• Quality Indicators/CTQs - Turn-around Times, Errors, Defect Rates, Defective Proportions, Physical parameters - Condenser Vacuum, Machine Start Times, Setup Times, Pressure, Temperature Readings taken periodically, chemical or drug concentrations (peak and trough levels),

• Personal data - Weight, heart rate,

• Financial Data - Salary Expense, Supply Costs, Sales, Volumes.

Construction of Line Graphs

1. Draw a vertical and a horizontal axis on a piece of graph paper.

2. Label the vertical axis with the variable being plotted.

3. Label the horizontal axis with the unit of time or order in which the numbers were collected (i.e. Day 1, 2, 3, . . ., Customer 1, 2, 3, . . . etc.).

4. Determine the scale of the vertical axis. The top of this axis should be about 20 percent larger than the largest data value. The bottom of this axis should be about 20 percent lower than the smallest data value. This let’s you see the best picture of the process’ variability. Label the axis in convenient intervals between these numbers.

5. Plot the data values on the graph number by number, preserving the order in which they occurred.

6. Connect the points on the graph.

7. (Optional) Calculate the mean of the data and draw this as a solid line through the data. This turns the line graph into a run chart - trends and patterns are often easier to see with a run chart.

Construction Notes

Try to get about twenty five (25) data points to get a line graph running. If you have less, go ahead and plot them anyway. It's good to start trending performance no matter how many points you currently have (i.e. if you process only “produces” one data per month - salary expense, supply costs, etc. - don’t wait two years to start your line graph!).

Now, here's how you actually get these 25 points for a line graph. If you are dealing with measurement data (time, cost, etc.), then each event you measure represents a data point to be plotted. Each patient’s temperature measurement could be plotted on a line graph: 98.5, 98.7, 98.6, 99.0, 98.4, etc.

If you are dealing with count data, though (or even worse, percentages made up of count data), then there are a few guidelines that may cause some data “heartburn.”

For typical count data, the guideline is that the mean of the data you plot should at least equal to 5, and no less than 1. Let's say you are counting the number of errors that occur on a daily basis. You get these numbers for a week's worth of errors: 7, 10, 6, 5, 8, 7, 6. The mean of errors (daily) is 7. This number is greater than 5, so you can plot the daily values as individual point.

We apply this rule for two reasons. First, to "see" variation in the process, we need to keep the data away from the horizontal (0 value) axis. The second reason lies in why you are taking the data in the first place: to take action. If you want to detect whether your change has had an effect, you’ll want to see its impact on the line graph.

Errors per 1000 Orders

Orders (1000) Errors

0 2 4 6 8 10

1 3 5 7 9 11 13 15 17 19

Now let's look at a different set of values. In counting orders for a particular specialty magazine (again, daily), a publications distributor finds that their first week gives this data: 0, 1, 0, 2, 1, 0, 1. Here, the daily mean value is less than 1. This mean doesn't meet the guidelines and plotting these data won’t produce a very useful line graph.

The distributor could group the data by combining enough days to make the mean equal or better than 5. In this case, there are 5 orders occurring per week. So, instead of plotting the daily occurrence of orders, they plot the weekly orders. To get a line graph going here, note that they are going to have to observe at least 125 events (25 points x 5 - mean).

This is difficult since it’s now going to take about 25 weeks to get a complete line graph instead of only 25 days. This kind of thing happens often when we start to stratify processes that are low volume to begin with for the company down to an individual department. The process just doesn't give us enough data for a line graph.

One way of getting around this problem is to plot the time between events. For example, one company was studying employee injuries. They measured the time between injuries. Since this is measurement data, it "only" took 26 injuries to get a good line graph going.

Percentage (or Proportion) Data - The guideline for plotting one point on a line graph (where the percentage is count data divided by count data, i.e. errors per 1000 orders) is that the numerator's mean should be greater than or equal to 3 and the denominator's mean should be greater than or equal to 50. You can see the implications of this on the amount of data and time needed to get a line graph going.

In document Six Sigma Green Belt Manual (Page 35-38)