3.6 Special measurement error cases
3.6.1 Heaping: rounding to favored distances
A chapter addressing measurement error in distance sampling would not be complete without a reference to heaping, an often reported effect of rounding to somewhat
preferred distances (e.g. Anderson and Pospahala, 1970; Rosenstock et al., 2002). If
the method for obtaining distances does not produce a precise measurement and there is any subjectivity involved in the process (e.g. visual estimation of distances), it is inevitable that the histogram of detected distances will show some distances that occur much more frequently than what would be expected, while some others are rarely recorded, if at all. An example of data with substantial heaping, collected as
part of a hare point transect survey4, is shown in figure 3.10.
Favored distances are usually multiples of 1, 5, 10 or 100, depending on the scale of measurements being made. Humans have a preference for round numbers, and in the absence of any better alternative, those are used. Although very frequent in real data, heaping has to be strong to have a clear effect on density estimates. Even the analysis of the raw data, if heaping is not severe, should not be problematic. If it is thought heaping might have a considerable influence in the results, we can use judicious grouping of distances so that heaped values are approximately at the center
of the distance intervals (Buckland et al., 2001, p. 109-110).
For illustration of the consequences of heaping, a small simulation example follows.
Radial distance (m) Density 0 50 100 150 0.00 0.01 0.02 0.03 0.04 0.05 0.06
Figure 3.10: A real life point transect data example (hares in Northern Ireland), showing strong heaping at multiples of 5 and 10 meters.
Distance Frequency 0 2 4 6 8 10 0 200 400 600 800 1000 1200 a) Distance 0 2 4 6 8 10 0 1000 2000 3000 4000 5000 b) Distance 0 2 4 6 8 10 0 2000 4000 6000 8000 10000 c)
Figure 3.11: Histograms of simulated data for testing the effect of analyzing distance data ignoring heaping. a) Original data; b) Distances recorded as the closest multiple of 0.5 with respect to original distance; c) Distances recorded as the closest multiple of 1 with respect to original distance. Note that in b) and c) the first histogram bar is (roughly) half the height of the second because it corresponds to values heaped at
0, i.e. X < 0.25 for b) and X < 0.5 for c), while all the other bars span values over
0.5 or 1 meter intervals, respectively for b) and c).
Using the same simulations as in section 3.3.1.4 (100 line transect data sets), I have introduced two levels of heaping in the original data (Figure 3.11). All distances were recorded as the closest multiple of 0.5 (strong heaping) and closest multiple of 1 (very strong heaping), and analyzed the data using the automated procedure described before. A note about figure 3.11b,c: because the first histogram bar corresponds to
true distances<0.5 m or<0.25 m, while all the others represent true distances in the
vicinity on either side of the corresponding number, the first bar is only accounting for half the width of the others, and hence it looks approximately half the height of the second bar.
The consequences of heaping for either density estimates or associated variance were negligible (Table 3.10).
Table 3.10: Results of simulation exercise to assess the effect of heaping in the es-
timation of density. Mean estimated density ˆD and respective CV for the analysis
performed in Distance 4 considering the true distances and distances with strong and very strong heaping. True density is 100 animals/ha.
Analysis Dˆ CV
True distances 99.02 0.0554
Strong heaping 98.92 0.0563
Very strong heaping 99.30 0.0562
responsible for some lack of independence between detections, and so special care must be given to the interpretation of goodness of fit measures. An example is
presented in Marqueset al. (in press).
However, heaping can be considered more like a hint for other problems, rather than a problem on its own. The presence of heaping usually means that the method used to measure distances was poor, and in such cases we can only hope that other errors, both random and systematic, were avoided.
The worst problem occurs when considerable heaping at 0 is present in the data, leading to a spiked detection function. This could be expected under several scenarios, like: (1) marine mammal surveys, where animals are detected at large distances from the observation platform, and rounding angles to 0 leads to heaping of perpendicular distances at 0; (2) Surveys on paths along which visibility is very good, coupled with animal movement, like flying birds on road surveys or small mammals along transects cut in dense forest; (3) incorrect definition of the measurement to be made, say a cluster recorded at 0 distance if any of the cluster members are on the line, when the true distance recorded should be the distance to the center of the cluster and (4) an imprecise definition of the transect. This might preclude appropriate estimation, usually leading to overestimation of abundance (provided the spike is
really an artifact). The best way to deal with this problem is to avoid it, since at (and close to) zero, distances should not be difficult to measure. It is important to plot your data at an early stage of data collection, to identify potential problems; it is much simpler to identify and remove heaping at the data collection stage than to salvage the data analysis.
Smearing was introduced by Butterworth (1982) as an attempt to deal with heap- ing (not only in sighting distances but also in the recording of angles). The idea is to replace the preferred distances by plausible distances, derived by sampling from the vicinity of the original distance in a sensible way. Several options have been proposed (e.g. Hammond, 1984; Buckland and Anganuzzi, 1988). Ideally, one would want to sample proportionally to the true detection function, in the vicinity from which the heaped data was actually coming, leading to the new values being more often larger than the original heaped values; however, the rounding error tends to increase with distance, which means that more often a larger value is rounded down than a smaller value rounded up to a preferred value, and the two tend to cancel out. This led to Buckland and Anganuzzi (1988) recommendation that a uniform smearing be used, as it is much simpler and not necessarily worse. The choice of smearing parameters
has been, so far, based on ad hoc methods.
It seems difficult to come up with simple models that represent heaping in an adequate way. Nonetheless, provided an experiment was designed to assess it, it seems to be possible to estimate an error model for heaping, namely using intervals rather than exact distances and estimating the probabilities of a distance being placed in given intervals, given the original interval the distance was in. Because it is unlikely for the presence of heaping without more general measurement error issues, it does not seem satisfactory to pursue methods which deal solely with this issue, rather than methods which deal with general measurement error issues and to some extent also
accommodate heaping.