• No results found

Modelling perceptual compensation for the effects of reverberation 1

4.3 Control of efferent suppression

In the previous section, the consequence of manually applying efferent attenuation was demonstrated using a model based around the efferent-DRNL implementation of Ferry and Meddis (cf. § 4.2.2, and Figure 4.8 in particular). As shown in Fig-ure 4.1b, the current model extends the work of Ferry and Meddis by introducing an ecosystemic control mechanism (Di Scipio, 2003) that determines the effective-ness of the efferent feedback loop. By varying the level of attenuation applied in the model in proportion to the amount of reverberation detected in the recent history of the acoustic surroundings, auditory efferent suppression is thus used as a candidate theory for explaining the effects of perceptual compensation for reverberation.

This section describes the conditions that such a control mechanism must satisfy if it is to allow simulation of perceptual compensation for the effects of reverberation that is observed in human speech perception. Two prospective metrics for reverber-ation estimreverber-ation are first described below, each of which makes an assessment of the amount of reverberation present in the signal, and then (automatically) uses this value to update the value of the efferent attenuation applied in the model. The se-lected reverberation measures are subsequently used to model human listener data collected by Watkins (2005a) in a series of experiments below (Experiments M1, M2 and M3).

In an ideal world, the metric deriving control for efferent attenuation would be based on human physiology and function. While behavioural data gathered in psychoacoustic studies is beginning to elucidate various mechanisms underlying perceptual compensation for reverberation, relatively little is yet known about the physiological factors influencing these processes. Thus compensation for rever-beration is not yet well-enough understood for a model to be biologically accu-rate. Rather, the current model makes a crude summarisation of known auditory processing in order to create a functional simulation of the compensation effect.

Here, all higher levels of the auditory system are lumped together and treated as a

‘black-box’ which outputs a single efferent signal. This efferent signal acts on the peripheral auditory system and thereby regulates the afferent processing behaviour.

This simplification is easily seen in Figure 3.1b; only the final descending pathway is modelled.

Functionally, then, the job of the feedback circuit is to make an assessment of the amount of reverberation present in the signal, and to use this value to automatically update the attenuation parameter in the DRNL filter bank. Two questions thus im-mediately arise. The first asks how reverberation should be quantified; the second queries the time-period over which this quantification should be done. These two areas are discussed in turn below.

It is generally accepted that the preceding context appears to inform a listener’s de-cision about a subsequent test-word. It is not yet fully understood, however, what the nature of this influential information may be. One theory, termed ‘modulation masking’, was put forward by Nielsen and Dau (2010). This explanation suggests that human listeners may adapt to the degree of modulation present in the preced-ing context signal; thus it would appear that a measure of dynamic range might prove useful in modelling compensation for the effects of reverberation. One such measure, the mean-to-peak ratio (MPR) is described below. On the other hand, Watkins (2005a) argues that listeners are informed by mechanisms that detect and compensate for the ‘reverberation tails’ present in a signal. More recently, Watkins et al. (2011) have suggested further that this compensation mechanism may be informed by an assessment of temporal envelopes within individual auditory

chan-nels. The second measure described below, the low-pass mask (LPM) reverberation estimator is based on these principles. Many other approaches could of course be investigated to quantify reverberation in the auditory modelling task but are not investigated further in the current work. One that has proved successful in ASR involves modelling the excitation signal for voiced speech with the linear predic-tion residual (Ananthapadmanabha and Yegnanarayana, 1979), and examining the higher order statistics (i.e. the kurtosis) to quantify the ‘peakiness’ of the resulting probability distribution (e.g., Gillespie et al., 2001). Another approach that looks interesting from a biological point of view would be to examine the auditory onset-rather than offset- response (see e.g., Heil, 2003; Longworth-Reed et al., 2009).

The time course of the monaural compensation effect has yet to be studied in de-tail1. Since Watkins (2005a) has repeatedly demonstrated compensation effects using single utterances, however, it seems that we are interested in a fairly rapid mechanism. However, in the earlier discussion of compensation for reverberation it became apparent that various constancies might be active on different time-scales (cf. § 2.4). Indeed, looking across recent studies that examined the timescales on which binaural compensation effects operate, it appears that the relevant timescale for contextual information may depend critically on the listener task. Long-term learning (over around 5 hours of exposure to a particular room condition) can im-prove performance in listeners’ localisation accuracy (Shinn-Cunningham, 2000).

On the other hand, listeners’ ability to determine the azimuth of a test pulse appears to be impeded by just seconds of inconsistent reverberation on the preceding con-text (Zahorik et al., 2009). For binaural speech-perception tasks, experiments have shown compensation occurring at the minimum temporal resolution of the anal-ysed data: in minutes for sentence sets in Longworth-Reed et al. (2009); within six sentences in Srinivasan and Zahorik (2013); and in just a few seconds in Bran-dewie and Zahorik (2010). BranBran-dewie and Zahorik (2013) recently designed a study specifically to measure the time course of the binaural effect, and reported that 850 ms2of room exposure was sufficient to achieve considerable speech intel-ligibility enhancement.

Having thus derived a measure of reverberation over a particular time period, whether it be based on dynamic-range or on reverberation tails, the measure is then used to linearly control the efferent attenuation applied in the model as de-scribed below, i.e. attenuation increases as the level of reverberation increases.

This is similar to the noise-based modelling strategies employed elsewhere, where

1Experiment H4 below directly addresses this point in Chapter 5.

2Interestingly, they found that the compensation mechanism appeared to slow down when the listener task involved dealing with an additional noise component in the stimuli.

an yan(n,c) STEPENV Amp.

yan(n,c)

c = 1

Σ

C

Freq. (Hz)

100 8000

100

50 150 200 250

Time (ms) 0

Figure 4.9: Demonstration of the mean-to-peak ratio (MPR) reverberation estimator. Above:

STEP simulated auditory nerve response, yan(n, c). Below: Envelope (ENV) of the across-channel summed auditory nerve response (as described by Equation 4.4). In the lower panel, the mean level of the ENV signal is shown with a solid line; the peak is shown dotted. An increase in reverberation would raise the noise floor and reduce the dynamic range. Thus the peak would be little changed, while the mean value would correspondingly increase.

attenuation increases as the level of noise increases (Brown et al., 2010; Lee et al., 2011; Messing et al., 2009).

4.3.1 Dynamic range estimation: mean-to-peak ratio (MPR)

The mean-to-peak ratio (MPR), related to the ‘blurredness’ metric of Palom¨aki et al. (2004), was proposed in Beeston and Brown (2010) as a method to monitor the dynamic range of the simulated auditory nerve signal (or more specifically, it’s temporal envelope), and thereby arrive at an estimate of the amount of reverbera-tion present in the signal. The method relies on the assumpreverbera-tion that late-arriving reflections add additional energy to a signal which reduces its dynamic range as the noise floor rises. While the peak value of the signal remains more-or-less un-changed, the mean value rises with additional reflected energy. Thus, with MPR defined simply as the ratio of the mean and peak values, an increase in the level of reverberation (raising the mean value) will bring about a corresponding increase in MPR value recorded.

The upper panel of Figure 4.9 shows the STEP simulated in the auditory nerve (AN) in the afferent pathway of the model in response to the spoken word ‘stir’

(this STEP was previously shown in the final panel of Figure 4.2). The response in all C frequency channels is summed at each time step n, giving a pooled estimate

of auditory nerve activity,

ENVan(n) =

C

X

c=1

yan(n, c). (4.4)

The estimated level of reverberation, Rmp(n), at time step n, is then quantified by the ratio of the mean and peak values of the previous AN temporal envelope computed over a windowed portion of duration Z time frames,

Rmp(n) =

1 Z

n

P

z=n−Z

ENVan(z) maxn

z=n−ZENVan(z) (4.5)

where z indexes time frames within this temporal window.

Posed in this way, later-arriving reflections can be regarded as contributing addi-tional energy to a signal, filling dips in its temporal envelope with reflected energy, raising the noise floor and thereby reducing its dynamic range. The previous chap-ter reviewed work suggesting that the efferent auditory system is involved in gain control, and appears to bring about a suppression of the auditory nerve response in situations of additive background noise (see § 3.4.1). Thus it seems that a close re-lationship may exist between reverberation, noise suppression and dynamic range control. A model that adjusts efferent suppression by detecting the effects of re-verberation on the signal’s dynamic range thus explores the idea that low-level mechanisms controlling dynamic range in the auditory nerve might be involved in compensation for reverberation.

4.3.2 Reverberation tails estimation: low-pass mask (LPM)

The low-pass mask (LPM) metric is not concerned with the dynamic range of the signal, but instead attempts to capture information regarding offsets in the signal since these are frequently prolonged by the presence of a reverberant tail (Beeston and Brown, 2013; Kallasjoki et al., 2014). Inspired by missing data approaches to robust speech processing, LPM considers individual areas in the spectro-temporal representation to be either informative (reliable) or not (unreliable) about the off-sets in individual frequency channels of the simulated auditory nerve response.

Subsequent processing is based on the informative parts of the signal only; other parts are discarded in the reverberation estimation calculation. A single-channel demonstration, from which it can be inferred that LPM attempts to judge the amount of reverberation present based on the proportion of energy present during tails (offsets) in the signal, is presented in Figure 4.10.

Time (ms) yanylpylpmlpm lpyan

100

50 150 200 250

0

Figure 4.10: Demonstration of the low-pass mask (LPM) offset capture technique in a single-channel of the simulated AN response. Top to bottom: simulated AN response, yan(n, c), in a single high-frequency channel (where c = 76); smoothed temporal envelope in this channel, ylp(n − τ, c);

derivative of the envelope, y0lp(n − τ, c); binary mask, mlp(n, c), which locates offsets via the negative portions of the derivative; the masked signal resulting in that channel, mlp(n, c) yan(n, c).

Figure 4.10 describes how the LPM method locates ‘tail-like’ regions in yan(n, c), the simulated STEP resulting in the auditory nerve simulation. First, the smoothed temporal envelope in each channel, ylp(n − τ, c), is estimated using a second-order low-pass Butterworth filter with cutoff frequency at 10 Hz, and is temporally corrected to remove the filter delay, τ . The derivative of the envelope, ylp0 (n − τ, c), is then calculated and the binary mask, mlp(n, c), subsequently locates its negative portions such that

mlp(n, c) =

 1 if ylp0 (n − τ, c) < 0,

0 otherwise. (4.6)

Finally, the amount of reverberation present in the signal at time frame n is es-timated with a single number, Rlp(n), by computing the mean masked signal strength present in each of the C channels over a preceding context window of Z frames duration (so that z indexes frames within the window), and taking the across-channel mean to summarise these values:

Rlp(n) = 1 C

1 Z

C

X

c=1 n

X

z=n−Z

mlp(z, c) yan(z, c). (4.7)

An increase in the level of reverberation would typically cause a longer reverbera-tion tail, thereby increasing the proporreverbera-tion of signal contributing toward the

mea-sure, i.e. with value 1 in the binary mask mlp(n, c). An increase in reverberation would thus likely give rise to a corresponding increase in the value of Rlp(n).

There is some support in the both psychoacoustic and speech-technology literature for a ‘reverberation tail’ based approach to dealing with reverberation. For exam-ple, Watkins and colleagues propose that listeners are informed by a reverberation tail-based perceptual mechanism (Watkins, 2005a; Watkins et al., 2011). Addi-tionally, Javed and Naylor (2014) have recently suggested a metric based on the detection of such tails which correlates with objective measures of room reverber-ation and aims to predict the perceived impact of reverberreverber-ation on a given speech signal. Since reverberation tail metrics are asymmetric in time, they hold the poten-tial to further examine and possibly explain findings regarding speech perception in time-reversed rooms.

4.4 Experiment M1: Application of the efferent model to