5 Empirical methods
5.4. Semiparametric estimation for censored models: CLAD estimator
The censored least absolute deviations (CLAD) estimation to a regression model where the dependent variable is constrained to be non-negative was first developed by James L. Powell (1984) as a generalisation of the least absolute deviations estimator (LAD). The CLAD
estimator is semiparametric, as it is only partly parameterised: the uncensored mean xi'β is parameterised, but the error distribution is not (Cameron and Trivedi 2005:564). As the CLAD estimator starts with a similar problem of a left-censored-at-zero linear model for the latent variable as the tobit model, which is appropriate for modelling the data on the amount of sent remittances, the CLAD estimator is suitable for this purpose, as well.
68
5.4.1. Standard model and its estimation
In the semiparametric literature, the linear model for the latent variable yi = xiβ +εi
' *
, which
is left-censored at zero, is usually written as
{
i i}
i x
y =max0, 'β0 +ε , (5.22) where the dependent variable yi and the independent variable vector xi are observed for each
I, while the parameter vector β0 and the error term εi are not observed. The definition of the LAD estimator starts with the notion that for any scalar random variable Z, the function
[
Z b Z]
E − − is minimised by choosing b to be a median of the distribution of Z. It is further assumed that the error term εi is continuously distributed and has a median zero, and that the density function is positive at zero. What this means is that the median function for yi takes the form m
(
xi,β0)
=max{
0,xi'β0}
. Thus, the probability that yi =0 when xi'β0 >0, is less than one-half, and the median of y i is xi'β0. Conversely, if xi'β0 ≤0, the probability that0
=
i
y is more than one-half, and the median of yi is zero. (Powell 1984:305)
The CLAD estimator βˆ minimizes the sum of absolute deviations of N y i from
{
0}
', 0
max xiβ
over all β in parameter space B, or
( )
∑
{
}
= − = N i i i N y x N S 1 ' , 0 max 1 β β . (5.23)The existence of the minimum is ensured by the parameter space B being compact. However,
the behaviour of the regression function 0 'β
i
x must be restricted to ensure that the LAD
estimator βˆ is unique for large samples: the censored sample median provides a consistent N estimate of the population median, if less than a half of the sample is censored. Moreover, the independent variables xi are required not to be collinear for the uncensored observations. (Powell 1984:305-6; Cameron and Trivedi 2005:564)
5.4.2. Assumptions
To sum up the assumptions made that ensure the consistency of the CLAD estimator βˆ , first N the parameter vector β0 must be an element of a compact parameter space B. Second, the
69
error terms εi must be independently and identically distributed, independent of the independent variables xi, and have a median zero. Moreover, the distribution function of εi must be continuously differentiable with density f that is positive at zero and bounded above, i.e.
( )
x >k>0f λ i , (5.24)
whenever λ <k, some k >0, all i.
Third, the independent variables xi must be independently distributed random variables with
0 3
K x
E i < for all I and some positive K0. Moreover, the smallest characteristic root v i of the matrix
(
)
∑
≥ i i i i x x x N E 0 0 ' ' 1 1 β ε (5.25)has vi >v0 whenever N >N0, some positive ε0,v0 and N0.
Clearly, the assumptions on the distribution of the error term εi are much weaker than those required for consistency of the maximum likelihood or least squares estimators for the censored regression model. Moreover, some of the assumptions made above may even be relaxed. It is sufficient that the conditional distribution of εi given xi has median zero for all
I, and the corresponding distribution functions for εi only need to be continuously
differentiable in a uniform neighbourhood of zero, with density functions as shown in equation (5.25). Importantly, even the assumption of homoskedasticity is not needed for the
consistency of βˆ , as the conditional median of the dependent variable will still be of the N
form
{
' 0}
, 0
max xiβ . Under some further assumptions,βˆ is also asymptotically normal, which N
holds even if the error terms εi are heteroskedastic. (Powell 1984:307-12).
The facts that the consistency of the CLAD estimator does not depend on the functional form
of the error terms, and that it is robust to heteroskedasticity of the error terms εi makes the estimator for the censored regression model an attractive alternative to the sensitive maximum likelihood estimator – employed by both the tobit model and the Heckman selection model –
70
which is rendered inconsistent when either non-normality or heteroskedasticity are detected. As both of these were indeed detected in this study after the estimation of the tobit model, CLAD estimations were carried out. However, the practical estimation proved troublesome due to non-convergence, and the estimation could be carried out properly only with the Senegalese data. Moreover, the censored sample median would not even have provided a consistent estimate of the population median in the case of Uganda, where more than a half of the sample is censored. The CLAD estimation of the probability and amount of remittances for Senegal are presented in appendix 4 and may be considered as a robustness check to the tobit, Heckman selection and two-part model estimates for Senegal discussed in sections 7.2.1, 7.2.2 and 7.2.3.