• No results found

4.4 Asymptotic Results

4.4.1 Estimation of the Drift Function

Note that for the use of modified kernels, extra attention has to be paid to the kernel moments. For simplicity assume Gj = [0, 1]. The numerator of Kh(u, v) is equal to one if v ∈ [h, 1 − h] and depends on v (but not on h) otherwise. Kernel constants are defined as

κl(u) = Z 1

0

(u − v)l Kh(u − v) R1

0 Kh(w − v) dwdv.

Easy calculations show that three cases have to be distinguished

κl(u) =







 R1

−1vlK(v) dv for u ∈ [2h, 1 − 2h]

R1

−1vlK(v) dv + O(hl+1) for u ∈ [h, 2h] ∪ [1 − 2h, 1 − h]

R1

0(u − v)lKh(u − v) dv + O(hl+1) for u ∈ [0, h] ∪ [1 − h, 1]

.

The modified kernels only have an influence at boundary points u ∈ [0, 2h] ∪ [1 − 2h, 1], where they differ from usual kernel constants. Analogously, kernel constants κ2l =R1

0(u − v)l(Kh(u, v))2dv are defined.

4.4.1 Estimation of the Drift Function

Without loss of generality the exposition is restricted to the case of estimating the first component of the drift vector µ1(x). The Nadaraya-Watson smooth backfitting estimators eµ1,jh (xj), j = 1, . . . , d are defined as the iterative solution of the set of equations (4.11) and the normalization (4.9). Their asymptotic properties are given in the following

Theorem 4.1. Let Assumptions 4.1 and 4.2 be fulfilled and the additive model (4.5) with centering (4.6) hold. For the bandwidth sequence it holds that h2 =

86 4. ESTIMATING ADDITIVE DIFFUSIONS

O((T h)−1/2) and nh3 → ∞. Then the algorithm (4.11) converges with geometric rate and for the estimators eµ1,j,N Wh (xj), j = 1, . . . , d it holds that

√T hµe1,j,NWh (xj) − µ1,j(xj) − b1,jh − βµ1,j,N W(xj) pv1(xj2(xj)/κ0(xj)2

−−→ N (0, 1),D

where

b1,jh = − Z Z

µ1,j(xj)Kh(xj, uj)f (uj) dujdxj

− h

Z Z

∂xjµ1,j(xj)κ1(xj)

κ0(xj)Kh(xj, uj)f (uj) dxj βµ1,j,N W(xj) = hκ1(xj)

κ0(xj)

∂xjµ1,j(xj) + h2βeµ1,j(xj) and

v1(xj) = (fj(xj))−1E(a11(X) | Xj = xj).

Note that the first part of the bias βµ1,j,N W(xj) is zero for xj ∈ [h, 1 − h] and therefore only present at the boundary. The second part is not given in explicit form, it is only defined as

( eβµ1,0, eβµ1,j(x1), . . . , eβµd,j(xd))

= arg min

βµ1,0,...,βµ1,d

Z

µ1(x) − βµ1,0− βµ1,1(x1) − · · · − βµ1,d(xd))2f (x) dx,

with

βµ(x) = κ2(xj) κ0(xj)

Xd j=1

(f (x))−1

∂xj1,j(xj))

∂xj(f (x)) +1 2

∂(xj)2µ1,j(xj).

Therefore the bias can be interpreted as the projection of βµ(x) on the space of additive functions with respect to the L2(f )-norm.

The term b1,jh converges to zero asymptotically since it holds that Z Z

µ1,j(xj)f (uj)Kh(xj, uj) dujdxj

= Z

uj∈[h,1−h]/

Z

µ1(xj)fj(xj)(Kh(xj, uj) − Kh(xj − uj)) dxjduj+ O(h2)

= O(h).

4.4 Asymptotic Results 87

and the second term is of order O(h2) because κ1(xj) is zero at interior points xj. However, this term is constant over xj and does therefore only affect the normalization of µ1,j(xj). It is generated by the difference between the empirical normalization (4.9) used in the algorithm and the theoretical normalization (4.6) and does not influence the shape of the estimator.

Convergence of the algorithm follows from consistency of the (one and two-dimensional) kernel density estimators. In particular, the unknown function do not have to be additive. If the additive model does not hold, the estimators will converge to a projection of the high-dimensional function onto the space of additive functions. In Chapter 5 this case is investigated for independent and identically distributed data.

The limit distribution of the vector (eµ1,1,N Wh (x1), . . . , eµ1,d,N Wh (xd)) is a multivari-ate (d-dimensional) normal distribution where the covariances are zero asymp-totically. Considering the joint estimation of the additive components of µi(x) and µi0(x) there are asymptotically non-vanishing covariances, given by

cov(

hT eµi,j,NWh (xj),√

hT eµih0,j,N W(xj)) = κ2(xj)

f (xj) E(aii0(X) | Xj = xj).

To judge the efficiency of the Smooth backfitting estimator it has to be compared to the oracle estimator, which is based on knowledge of all other µ1,i(xi), i 6= j.

With this knowledge, the response variables could be modified to Yk∆? = ∆−1(X(k+1)∆1 − Xk∆1 ) −X

i6=j

µ1,i(Xk∆i )

=

Z (k+1)∆

k∆

µ1,j(Xsj) ds + Xd

i=1

Z (k+1)∆

k∆

σ1i(Xs) dWsl

+X

i6=j

Z (k+1)∆

k∆

1,i(Xsi) − µ1,i(Xk∆i )) ds.

Then, the infeasible oracle estimator is given by ˇ

µ1,j,N Wh (xj) =

PnT −1

k=0 Kh(xj, Xk∆i )Yk∆? PnT −1

k=0 Kh(xj, Xk∆i ) .

The knowledge of the other components allows to estimate µ1,j(xj) from discrete data only. The discretization errors are of order OP(n−3/2) and therefore do not affect the estimation asymptotically. Therefore it holds that (even under weaker assumptions than Theorem 4.1)

√T hµˇ1,j,NWh (xj) − µ1,j(xj) − ˇβ1,j(xj) pκ2(xj)v1(xj)/κ0(xj)

−−→ N (0, 1),D

88 4. ESTIMATING ADDITIVE DIFFUSIONS

where

βˇ1,j(xj) = hκ1(xj) κ0(xj)

∂xj1,j(xj)) + h2

µκ2(xj) κ0(xj)

∂xj1,j(xj))∂xj(f (xj)) f (xj) +1

2

2

∂(xj)2µ1,j(xj)

.

This follows from Lemmata 4.2 and 4.4 in the appendix and therefore, the smooth backfitting estimator eµ1,j,N Wh (xj) achieves the same variance as the oracle estima-tor, but has a different bias. This is the same efficiency result as in the classical regression setting, which was shown by Mammen, Linton and Nielsen (1999). To understand the bias behavior recall that the smooth backfitting estimator can be regarded as a projection of the full dimensional Nadaraya-Watson estimator onto the space of additive functions. Theorem 4.1 shows that the bias of the smooth backfitting estimator is the additive projection of the bias of the full-dimensional estimator. But this is not additive because the stationary density of the process f (x) is in general not additive. In contrast, the bias of a full-dimensional local linear estimator is additive and consists of the sum of the second derivatives of the additive components (times a constant). Smooth backfitting based on the lo-cal linear estimator can again be regarded as a projection of the full-dimensional local linear estimator. In the next theorem it will be shown that the bias of the local linear smooth backfitting estimator is again the projection of the bias of the full-dimensional estimator and therefore local linear backfitting is fully oracle ef-ficient. The design independence of local linear estimation, which means that the bias is independent of the density of the regressors, carries over to the projected estimators and drives the efficiency result.

The next theorem states the asymptotic properties of the local linear smooth backfitting estimators, defined in equations (4.16) and (4.17).

Theorem 4.2. Let Assumptions 4.1 and 4.2 be fulfilled and the additive model (4.5) with centering (4.6) hold. For the bandwidth sequence it holds that h2 = O((T h)−1/2) and nh3 → ∞. Then the algorithm (4.16) converges with geometric rate and for the estimators eµ1,j,LLh (xj), j = 1, . . . , d it holds that

√T hµe1,j,LLh (xj) − µ1,j(xj) − b1,jh − βµLL(xj) pv1(xj)eκ(xj)

−−→ N (0, 1),D

4.4 Asymptotic Results 89

with b1,jh and v1(xj) as given in Theorem 4.1 and where βµLL(xj) = h21

2 κ2(xj) κ0(xj)

2

∂(xj)2µ1(xj) e

κ(xj) = κ20(xj2(xj) − κ1(xj21(xj) κ0(xj2(xj) − (κ1(xj))2 .

For interior points, the variance reduces to eκ(xj) = κ20(xj) because all other kernel constants are zero or one. In contrast to local constant smooth backfitting, the bias is given in explicit form. To derive the oracle efficiency, consider the unfeasible local linear estimator based on the data Yk∆ . Applying Lemmata 4.2 and 4.4, the asymptotic properties of the oracle estimator (under Assumptions 4.1 and 4.2) are given by

√T hµˇ1,j,LLh (xj) − µ1,j(xj) − βµLL(xj) pv1(xj)eκ(xj)

−−→ N (0, 1).D

From this it can be seen that both bias and variance are identical to the ex-pressions in Theorem 4.2. Therefore the local linear estimators are fully oracle efficient.