Abstract
Let $X = [ ξ_1, \,\, ξ_2,...\,\, ,ξ_d]^\top$ be a zero-mean random vector of large dimension $d$ ($d \rightarrow \infty$) with (hidden) covariance matrix $M = (m_{ij})_{1 \leq i, j \leq d},$ where $m_{ij} = m_{ji} = \textbf{Cov}(ξ_i, ξ_j).$ Let $X_1, X_2, \dots, X_n$ be $n$ iid samples of $X$. Consider the sample covariance matrix $$\textstyle \tilde{M} := \frac{1}{n} \sum_{i=1}^{n} X_i X_i^\top.$$ In practice, one frequently uses the eigenvectors and eigenspaces of $\tilde M$ as estimators for those of $M$. A central task is to provide an error analysis for these estimators. In this paper, we provide an optimal error analysis, obtaining upper and lower bounds of matching order of magnitude, for a wide range of parameters $d$ and $n$, under mild assumptions on $M$. As corollaries, we obtain new necessary and sufficient conditions for the consistency of the estimators. In these conditions, we only require the number of samples $n$ to depend linearly on the effective rank of $M$, which can be much smaller than the dimension $d$.
摘要
设$X = [ ξ_1, \,\, ξ_2,...\,\, ,ξ_d]^\top$是一个大维数$d$($d \rightarrow \infty$)的零均值随机向量,具有(隐藏)协方差矩阵$M = (m_{ij})_{1 \leq i, j \leq d},$其中$m_{ij} = m_{ji} = \textbf{Cov}(ξ_i, ξ_j).$ 设$X_1, X_2, \dots, X_n$是$X$的$n$个iid样本。考虑样本协方差矩阵$$\textstyle \tilde{M} := \frac{1}{n} \sum_{i=1}^{n} X_i X_i^\top.$$ 在实践中,我们常用$\tilde M$的特征向量和特征空间作为$M$的估计。核心任务是为这些估计提供误差分析。在本文中,我们在对$M$施加温和假设的条件下,针对广泛的参数$d$和$n$,给出了最优误差分析,得到了匹配阶的上界和下界。作为推论,我们得到了估计量一致性的新的充要条件。在这些条件中,我们仅要求样本数$n$线性依赖于$M$的有效秩,而有效秩可以远小于维数$d$。