在医学实验中,常用的两种研究设计分别为队列研究和病例-对照研究。队列研究多为前瞻性,其结果可作为较强的医学证据,但面对罕见疾病时,队列研究往往需要较大的样本量才能获得足够的病例个体,研究成本高昂。病例-对照研究为追溯性,可以自由控制病例组和对照组的样本量,但病例-对照研究易受人为因素干扰,因此其结果被认为是较弱的医学证据。为了综合两者的优势,研究者提出在队列研究中使用二阶段抽样方法。通过过滤部分信息含量较少的对照个体,来缩小样本量,进而降低研究成本。传统二阶段抽样设计,包括病例-队列设计和病例-对照设计等,对于对照个体均采用等概率抽样的方式,忽略了观察时间中的信息。为了提高一阶段样本利用率,研究者提出了一种新的抽样设计,终端抽样设计。终端抽样设计采用不等概率抽样,对照个体的入样概率随着其观察时间的增加而增加,由此使协变量的分布更为广泛,从而达到改善参数估计效率的目的。 关于终端抽样设计,当前研究主要针对比例危险率模型,在边际独立和条件独立假设下,分别利用极大似然和逆概率加权方法,得到了回归参数的估计,并证明了相应的大样本性质。然而,现有方法仅适用于比例危险率模型。在实际应用中,个体间危险率函数成比例的假定有时难以满足。并且随着数据维度的提高,变量选择的重要性与日俱增,而现有估计不具备稀疏性。另外,当前研究均假定模型正确识别,没有考虑模型错误识别时的推断问题。论文首先将终端抽样设计从比例危险率模型推广至更一般的线性变换模型。其次在线性变换模型下提出了基于稀疏惩罚项的变量选择方法。最后考虑了当模型错误识别时,比例危险率模型针对包括终端抽样数据在内的二阶段抽样数据的稳健推断问题。 对于参数估计问题,论文通过对全量数据的似然函数进行逆概率加权调整,消除了二阶段样本中观测不到的协变量,使其能够正确处理二阶段抽样数据。在一定的正则性条件下,最大化逆概率加权似然函数得到的估计具有一致性,并且收敛到一个高斯过程。仿真结果显示,在样本量接近的情况下,相比病例-队列设计和病例-对照设计,终端抽样设计对应估计量的经验标准误更小,估计效率更高。 对于变量选择问题,由于线性变换模型无法像比例危险率模型一样,通过截面似然方法得到一个只包含回归参数的目标函数,因此需要将其视作一个脆弱模型。利用EM算法求解极大似然估计过程中M步的目标函数,配合自适应Lasso惩罚项,可以获得回归参数的稀疏估计,并且该估计具有一致性和神谕性(Oracle)。在仿真实验中,对比平均选中特征数,平均误选特征数和真模型选中比例三个指标,终端抽样设计的表现均优于两种传统抽样设计。 对于稳健推断问题,论文利用逆概率加权方法得到了回归参数的估计。由于模型是错误识别的,不存在参数真值,但该估计关于伪真值仍具有一致性和渐近正态性。相应的协方差矩阵估计可由插入法得到。仿真结果显示,即使模型识别错误,终端抽样设计的估计效率依然高于病例-队列设计和病例-对照设计。 在推广至线性变换模型后,终端抽样设计的应用得到了拓展;基于稀疏惩罚项的变量选择方法可以根据实际问题,利用不同的惩罚项灵活调整,适用面非常广泛,同时其也为未来处理高维数据做好了准备;稳健推断方法解决了二阶段抽样与错误识别同时存在时,关于比例危险率模型回归参数的统计推断问题,完善了二阶段抽样数据的分析框架。 关键词:队列研究,终端抽样设计,二阶段抽样设计,变量选择,稳健推断
In medical experiments, cohort study and case-control study are two common research designs. Cohort study is usually prospective and its results are viewed as strong evidence. However, when the disease outcome of interest is rare, sample size must be very large in order to obtain adequate cases, which makes the study quite costly. Case-control study is retrospective. Researchers can determine the sample size of case group and control group separately. But case-control study is susceptible to human factors. Thus, its result are considered as weak evidence. To combine the advantages of both studies, researchers suggest using two-phase sampling approach in cohort study. By excluding part of controls with less information, the sample size can be reduced and the cost is lower. In traditional two-phase sampling design, including case-cohort design and case-control design, controls are drawn with equal probability, which means the information in observed time is ignored. To make best use of the full cohort, end-point sampling design is proposed. In end-point sampling design, the inclusion probabilities are unequal. Controls with larger observed time are more likely to be included into the two-phase sample. This makes covariates more spread out, so the efficiency for parameter estimation can be improved. Current research on end-point sampling design mainly focus on the proportional hazard model. Estimator for regression parameter can be obtained by the maximum likelihood method under marginal independent assumption and the inverse probabil-ity weighting method under conditional independent assumption, respectively. The large sample properties are developed. However, the existing approach is only avail-able for the proportional hazard model. In real life application, the constant hazard ratio assumption is sometimes violated. Also, with the increase of data dimension, variable selection becomes more and more important, but the existing estimator is not sparse. Besides, current research assume that the model is correctly specified. The misspecified situation is not taken into account. This paper firstly generalize end-point sampling design from the proportional hazard model to the linear trans-formation model. Next, under the linear transformation model, a variable selection method based on sparse penalty is proposed. Finally, the robust inference for two-phase sampling data, including end-point sampling data, under the misspecified proportional hazard model is considered. For parameter estimation, the likelihood function is adjusted by inverse prob-ability weighting. After excluding the unobservable covariates, the function is ap-plicable to two-phase sampling data. Under suitable regularity conditions, it can be proved that the estimator obtained by maximizing inverse probability weighting likelihood function, is consistent and converges to a Gaussian process. Simulation shows that the estimator from end-point sampling data has smaller empirical stan-dard errors than that from case-cohort or case-control data with comparable size of covariates ascertainment. For variable selection, unlike the proportional hazard model, profile likelihood method cannot be applied to the linear transformation model to derive an objec-tive function with regression parameter only. Therefore, the linear transformation model is treated as a frailty model. By utilizing the objective function in M-step of EM algorithm and adaptive Lasso penalty, the sparse estimator for the regres-sion parameter can be obtained. Asymptotic properties, such as consistency and oracle property, are also derived. In simulation, end-point design outperforms two traditional designs in average number of selected features, average number of falsely selected features and percentage of true model being selected. For robust inference, the inverse probability weighting approach is adopted to estimate the regression parameter. Since the model is misspecified, true value does not exist. However, the estimator is still consistent with pseudo-true value, and converges to the normal distribution. Moreover, the variance-covariance matrix can be consistently estimated by using the plugged-in approach. Simulation shows that under the misspecified proportional hazard model, end-point sampling design has higher efficiency than case-cohort and case-control design. After being generalized to the linear transformation model, the application of end-point sampling design is extended. The variable selection method based on sparse penalty is flexible, since the penalty can be adjusted for specific purpose.The robust inference procedure is developed for the proportional hazard model under the situation where biased sampling and model misspecification exist simultaneously, which completes the analysis framework for two-phase sampling data. Keywords: Cohort Study,End-point Sampling Design,Two-phase Sampling De-sign,Variable Selection,Robust Inference