Robust Variable Selection and Estimation Based on Kernel Modal Regression

Changying Guo; Biqin Song; Yingjie Wang; Hong Chen; Huijuan Xiong

doi:10.3390/e21040403

Robust Variable Selection and Estimation Based on Kernel Modal Regression

Entropy ◽

10.3390/e21040403 ◽

2019 ◽

Vol 21 (4) ◽

pp. 403 ◽

Cited By ~ 1

Author(s):

Changying Guo ◽

Biqin Song ◽

Yingjie Wang ◽

Hong Chen ◽

Huijuan Xiong

Keyword(s):

Variable Selection ◽

Gaussian Noise ◽

Free Variable ◽

Noise Condition ◽

New Model ◽

Information Theoretic Learning ◽

Model Free ◽

Modal Regression ◽

Gradient Based ◽

Non Gaussian

Model-free variable selection has attracted increasing interest recently due to its flexibility in algorithmic design and outstanding performance in real-world applications. However, most of the existing statistical methods are formulated under the mean square error (MSE) criterion, and susceptible to non-Gaussian noise and outliers. As the MSE criterion requires the data to satisfy Gaussian noise condition, it potentially hampers the effectiveness of model-free methods in complex circumstances. To circumvent this issue, we present a new model-free variable selection algorithm by integrating kernel modal regression and gradient-based variable identification together. The derived modal regression estimator is related closely to information theoretic learning under the maximum correntropy criterion, and assures algorithmic robustness to complex noise by replacing learning of the conditional mean with the conditional mode. The gradient information of estimator offers a model-free metric to screen the key variables. In theory, we investigate the theoretical foundations of our new model on generalization-bound and variable selection consistency. In applications, the effectiveness of the proposed method is verified by data experiments.

Download Full-text

A gradient-based eigenspace approach to dealing with occlusions and non-Gaussian noise

Object recognition supported by user interaction for service robots ◽

10.1109/icpr.2002.1048469 ◽

2003 ◽

Author(s):

H. Wildenauer ◽

T. Melzer ◽

H. Bischof

Keyword(s):

Gaussian Noise ◽

Gradient Based ◽

Non Gaussian

Download Full-text

A model-free variable selection method for reducing the number of redundant variables

Statistics ◽

10.1080/02331888.2018.1515949 ◽

2018 ◽

Vol 52 (6) ◽

pp. 1212-1248

Author(s):

Anchao Song ◽

Tiefeng Ma ◽

Shaogao Lv ◽

Changsheng Lin

Keyword(s):

Variable Selection ◽

Free Variable ◽

Selection Method ◽

Variable Selection Method ◽

Model Free

Download Full-text

On dual model-free variable selection with two groups of variables

Journal of Multivariate Analysis ◽

10.1016/j.jmva.2018.06.003 ◽

2018 ◽

Vol 167 ◽

pp. 366-377

Author(s):

Ahmad Alothman ◽

Yuexiao Dong ◽

Andreas Artemiou

Keyword(s):

Variable Selection ◽

Free Variable ◽

Dual Model ◽

Model Free

Download Full-text

Knockoff boosted tree for model-free variable selection

Bioinformatics ◽

10.1093/bioinformatics/btaa770 ◽

2020 ◽

Author(s):

Tao Jiang ◽

Yuanyuan Li ◽

Alison A Motsinger-Reif

Keyword(s):

Variable Selection ◽

Principal Component ◽

Free Variable ◽

Supplementary Information ◽

Type I ◽

Test Statistics ◽

Linear Regression Models ◽

Model Free ◽

Tree Models ◽

Boosted Tree

Abstract Motivation The recently proposed knockoff filter is a general framework for controlling the false discovery rate (FDR) when performing variable selection. This powerful new approach generates a ‘knockoff’ of each variable tested for exact FDR control. Imitation variables that mimic the correlation structure found within the original variables serve as negative controls for statistical inference. Current applications of knockoff methods use linear regression models and conduct variable selection only for variables existing in model functions. Here, we extend the use of knockoffs for machine learning with boosted trees, which are successful and widely used in problems where no prior knowledge of model function is required. However, currently available importance scores in tree models are insufficient for variable selection with FDR control. Results We propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. We extend the current knockoff method to model-free variable selection through the use of tree-based models. Additionally, we propose and evaluate two new sampling methods for generating knockoffs, namely the sparse covariance and principal component knockoff methods. We test and compare these methods with the original knockoff method regarding their ability to control type I errors and power. In simulation tests, we compare the properties and performance of importance test statistics of tree models. The results include different combinations of knockoffs and importance test statistics. We consider scenarios that include main-effect, interaction, exponential and second-order models while assuming the true model structures are unknown. We apply our algorithm for tumor purity estimation and tumor classification using Cancer Genome Atlas (TCGA) gene expression data. Our results show improved discrimination between difficult-to-discriminate cancer types. Availability and implementation The proposed algorithm is included in the KOBT package, which is available at https://cran.r-project.org/web/packages/KOBT/index.html. Supplementary information Supplementary data are available at Bioinformatics online.

Download Full-text

An improved G-music algorithm for non-Gaussian noise condition direction-of-arrival estimation

2015 23rd Iranian Conference on Electrical Engineering ◽

10.1109/iraniancee.2015.7146261 ◽

2015 ◽

Author(s):

Mahmoud Ahmadi ◽

Ehsan Yazdian ◽

Ali A. Tadaion

Keyword(s):

Gaussian Noise ◽

Direction Of Arrival ◽

Direction Of Arrival Estimation ◽

Music Algorithm ◽

Noise Condition ◽

Non Gaussian

Download Full-text

Model-free variable selection

Journal of the Royal Statistical Society Series B (Statistical Methodology) ◽

10.1111/j.1467-9868.2005.00502.x ◽

2005 ◽

Vol 67 (2) ◽

pp. 285-299 ◽

Cited By ~ 47

Author(s):

Lexin Li ◽

R. Dennis Cook ◽

Christopher J. Nachtsheim

Keyword(s):

Variable Selection ◽

Free Variable ◽

Model Free

Download Full-text

Gradient-induced Model-free Variable Selection with Composite Quantile Regression

Statistica Sinica ◽

10.5705/ss.202016.0222 ◽

2018 ◽

Author(s):

Shaogao Lv ◽

Xin He ◽

Junhui Wang

Keyword(s):

Variable Selection ◽

Quantile Regression ◽

Free Variable ◽

Composite Quantile Regression ◽

Model Free

Download Full-text

Multiple Loci Mapping via Model-free Variable Selection

Biometrics ◽

10.1111/j.1541-0420.2011.01650.x ◽

2011 ◽

Vol 68 (1) ◽

pp. 12-22

Author(s):

Wei Sun ◽

Lexin Li

Keyword(s):

Variable Selection ◽

Free Variable ◽

Model Free ◽

Multiple Loci

Download Full-text

Error Bound of Mode-Based Additive Models

Entropy ◽

10.3390/e23060651 ◽

2021 ◽

Vol 23 (6) ◽

pp. 651

Author(s):

Hao Deng ◽

Jianghong Chen ◽

Biqin Song ◽

Zhibin Pan

Keyword(s):

Variable Selection ◽

Hilbert Spaces ◽

Regression Models ◽

Reproducing Kernel ◽

Additive Models ◽

Generalization Error ◽

Proposed Model ◽

Modal Regression ◽

Kernel Hilbert Spaces ◽

Non Gaussian

Due to their flexibility and interpretability, additive models are powerful tools for high-dimensional mean regression and variable selection. However, the least-squares loss-based mean regression models suffer from sensitivity to non-Gaussian noises, and there is also a need to improve the model’s robustness. This paper considers the estimation and variable selection via modal regression in reproducing kernel Hilbert spaces (RKHSs). Based on the mode-induced metric and two-fold Lasso-type regularizer, we proposed a sparse modal regression algorithm and gave the excess generalization error. The experimental results demonstrated the effectiveness of the proposed model.

Download Full-text