Pub Date : 2026-07-01Epub Date: 2024-05-01DOI: 10.1016/j.ecosta.2024.04.004
David Donoho , Behrooz Ghorbani
<div><div>Consider estimation of the covariance matrix under relative condition number loss <span><math><mrow><mi>κ</mi><mo>(</mo><msup><mstyle><mi>Σ</mi></mstyle><mrow><mo>−</mo><mn>1</mn><mo>/</mo><mn>2</mn></mrow></msup><mover><mstyle><mi>Σ</mi></mstyle><mo>^</mo></mover><msup><mstyle><mi>Σ</mi></mstyle><mrow><mo>−</mo><mn>1</mn><mo>/</mo><mn>2</mn></mrow></msup><mo>)</mo></mrow></math></span>, where <span><math><mrow><mi>κ</mi><mo>(</mo><mstyle><mi>Δ</mi></mstyle><mo>)</mo></mrow></math></span> is the condition number of matrix <span><math><mstyle><mi>Δ</mi></mstyle></math></span>, and <span><math><mover><mstyle><mi>Σ</mi></mstyle><mo>^</mo></mover></math></span> and <span><math><mstyle><mi>Σ</mi></mstyle></math></span> are the estimated and theoretical covariance matrices. Recent advances in understanding the so-called <em>spiked covariance model</em> for <span><math><mstyle><mi>Σ</mi></mstyle></math></span>, are used here to derive a nonlinear shrinker which is asymptotically optimal among orthogonally-covariant procedures. These advances apply in an asymptotic setting, where the number of variables <span><math><mi>p</mi></math></span> is comparable to the number of observations <span><math><mi>n</mi></math></span>. The form of the optimal nonlinearity depends on the aspect ratio <span><math><mrow><mi>γ</mi><mo>=</mo><mi>p</mi><mo>/</mo><mi>n</mi></mrow></math></span> of the data matrix and on the top eigenvalue of <span><math><mstyle><mi>Σ</mi></mstyle></math></span>. For <span><math><mrow><mi>γ</mi><mo>></mo><mrow><mn>0.618033</mn><mo>.</mo><mo>.</mo><mo>.</mo></mrow></mrow></math></span>, even dependence on the top eigenvalue can be avoided. The optimal shrinker has three notable properties. First, when <span><math><mrow><mi>p</mi><mo>/</mo><mi>n</mi><mo>→</mo><mi>γ</mi><mo>≫</mo><mn>1</mn></mrow></math></span> is moderately large, it shrinks even very large eigenvalues substantially, by a factor <span><math><mrow><mn>1</mn><mo>/</mo><mo>(</mo><mn>1</mn><mo>+</mo><mi>γ</mi><mo>)</mo></mrow></math></span>. Second, even for moderate <span><math><mi>γ</mi></math></span>, certain highly statistically significant eigencomponents will be completely suppressed.Third, when <span><math><mrow><mi>γ</mi><mo>≫</mo><mn>1</mn></mrow></math></span> is very large, the optimal covariance estimator can be purely diagonal, despite the top theoretical eigenvalue being large and the empirical eigenvalues being highly statistically significant. This aligns with practitioner experience. Alternatively, certain non-optimal intuitively reasonable procedures can have small worst-case relative regret - the simplest being generalized soft thresholding having threshold at the bulk edge and slope <span><math><msup><mrow><mo>(</mo><mn>1</mn><mo>+</mo><mi>γ</mi><mo>)</mo></mrow><mrow><mo>−</mo><mn>1</mn></mrow></msup></math></span> above the bulk. For <span><math><mrow><mi>γ</mi><mo><</mo><mn>2</mn></mrow></math></span> this has at most a few percent relative regr
{"title":"Optimal Covariance Estimation for Condition Number Loss in the Spiked model","authors":"David Donoho , Behrooz Ghorbani","doi":"10.1016/j.ecosta.2024.04.004","DOIUrl":"10.1016/j.ecosta.2024.04.004","url":null,"abstract":"<div><div>Consider estimation of the covariance matrix under relative condition number loss <span><math><mrow><mi>κ</mi><mo>(</mo><msup><mstyle><mi>Σ</mi></mstyle><mrow><mo>−</mo><mn>1</mn><mo>/</mo><mn>2</mn></mrow></msup><mover><mstyle><mi>Σ</mi></mstyle><mo>^</mo></mover><msup><mstyle><mi>Σ</mi></mstyle><mrow><mo>−</mo><mn>1</mn><mo>/</mo><mn>2</mn></mrow></msup><mo>)</mo></mrow></math></span>, where <span><math><mrow><mi>κ</mi><mo>(</mo><mstyle><mi>Δ</mi></mstyle><mo>)</mo></mrow></math></span> is the condition number of matrix <span><math><mstyle><mi>Δ</mi></mstyle></math></span>, and <span><math><mover><mstyle><mi>Σ</mi></mstyle><mo>^</mo></mover></math></span> and <span><math><mstyle><mi>Σ</mi></mstyle></math></span> are the estimated and theoretical covariance matrices. Recent advances in understanding the so-called <em>spiked covariance model</em> for <span><math><mstyle><mi>Σ</mi></mstyle></math></span>, are used here to derive a nonlinear shrinker which is asymptotically optimal among orthogonally-covariant procedures. These advances apply in an asymptotic setting, where the number of variables <span><math><mi>p</mi></math></span> is comparable to the number of observations <span><math><mi>n</mi></math></span>. The form of the optimal nonlinearity depends on the aspect ratio <span><math><mrow><mi>γ</mi><mo>=</mo><mi>p</mi><mo>/</mo><mi>n</mi></mrow></math></span> of the data matrix and on the top eigenvalue of <span><math><mstyle><mi>Σ</mi></mstyle></math></span>. For <span><math><mrow><mi>γ</mi><mo>></mo><mrow><mn>0.618033</mn><mo>.</mo><mo>.</mo><mo>.</mo></mrow></mrow></math></span>, even dependence on the top eigenvalue can be avoided. The optimal shrinker has three notable properties. First, when <span><math><mrow><mi>p</mi><mo>/</mo><mi>n</mi><mo>→</mo><mi>γ</mi><mo>≫</mo><mn>1</mn></mrow></math></span> is moderately large, it shrinks even very large eigenvalues substantially, by a factor <span><math><mrow><mn>1</mn><mo>/</mo><mo>(</mo><mn>1</mn><mo>+</mo><mi>γ</mi><mo>)</mo></mrow></math></span>. Second, even for moderate <span><math><mi>γ</mi></math></span>, certain highly statistically significant eigencomponents will be completely suppressed.Third, when <span><math><mrow><mi>γ</mi><mo>≫</mo><mn>1</mn></mrow></math></span> is very large, the optimal covariance estimator can be purely diagonal, despite the top theoretical eigenvalue being large and the empirical eigenvalues being highly statistically significant. This aligns with practitioner experience. Alternatively, certain non-optimal intuitively reasonable procedures can have small worst-case relative regret - the simplest being generalized soft thresholding having threshold at the bulk edge and slope <span><math><msup><mrow><mo>(</mo><mn>1</mn><mo>+</mo><mi>γ</mi><mo>)</mo></mrow><mrow><mo>−</mo><mn>1</mn></mrow></msup></math></span> above the bulk. For <span><math><mrow><mi>γ</mi><mo><</mo><mn>2</mn></mrow></math></span> this has at most a few percent relative regr","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 216-267"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"141063561","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-04-06DOI: 10.1016/j.ecosta.2023.03.004
Gaëtan Louvet, Jakob Raymaekers, Germain Van Bever, Ines Wilms
The precision matrix that encodes conditional linear dependency relations among a set of variables forms an important object of interest in multivariate analysis. Sparse estimation procedures for precision matrices such as the graphical lasso (Glasso) gained popularity as they facilitate interpretability, thereby separating pairs of variables that are conditionally dependent from those that are independent (given all other variables). Glasso lacks, however, robustness to outliers. To overcome this problem, one typically applies a robust plug-in procedure where the Glasso is computed from a robust covariance estimate instead of the sample covariance, thereby providing protection against outliers. These estimators are studied theoretically, by deriving and comparing their influence function, sensitivity curve and asymptotic variance.
{"title":"The Influence Function of Graphical Lasso Estimators","authors":"Gaëtan Louvet, Jakob Raymaekers, Germain Van Bever, Ines Wilms","doi":"10.1016/j.ecosta.2023.03.004","DOIUrl":"10.1016/j.ecosta.2023.03.004","url":null,"abstract":"<div><div><span>The precision matrix that encodes conditional linear dependency relations among a set of variables forms an important object of interest in </span>multivariate analysis. Sparse estimation procedures for precision matrices such as the graphical lasso (Glasso) gained popularity as they facilitate interpretability, thereby separating pairs of variables that are conditionally dependent from those that are independent (given all other variables). Glasso lacks, however, robustness to outliers. To overcome this problem, one typically applies a robust plug-in procedure where the Glasso is computed from a robust covariance estimate instead of the sample covariance, thereby providing protection against outliers. These estimators are studied theoretically, by deriving and comparing their influence function, sensitivity curve and asymptotic variance.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 268-280"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"76567416","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
The information matrix test for a normal random vector is shown to coincide with the sum of the moment tests for all third- and fourth-order multivariate Hermite polynomials. The statistic is decomposed as the sum of the marginal information matrix test for a subvector, the conditional information matrix test for the complementary subvector, and a third leftover component. It is also shown that exact finite sample distributions can be obtained by drawing spherical Gaussian vectors and orthogonalising them using sample moments. These tests are applied to assess the implications of Gibrat’s law for US city sizes using the three most recent censuses.
{"title":"Multivariate Hermite polynomials and information matrix tests","authors":"Dante Amengual, Gabriele Fiorentini, Enrique Sentana","doi":"10.1016/j.ecosta.2024.01.005","DOIUrl":"10.1016/j.ecosta.2024.01.005","url":null,"abstract":"<div><div>The information matrix test for a normal random vector is shown to coincide with the sum of the moment tests for all third- and fourth-order multivariate Hermite polynomials. The statistic is decomposed as the sum of the marginal information matrix test for a subvector, the conditional information matrix test for the complementary subvector, and a third leftover component. It is also shown that exact finite sample distributions can be obtained by drawing spherical Gaussian vectors and orthogonalising them using sample moments. These tests are applied to assess the implications of Gibrat’s law for US city sizes using the three most recent censuses.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 22-48"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139771937","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2024-03-05DOI: 10.1016/j.ecosta.2024.02.005
Christine De Mol, Domenico Giannone, Lucrezia Reichlin
The asymptotic properties of ridge regression in large dimension are studied. Two key results are established. First, consistency and rates of convergence for ridge regression are obtained under assumptions which impose different rates of increase in the dimension between the first and the remaining eigenvalues of the population covariance of the predictors. Second, it is proved that under the special and more restrictive case of an approximate factor structure, principal component and ridge regression have the same rate of convergence and the rate is faster than the one previously established for ridge.
{"title":"The Asymptotic Equivalence of Ridge and Principal Component Regression with Many Predictors","authors":"Christine De Mol, Domenico Giannone, Lucrezia Reichlin","doi":"10.1016/j.ecosta.2024.02.005","DOIUrl":"10.1016/j.ecosta.2024.02.005","url":null,"abstract":"<div><div><span>The asymptotic properties<span> of ridge regression in large dimension are studied. Two key results are established. First, consistency and rates of convergence for ridge regression are obtained under assumptions which impose different rates of increase in the dimension </span></span><span><math><mi>n</mi></math></span> between the first <span><math><msub><mi>n</mi><mn>1</mn></msub></math></span> and the remaining <span><math><mrow><mi>n</mi><mo>−</mo><msub><mi>n</mi><mn>1</mn></msub></mrow></math></span><span> eigenvalues of the population covariance of the predictors. Second, it is proved that under the special and more restrictive case of an approximate factor structure, principal component and ridge regression have the same rate of convergence and the rate is faster than the one previously established for ridge.</span></div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 49-60"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140085412","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2024-01-09DOI: 10.1016/j.ecosta.2024.01.001
Dante Amengual, Xinyue Bei, Enrique Sentana
Tests are developed for neglected serial correlation when the information matrix is repeatedly singular under the null hypothesis. Specifically, consideration is given to white noise against a multiplicative seasonal Ar model, and a local-level model against a nesting Ucarima one. The proposed tests, which involve higher-order derivatives, are asymptotically equivalent to the likelihood ratio test but only require estimation under the null. It is shown that the tests effectively check that certain autocorrelations of the observations are zero, so their asymptotic distribution is standard. Monte Carlo exercises examine finite sample size and power properties, with comparisons made to alternative approaches.
在零假设条件下,当信息矩阵重复奇异时,对被忽视的序列相关性进行检验。具体而言,考虑了针对乘法季节性 Ar 模型的白噪声,以及针对嵌套 Ucarimaone 的局部模型。所提出的检验涉及高阶导数,在渐近上等同于似然比检验,但只需要在零假设下进行估计。结果表明,这些检验有效地检验了观测数据的某些自相关性为零,因此其渐近分布是标准的。蒙特卡罗练习检验了有限样本大小和功率特性,并与其他方法进行了比较。
{"title":"Highly irregular serial correlation tests","authors":"Dante Amengual, Xinyue Bei, Enrique Sentana","doi":"10.1016/j.ecosta.2024.01.001","DOIUrl":"10.1016/j.ecosta.2024.01.001","url":null,"abstract":"<div><div>Tests are developed for neglected serial correlation when the information matrix is repeatedly singular under the null hypothesis. Specifically, consideration is given to white noise against a multiplicative seasonal <span>Ar</span> model, and a local-level model against a nesting <span>Ucarima</span><span><span> one. The proposed tests, which involve higher-order derivatives, are asymptotically equivalent to the likelihood ratio test but only require estimation under the null. It is shown that the tests effectively check that certain </span>autocorrelations<span><span> of the observations are zero, so their asymptotic distribution is standard. Monte Carlo exercises examine finite </span>sample size and power properties, with comparisons made to alternative approaches.</span></span></div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 4-21"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139411045","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-04-25DOI: 10.1016/j.ecosta.2023.04.004
Matias Salibian-Barrera
Nonparametric regression models offer a way to understand and quantify relationships between variables without having to identify an appropriate family of possible regression functions. Although many estimation methods for these models have been proposed in the literature, most of them can be highly sensitive to the presence of a small proportion of atypical observations in the training set. A review of outlier robust estimation methods for nonparametric regression models is provided, paying particular attention to practical considerations. Since outliers can also influence negatively the regression estimator by affecting the selection of bandwidths or smoothing parameters, a discussion of robust alternatives for this task is also included. Using many of the “classical” nonparametric regression estimators (and their robust counterparts) can be very challenging in settings with a moderate or large number of explanatory variables, so recently proposed robust nonparametric regression methods that scale well with a growing number of covariates are also discussed.
{"title":"Robust nonparametric regression: Review and practical considerations","authors":"Matias Salibian-Barrera","doi":"10.1016/j.ecosta.2023.04.004","DOIUrl":"10.1016/j.ecosta.2023.04.004","url":null,"abstract":"<div><div>Nonparametric regression models offer a way to understand and quantify relationships between variables without having to identify an appropriate family of possible regression functions<span>. Although many estimation methods for these models have been proposed in the literature, most of them can be highly sensitive to the presence of a small proportion of atypical observations in the training set. A review of outlier robust estimation methods for nonparametric regression models is provided, paying particular attention to practical considerations. Since outliers can also influence negatively the regression estimator<span> by affecting the selection of bandwidths or smoothing parameters, a discussion of robust alternatives for this task is also included. Using many of the “classical” nonparametric regression estimators (and their robust counterparts) can be very challenging in settings with a moderate or large number of explanatory variables<span>, so recently proposed robust nonparametric regression methods that scale well with a growing number of covariates are also discussed.</span></span></span></div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 294-309"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"82760056","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-07-16DOI: 10.1016/j.ecosta.2023.07.002
Mario Forni, Luca Gambetti, Luca Sala
A procedure to estimate measures of macroeconomic uncertainty and compute the effects of uncertainty shocks based on standard VARs is proposed. Uncertainty and its effects are estimated using a single model so to ensure internal consistency. Under suitable assumptions, the procedure is equivalent to using the square of the VAR forecast error as an external instrument in a proxy SVAR. The procedure allows to add orthogonality constraints to the standard proxy SVAR identification scheme. The method is applied to a US data set; results show that macroeconomic uncertainty is responsible of a large fraction of business-cycle fluctuations while financial uncertainty plays a modest role.
{"title":"Macroeconomic uncertainty and vector autoregressions","authors":"Mario Forni, Luca Gambetti, Luca Sala","doi":"10.1016/j.ecosta.2023.07.002","DOIUrl":"10.1016/j.ecosta.2023.07.002","url":null,"abstract":"<div><div>A procedure to estimate measures of macroeconomic uncertainty and compute the effects of uncertainty shocks based on standard VARs is proposed. Uncertainty and its effects are estimated using a single model so to ensure internal consistency. Under suitable assumptions, the procedure is equivalent to using the square of the VAR forecast error as an external instrument in a proxy SVAR. The procedure allows to add orthogonality constraints to the standard proxy SVAR identification scheme. The method is applied to a US data set; results show that macroeconomic uncertainty is responsible of a large fraction of business-cycle fluctuations while financial uncertainty plays a modest role.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 61-80"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"84966921","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-06-08DOI: 10.1016/j.ecosta.2023.06.002
Ioannis Kalogridis
A novel family of multivariate robust smoothers based on the thin-plate (Sobolev) penalty that is particularly suitable for the analysis of spatial data is proposed. The proposed family of estimators can be expediently computed even in high dimensions, is invariant with respect to rigid transformations of the coordinate axes and can be shown to possess optimal theoretical properties under mild assumptions. The competitive performance of the proposed thin-plate spline estimators relative to their non-robust counterpart is illustrated in a simulation study and a real data example involving two-dimensional geographical data on ozone concentration.
{"title":"Robust thin-plate splines for multivariate spatial smoothing","authors":"Ioannis Kalogridis","doi":"10.1016/j.ecosta.2023.06.002","DOIUrl":"10.1016/j.ecosta.2023.06.002","url":null,"abstract":"<div><div>A novel family of multivariate robust smoothers based on the thin-plate (Sobolev) penalty that is particularly suitable for the analysis of spatial data is proposed. The proposed family of estimators can be expediently computed even in high dimensions, is invariant with respect to rigid transformations of the coordinate axes and can be shown to possess optimal theoretical properties under mild assumptions. The competitive performance of the proposed thin-plate spline estimators relative to their non-robust counterpart is illustrated in a simulation study and a real data example involving two-dimensional geographical data on ozone concentration.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 281-293"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"76840859","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-04-28DOI: 10.1016/j.ecosta.2023.04.003
Marcus Mayrhofer, Peter Filzmoser
For the purpose of explaining multivariate outlyingness, it is shown that the squared Mahalanobis distance of an observation can be decomposed into outlyingness contributions originating from single variables. The decomposition is obtained using the Shapley value, a well-known concept from game theory that became popular in the context of Explainable AI. In addition to outlier explanation, this concept also relates to the recent formulation of cellwise outlyingness, where Shapley values can be employed to obtain variable contributions for outlying observations with respect to their “expected” position given the multivariate data structure. In combination with squared Mahalanobis distances, Shapley values can be calculated at a low numerical cost, making them an even more attractive tool for outlier interpretation. Simulations and real-world data examples demonstrate the usefulness of these concepts.
{"title":"Multivariate outlier explanations using Shapley values and Mahalanobis distances","authors":"Marcus Mayrhofer, Peter Filzmoser","doi":"10.1016/j.ecosta.2023.04.003","DOIUrl":"10.1016/j.ecosta.2023.04.003","url":null,"abstract":"<div><div>For the purpose of explaining multivariate outlyingness, it is shown that the squared Mahalanobis distance of an observation can be decomposed into outlyingness contributions originating from single variables. The decomposition is obtained using the Shapley value, a well-known concept from game theory that became popular in the context of Explainable AI. In addition to outlier explanation, this concept also relates to the recent formulation of cellwise outlyingness, where Shapley values can be employed to obtain variable contributions for outlying observations with respect to their “expected” position given the multivariate data structure. In combination with squared Mahalanobis distances, Shapley values can be calculated at a low numerical cost, making them an even more attractive tool for outlier interpretation. Simulations and real-world data examples demonstrate the usefulness of these concepts.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 179-199"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"86773664","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-07-01Epub Date: 2023-04-14DOI: 10.1016/j.ecosta.2023.04.001
Ana M. Bianco , Graciela Boente
Proposals given in the field of ROC curves focusing on their robust aspects and contributions are considered. The motivation is the extended belief that ROC curves are robust. Without being exhaustive, some recent advances in the area are mentioned. The attention is placed on those situations where the presence of covariates related to the diagnostic marker may increase the discriminating power of the ROC curve. Recent robust procedures given in the framework of the induced methodology are extended to the situation where functional covariates are also present. Consistency results for this proposal are derived under mild conditions. The reported numerical study illustrates that the robust estimators of the covariate specific ROC curve improve the performance of the classical ones for contaminated samples.
{"title":"Addressing robust estimation in covariate–specific ROC curves","authors":"Ana M. Bianco , Graciela Boente","doi":"10.1016/j.ecosta.2023.04.001","DOIUrl":"10.1016/j.ecosta.2023.04.001","url":null,"abstract":"<div><div>Proposals given in the field of ROC curves focusing on their robust aspects and contributions are considered. The motivation is the extended belief that ROC curves are robust. Without being exhaustive, some recent advances in the area are mentioned. The attention is placed on those situations where the presence of covariates related to the diagnostic marker may increase the discriminating power of the ROC curve. Recent robust procedures given in the framework of the induced methodology are extended to the situation where functional covariates are also present. Consistency results for this proposal are derived under mild conditions. The reported numerical study illustrates that the robust estimators of the covariate specific ROC curve improve the performance of the classical ones for contaminated samples.</div></div>","PeriodicalId":54125,"journal":{"name":"Econometrics and Statistics","volume":"39 ","pages":"Pages 353-370"},"PeriodicalIF":3.0,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"80124976","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}