Most Downloaded Articles

Open access

ISSN: 3051-3901

Robust spectral watermark for synthetic tabular data

The rise of generative AI has enabled the production of high-fidelity synthetic tabular data across fields such as healthcare, finance, and public policy, raising growing concerns about data provenance...

Statistics in the next quarter-century: Playing also in the frontyard?

Statistics has undergone remarkable development over the past decades while interacting increasingly with Data Science, Machine Learning, and Artificial Intelligence (AI). This perspective discusses...

Autocorrelation test under frequent mean shifts

Testing for the presence of autocorrelation is a fundamental problem in time series analysis. Classical methods such as the Box–Pierce test rely on the assumption of stationarity, necessitating the...

Varying coefficient tensor regression

We propose a new varying coefficient model for tensor data regression analysis. To manage the complexity of multi-dimensional tensors, we first employ a tensor partitioning strategy to reduce dimensionality,...

Network perturbation aggregation for graphon estimation

In recent years, various methods have been proposed to estimate the edge probability under the graphon model given a single observed network. Since the presence or absence of edges in the observed network...

Sublinearly structured deep neural networks achieve feature learning consistency for compositional functions

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical...

Distributed inference for high-dimensional convoluted rank regression

Convoluted rank regression is recently developed as a powerful tool to deal with outliers and heavy-tailed noise data. However, there is still a lack of suitable methods for convoluted rank regression...

Non-asymptotic analysis of median-of-means estimation for high-dimensional time series

This study addresses the challenges in estimating mean vectors and autocovariance matrices in modern data settings, which are often affected by three key issues: high-dimensionality, heavy-tailed distributions,...

LSD of sample covariances of superposition of matrices with separable covariance structure

We study the asymptotic behavior of the spectra of matrices of the form Sn=1nXX∗, where X=∑r=1KXr and Xr=Ar12ZrBr12, K∈N. Here, {Zr:r∈[K]} are p×n matrices containing zero-mean, unit-variance innovation...

Quantitative analysis of rightmost eigenvalue for a large chiral non-Hermitian random matrix

This paper provides a quantitative analysis of the rightmost eigenvalue for a chiral non-Hermitian random matrix in the maximally non-Hermitian regime (τ=0). Let (σi)1≤i≤n be the eigenvalues with positive...

Linear discriminant analysis with high-dimensional mixed variables

Datasets containing both categorical and continuous variables are frequently encountered in many areas. The dimensions of these variables can be very high especially in modern data analysis. Despite...

Inductive Node-split Cross-Validation in networks

In network literature, cross-validation (CV) has been used for community detection via node or edge splitting, but with half-consistency that can only prevent underfitting. While penalized CV can rescue,...

Stay Informed

Register your interest and receive email alerts tailored to your needs. Sign up below.