Autocorrelation test under frequent mean shifts
September 2026
Testing for the presence of autocorrelation is a fundamental problem in time series analysis. Classical methods such as the Box–Pierce test rely on the assumption of stationarity, necessitating the...
Distributed inference for high-dimensional convoluted rank regression
September 2026
Convoluted rank regression is recently developed as a powerful tool to deal with outliers and heavy-tailed noise data. However, there is still a lack of suitable methods for convoluted rank regression...
Network perturbation aggregation for graphon estimation
September 2026
In recent years, various methods have been proposed to estimate the edge probability under the graphon model given a single observed network. Since the presence or absence of edges in the observed network...
Statistics in the next quarter-century: Playing also in the frontyard?
September 2026
Statistics has undergone remarkable development over the past decades while interacting increasingly with Data Science, Machine Learning, and Artificial Intelligence (AI). This perspective discusses...
Sublinearly structured deep neural networks achieve feature learning consistency for compositional functions
September 2026
Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical...
Prompt Perturbation for reliable LLM evaluation over comparison graphs
Available online 14 August 2026
Evaluating large language models (LLMs) is important for understanding their capabilities, comparing competing systems, and supporting the deployment of reliable models in practice. For open-ended tasks,...
Privacy-preserving reinforcement learning from human feedback via decoupled reward modeling
Available online 13 August 2026
Preference-based fine-tuning has become an important component in training large language models, and the data used at this stage may contain sensitive user information. A central question is how to...
AdAdaGrad: Adaptive batch size schemes for adaptive gradient methods
Available online 11 August 2026
The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-scale model training. Although large-batch training is...
Linear discriminant analysis with high-dimensional mixed variables
Available online 30 July 2026
Datasets containing both categorical and continuous variables are frequently encountered in many areas. The dimensions of these variables can be very high especially in modern data analysis. Despite...
Inductive Node-split Cross-Validation in networks
Available online 29 July 2026
In network literature, cross-validation (CV) has been used for community detection via node or edge splitting, but with half-consistency that can only prevent underfitting. While penalized CV can rescue,...
Quantitative analysis of rightmost eigenvalue for a large chiral non-Hermitian random matrix
Available online 22 July 2026
This paper provides a quantitative analysis of the rightmost eigenvalue for a chiral non-Hermitian random matrix in the maximally non-Hermitian regime (τ=0). Let (σi)1≤i≤n be the eigenvalues with positive...
LSD of sample covariances of superposition of matrices with separable covariance structure
Available online 15 July 2026
We study the asymptotic behavior of the spectra of matrices of the form Sn=1nXX∗, where X=∑r=1KXr and Xr=Ar12ZrBr12, K∈N. Here, {Zr:r∈[K]} are p×n matrices containing zero-mean, unit-variance innovation...
Non-asymptotic analysis of median-of-means estimation for high-dimensional time series
Available online 7 July 2026
This study addresses the challenges in estimating mean vectors and autocovariance matrices in modern data settings, which are often affected by three key issues: high-dimensionality, heavy-tailed distributions,...
Varying coefficient tensor regression
Available online 3 June 2026
We propose a new varying coefficient model for tensor data regression analysis. To manage the complexity of multi-dimensional tensors, we first employ a tensor partitioning strategy to reduce dimensionality,...
Robust spectral watermark for synthetic tabular data
Available online 28 April 2026
The rise of generative AI has enabled the production of high-fidelity synthetic tabular data across fields such as healthcare, finance, and public policy, raising growing concerns about data provenance...