arXiv:2609.26951v1 Announce Type: new
Abstract: We introduce Rolling Conformal Prediction (rolling-CP), a distribution-free predictive inference method for the setting of sequential model training. Specifically, given a data stream $(X_1,Y_1),(X_2,Y_2),\dots$, at each time $n$ the trained model may depend on the observed history $\{(X_i,Y_i)\}_{i
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
oai:arXiv.org:2609.26978v1
arXiv:2609.26978v1 Announce Type: new
Abstract: We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $\Omega(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.
CVaR anchor regression protects against rare shifts
oai:arXiv.org:2609.27034v1
arXiv:2609.27034v1 Announce Type: new
Abstract: We study prediction in new environments when training data contain rare, large shifts. Anchor regression penalizes the average of the squared mean residual across environments. It protects against shifts in an ellipsoid determined by the second moment of the training shifts. Covering rare shifts may therefore require a large penalty, expanding the ellipsoid in every direction and reducing accuracy on common environments. We propose CVaR anchor regression, which replaces the average of the squared mean residuals with a tail average. Unlike CVaR or GroupDRO applied directly to prediction risks, it does not give environments more weight solely because their noise levels are high. We prove an exact worst-case risk guarantee under a linear structural model that allows for heteroscedastic noise. For discrete environments, decreasing the CVaR tail fraction expands the robustness set from an ellipsoid to a scaled convex hull of the training shifts and their negatives. A separate parameter controls its scale. Examples show how the method can improve protection against rare shifts while retaining accuracy on common environments. We illustrate the method on New York City taxi data.
Formulating Cross-World Mediation Estimands Through Single-World Mixtures
oai:arXiv.org:2609.27075v1
arXiv:2609.27075v1 Announce Type: new
Abstract: In causal mediation analysis, natural direct and indirect effects are defined through nested counterfactuals that combine the outcome under one exposure level with the mediator value under another and are therefore inherently cross-world. Their canonical identification additionally relies on cross-world independence assumptions. Consequently, both the estimands and their identifying assumptions remain controversial. In this paper, we ask what additional single-world structure would be required to re-express cross-world estimands as single-world quantities, identifiable under single-world assumptions. We show that these quantities can be written as mixtures of controlled single-world effects under assumptions involving a susceptibility marker, a possibly latent baseline variable that encodes the "would-be" mediator value under a reference treatment arm. If observed, the marker would make several implications of the marker restrictions empirically testable. These results provide a transparent single-world formulation of cross-world mediation effects while making explicit the assumptions required. Although these conditions clarify how cross-world estimands can be interpreted within a single-world framework, their practical relevance depends on whether such susceptibility markers can be justified in specific applications. We discuss implications for principal stratification and show in the Supplementary Material how the formulation extends to ordered mediators and settings with exposure-induced mediator-outcome confounding.
A Bayesian framework for multilevel data under model mis-specification
oai:arXiv.org:2609.27112v1
arXiv:2609.27112v1 Announce Type: new
Abstract: We propose a Bayesian framework for uncertainty quantification from the perspective that the working model is mis-specified in settings of a multilevel data-generating process. We focus on settings in which the mis-specification fails to match the functional form of the mean structure, and discuss Bayesian estimation of target parameters under dependence induced by a mismatch between working and data-generating models. The proposal represents a Bayesian semi-parametric procedure aimed at estimating population-level parameters while accounting for cluster- and unit-level variation in the estimating function. The proposal extends the regular Bayesian bootstrap to account for cluster- and unit-level variation using multilevel weights from an enriched Dirichlet model. Simulation studies indicate that the proposed approach has good frequentist properties when the data-generating process and the proposed model induce a partially exchangeable sequence associated with the unknown quantity of interest. Applications to radon \citep{gelman2007data}, Programme for International Student Assessment 2022 \citep{OECD2023PISA}, and tuberculosis \citep{nobre2023impact} datasets are presented for illustrative purposes. The results demonstrate that the proposed method is competitive with variations of multilevel models, with major differences observed in the range of credible intervals, which are justified by the nonparametric assumptions underlying the proposed method.
Bayesian Nonparametric Approaches to Ordinal Drought Modeling in the United States
oai:arXiv.org:2609.27135v1
arXiv:2609.27135v1 Announce Type: new
Abstract: Data observed over space and time can exhibit dependence across both dimensions, and datasets can grow very large in both the number of locations and the number of time periods. This dependence and data size mean that many statistical models cannot be fit at a reasonable computational cost as a result of dense matrix inversions, large parameter spaces, or memory and storage challenges. When the data are ordinal, this only adds to the computational complexity of model fitting. Many ordinal models rely on a latent continuous variable designed to capture dependence in a computationally efficient manner, and then partitioned according to cutoff parameters to yield the observed ordinal response data. A very common choice is a latent Gaussian distribution, which can accommodate different dependence structures and permits Gibbs sampling in a Bayesian framework. Unfortunately, this model can be overly restrictive and fails to provide the flexibility needed to capture a range of ordinal outcomes. In this paper, we demonstrate the use of Bayesian nonparametric (BNP) methods to enhance the model flexibility of a Gaussian latent model for ordinal drought data observed over space and time, with Dirichlet process priors inducing clustering among time periods within each spatial location. In this work, we model ordinal drought data separately at each spatial location while accounting for temporal dependence, but we do not model spatial dependence across locations. We show that these BNP models often outperform Bayesian parametric approaches at a reasonable computational cost.
Elastic Multi-Fidelity Bayesian Model Calibration
oai:arXiv.org:2609.27174v1
arXiv:2609.27174v1 Announce Type: new
Abstract: Bayesian calibration of functional-output computer models typically relies on dimension reduction techniques, such as functional principal component analysis, which assume that differences among simulator realizations arise only from amplitude variation. When simulator output also exhibits phase variation such as shifts in the timing or location of key features, this assumption is violated. Recent work has addressed this issue through elastic calibration, which aligns functional computer model realizations with observed experimental data prior to dimension reduction. Separately, multi-fidelity methods reduce the cost of calibration by supplementing a small number of expensive high-fidelity simulator runs with a larger ensemble of cheap low-fidelity runs. This is typically done through either a mapping strategy, which corrects low-fidelity predictions toward high-fidelity output, or a fusion strategy, which builds a shared basis across both fidelities. This paper combines these two lines of work, introducing elastic multi-fidelity Bayesian model calibration, which aligns high- and low-fidelity functional output to a common reference before applying multi-fidelity mapping or fusion. On a synthetic two-dimensional calibration problem and a dynamic material properties equation-of-state problem, both elastic multi-fidelity strategies match or improve on the leave-one-out predictive accuracy of a mono-fidelity elastic emulator, with the fusion approach achieving the lowest error. Both strategies also produce tighter calibrated posteriors than the mono-fidelity baseline, with the fusion approach providing the best coverage and parameter estimates closest to the true values.
Change detection with conformal martingales: new optimal constructions, and suboptimality of existing methods
oai:arXiv.org:2609.27179v1
arXiv:2609.27179v1 Announce Type: new
Abstract: We study distribution-free sequential changepoint detection for independent observations with unknown and unrestricted pre- and post-change laws. We build on the conformal test martingales and associated e-detectors of Vovk(2021), which control the probability of false alarm (PFA) and the average run length (ARL) respectively. The majority of these works focus on validity, with statistical efficiency usually left for simulations. We develop a comprehensive theory of how conformal p-values behave under non-exchangeable data with a changepoint at an unknown time $T$. We use this to analyze the post-change growth and resulting detection delay of conformal martingale methods, and prove that the standard existing methods are suboptimal for PFA and ARL control, and can lead to delays that are $\Omega(T)$ and $\Omega(\sqrt{\text{ARL}})$ respectively. We propose different conformal e-processes and e-detectors that are provably minimax optimal, with delays $\Theta(\log T)$ and $\Theta(\log \text{ARL})$ respectively, and have much shorter delays in simulations.
Artificial intelligence surrogates for treatment effect estimation with before-and-after data
oai:arXiv.org:2609.27180v1
arXiv:2609.27180v1 Announce Type: new
Abstract: Estimating the causal effects of medical treatments is difficult when clinically important outcomes are costly to measure or require long follow-up. Short-term or inexpensive surrogate outcomes offer a potential alternative, but surrogate biomarkers may be unavailable or difficult to identify. Advances in artificial intelligence (AI) have enabled increasingly accurate prediction of clinical outcomes from inexpensive, high-dimensional measurements, which creates an opportunity to use AI predictions themselves as surrogates. To this end, we develop a framework for estimating treatment effects from paired measurements obtained before and after treatment for each treated individual. A pretrained AI model is applied to the before and after measurements, and our estimator compares the resulting outcome predictions. We characterize the technical assumptions under which this within-person contrast identifies the average treatment effect on the treated, even when clinical outcomes are never observed for treated individuals. When these assumptions cannot be justified, we use prediction-powered inference to correct bias using a small number of observed clinical outcomes and obtain valid inference. Synthetic and real-world cardio-oncology experiments demonstrate the validity and accuracy of the approach.
Model Specification Test for Stationary Functional Time Series
oai:arXiv.org:2609.27182v1
arXiv:2609.27182v1 Announce Type: new
Abstract: We develop a general framework for model specification testing in stationary functional time series. The approach is based on an autoregressive approximation that represents a broad class of stationary functional processes through coefficient kernels whose dimension and autoregressive order may increase with the sample size. Different model assumptions induce different structural restrictions on these kernels, and our tests are constructed by measuring deviations from the corresponding restrictions. We illustrate this principle for three problems: testing a prescribed order of a functional autoregressive model, testing a functional autoregressive moving-average specification, and testing separability of autoregressive coefficient kernels. The resulting statistics are based on weighted $\mathcal{L}^2$-distances, and the critical values are obtained by a multiplier bootstrap. We establish a quantitative bootstrap approximation that is uniform over a class of weight functions and prove asymptotic validity and consistency of the proposed tests. The methodology allows for data-adaptive weighting and is illustrated by simulations and a data example.
Local Optimality and Rigidity of Frobenius Tests for Dense High-Dimensional Covariance Alternatives
oai:arXiv.org:2609.27200v1
arXiv:2609.27200v1 Announce Type: new
Abstract: We study identity testing for high-dimensional covariance matrices against dense alternatives of unknown direction, with $p/n \to \gamma$. Along a globally positive quadratic precision path, mixing Gaussian alternatives over a Gaussian Orthogonal Ensemble direction yields a contiguous experiment whose log likelihood reduces to the corrected Frobenius statistic; its upper-tail test attains the limiting weighted-power envelope at every fixed strength. Fixing the prior's Frobenius radius perturbs the mixture by only $O(p^{-1/2})$ in total variation, and exact whitening carries the experiment, the statistic, and its null law to any known null covariance. Separately, under a product-coordinate null, feasibility needs only $4+\eta$ moments, plus identical distributions over time when means are estimated; studentization and an exact degrees-of-freedom correction preserve the local power. A stability inequality turns near-envelope attainment into null agreement with the Frobenius rule, so uniform noninferiority on the typical dense bulk precludes gains at any contiguous alternative. For trace-matched rank-one alternatives, the corrected statistic is the first likelihood direction when $\vartheta_n \to 0$ and $n\vartheta_n \to \infty$; at fixed strength, the log likelihood ratio in the Onatski-Moreira-Hallin fixed-spike benchmark is governed by a richer linear spectral statistic below the Baik-Ben Arous-Peche threshold, while eigenvalue separation permits cost-free largest-eigenvalue enhancement above it. Simulations illustrate the theory.
Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
oai:arXiv.org:2609.27206v1
arXiv:2609.27206v1 Announce Type: new
Abstract: Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
On the Sample Complexity of Active Learning with Membership Queries
oai:arXiv.org:2609.27241v1
arXiv:2609.27241v1 Announce Type: new
Abstract: This work revisits a fundamental question in active learning: how powerful is the ability to synthesize arbitrary queries? Compared to pool-based active learning, where the learner only selects queries from a given unlabeled pool, we find that this seemingly mild change in query ability may dramatically alter the difficulty of statistical learning. In particular, some hypothesis classes that are inherently slow to learn in the pool-based setting, achieving only polynomial error decay in the number of samples, become exponentially learnable once synthesized queries are allowed. This striking gap suggests that membership query synthesis induces a fundamentally different mode of learning, one that is not adequately captured by existing active learning theory and calls for new analytical tools to characterize its complexity. Motivated by this phenomenon, we develop several sufficient conditions, present intriguing examples, and propose a conjectural perspective toward understanding which hypothesis classes admit efficient learning through synthesized queries.
From Metrics to Decisions in NBA Analytics: A Critical Integrative Review and Decision-Readiness Framework
oai:arXiv.org:2609.27245v1
arXiv:2609.27245v1 Announce Type: new
Abstract: National Basketball Association (NBA) teams have increasingly detailed metrics, but better predictions do not necessarily improve decisions. This critical integrative review draws on prior reviews, citation tracing, and topic searches across seven research streams: on-court action, player value, role, lineup synergy, availability, draft and development, and contracts and roster construction. An observation-state-action-decision-evaluation chain organizes the synthesis. Six decision-readiness gates guide our assessment: point-in-time validity, uncertainty, context portability, action feasibility, opportunity-set observability, and evaluation, with requirements matched to each claim. The reviewed literature is strongest in measuring and predicting individual components of a decision. Evidence is less developed at interfaces that combine components, transfer them across settings, and compare feasible actions. We outline a proposed deployment workflow, a reporting contract, and a research agenda covering player transport, role substitution, roster fragility, legal action generation, and asset valuation. The 2023 collective bargaining agreement and forthcoming 3-2-1 Draft Lottery illustrate how institutional changes generate research questions. Models should inform evaluable comparisons of feasible choices. While its effect on organizational decision quality remains an empirical question, the framework provides a diagnostic and reporting structure for matching decision claims to evidence requirements.
Functional Causal Discovery via Conditional Covariance Ordering
oai:arXiv.org:2609.27256v1
arXiv:2609.27256v1 Announce Type: new
Abstract: We study causal discovery where each node is a random function. Previous studies on this topic rely on structural assumptions, e.g., linearity or non-linearity, and distributional assumptions, e.g., Gaussianity or non-Gaussianity. In contrast, we make use of covariance operators to avoid these assumptions. Under functional additive noise models, we propose a new sufficient condition to identify a valid topological ordering based on comparing norms of conditional covariance operators. Taking advantage of this identifiability condition, we develop a new mixed regression model that subsumes linear and non-linear models. Together with variable selection, our procedure yields an estimation of the causal directed acyclic graph (DAG) for functional variables. In theory, we develop the least-squares-type theory of this regression model, and derive asymptotic consistency of order determination, sparse regression, as well as identifying the DAG. Computational algorithms based on discrete observations are provided. Applied to simulated data, our approach performs satisfactorily among existing approaches. A real data example of brain effective connectivity is also presented.
Multitask Regression with Pairwise Fusion
oai:arXiv.org:2609.27280v1
arXiv:2609.27280v1 Announce Type: new
Abstract: We study multitask regression when coefficient sharing can differ by predictor. For a given predictor, many tasks may have the same coefficient while a few differ, and the exceptional tasks need not be the same for another predictor. We describe this structure by two quantities: the number of active predictors and the total number of task coefficients that differ from the most common value for their predictor. We estimate the coefficient matrix by penalizing all pairwise coefficient differences across tasks, with an additional group penalty when predictor selection is needed. The resulting upper and lower bounds have the same dependence on these two quantities. We also consider the stronger setting in which a large set of tasks shares one entire coefficient vector. Under explicit sample-size conditions, the same pairwise estimator pools those tasks exactly, while allowing the remaining tasks to differ. Simulations and household energy data illustrate the transition between broad sharing and task-specific coefficients.
Bayesian calibration of adaptive-behavior SIR models for multi-wave COVID-19 incidence in New York City
oai:arXiv.org:2609.27283v1
arXiv:2609.27283v1 Announce Type: new
Abstract: Epidemic incidence reflects both transmission dynamics and adaptive human behavior, yet these mechanisms may be difficult to distinguish from aggregate case data alone. We calibrated four susceptible--infected--recovered (SIR) specifications to weekly confirmed COVID-19 incidence in New York City from June to December 2020, comparing a single continuous SIR trajectory, a wave-initialized SIR model, and two adaptive-behavior models with either shared or wave-specific transmission. Inference was performed using rejection Approximate Bayesian Computation (ABC), and in-sample reconstruction was assessed using root mean squared error (RMSE) and the weighted interval score (WIS). Reinitializing the epidemic state by wave produced the largest structural improvement over the continuous SIR trajectory, reducing mean-based RMSE by 48.6\% and WIS by 17.6\%. Adding delayed prevalence-dependent behavioral adaptation with shared transmission further reduced mean-based RMSE by 23.4\%, but yielded essentially unchanged WIS relative to the wave-initialized SIR model. Allowing transmission to vary by wave did not provide a consistent additional advantage and produced strongly asymmetric posterior-simulation trajectories. Behavioral sensitivity, response midpoint, and delay remained only weakly to partially identified. The clearest posterior structure was a negative association between transmission intensity and the behavioral midpoint, indicating that higher transmission could be compensated by behavioral responses activated at lower prevalence. Sensitivity to ordered behavioral priors further showed that reconstruction and behavioral inference depend materially on structural prior assumptions. These results suggest that adaptive mechanisms can improve multi-wave incidence reconstruction, while aggregate incidence alone is insufficient to sharply separate transmission from behavioral adaptation.
Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS
oai:arXiv.org:2609.27299v1
arXiv:2609.27299v1 Announce Type: new
Abstract: Information systems (IS) researchers increasingly use quasi-experimental methods such as difference-in-differences (DiD) and instrumental variables (IV) to recover causal effects from observational panel data. Power calculations that justify these designs assume i.i.d. errors, but the deeper problem is what even a cluster-robust calculator cannot see. We report a Monte Carlo study over 9837 parameter conditions (approx 9.8 million datasets) and decompose the planned-versus-achieved power gap. The serial-correlation component is recoverable by an AR(1)-aware calculator when rho is known, and partially when rho must be estimated from short pre-periods, but panel attrition, staggered-adoption bias, and parallel-trends pretesting are captured by no closed-form formula; exogenous attrition alone costs approx 8 to 11 percentage points at the few-hundred-to-thousand sample sizes IS studies use. Treatment-correlated, outcome-dependent attrition instead induces bias, not just power loss. For IV, holding first-stage F fixed, larger N neither raises power nor curbs exclusion bias, though with a fixed instrument more data does sharpen the first stage, so identification rests on instrument strength, not sample size.
Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel
oai:arXiv.org:2609.27384v1
arXiv:2609.27384v1 Announce Type: new
Abstract: A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.
Bayesian inference, on-line forecasting and model choice for large VAR models with Cholesky stochastic volatility
oai:arXiv.org:2609.27431v1
arXiv:2609.27431v1 Announce Type: new
Abstract: We consider $K$-dimensional Bayesian vector autoregressions (BVARs) with Cholesky stochastic volatility (SV), in which the innovation covariance matrix is a lower-triangular linear transform of $K$ independent univariate SV processes. Such models are widely used in empirical macroeconomics to capture time-varying uncertainty and improve forecast accuracy, but the cost of posterior simulation is the binding constraint on the size of the system, and a major bottleneck for empirical work. We introduce a Markov chain Monte Carlo (MCMC) kernel that mixes better than existing samplers at the same computational complexity. It rests on a reparametrisation that makes the $K$ volatility trajectories conditionally independent, and on Particle Gibbs to update each trajectory. The kernel targets the exact posterior, rather than an approximation of it, and in an application with $K = 15$ it raises the mean effective sample size per second, relative to the benchmark corrected triangular algorithm, by a factor of approximately $14$ for the VAR coefficients and $3.4$ for the volatilities. We also introduce a a Sequential Monte Carlo squared (\smcsq{}) sampler, which used our MCMC kernel as a building block, and which delivers at every $t$ the one-step-ahead predictive density and the marginal likelihood of the data up to $t$, and hence on-line forecasting and model choice. To our knowledge, this is the first algorithm that delivers sequential marginal likelihoods for Cholesky-SV BVARs with static contemporaneous coefficients and a non-conjugate prior. We illustrate both on US monthly macroeconomic data, with $K=15$ for posterior inference and $K=6$ for the sequential sampler and model choice.
Short-term rental market occupancy - daily time series for 2017-2022 on 500 markets worldwide
oai:arXiv.org:2609.27445v1
arXiv:2609.27445v1 Announce Type: new
Abstract: Short-term vacation rentals, as promoted by platforms such as Airbnb, Homeaway, Vrbo, etc., are a growing component of the travel industry. This paper provides a unique, large dataset on global market occupancy for the short-term rental market using data from the American company Wheelhouse. The dataset consists of data for $500$ markets around the world. For each market, a daily occupancy time series from January $2017$ to December $2022$ is provided, allowing for studies of local and global patterns in the evolution of the short-term rental market. Additionally, the dataset includes curves representing the booking trajectory of each market and stay date up to one year prior to the stay date. This large dataset comprises a unique combination of time series and survival analysis data, and is suitable as a methodological benchmark for both classical statistical and machine learning models.
Pooling Sequential Evidence Across Hypotheses: Rate-Optimal Multiple Testing at a Fixed Horizon
oai:arXiv.org:2609.27465v1
arXiv:2609.27465v1 Announce Type: new
Abstract: We study sequential testing of a fixed family of hypotheses when observations are costly and a sampling horizon is specified in advance. The challenge is to pool evidence for earlier decisions when the number and identities of false hypotheses are unknown, while controlling the probability of any false rejection at level $\alpha$. Existing merges attain the pooled growth rate at a single number of false hypotheses: averaging when one is false, multiplying when all are. Under an independent-stream model with common simple null and alternative distributions, we test each intersection with a prior-weighted mixture of products of marginal likelihood ratios. Closed testing combines these elementary-symmetric-polynomial mixtures to identify individual false hypotheses. Design-specific boundary calibration gives finite-horizon family-wise error control, with exact finite-state guarantees or a confidence qualification for Monte Carlo calibration. The prior-matched mixture uniquely maximizes expected log evidence at each horizon. Mixtures assigning positive weight to every nonempty subset of streams attain log-growth rate $lD$ when $l$ streams follow the alternative. Here $D$ is the mean log likelihood ratio per alternative observation, and a round supplies one observation per stream. This rate attains the first-order intersection-delay lower bound as $\alpha\downarrow0$ at fixed dimension, configuration, weights, and a long enough horizon. Power for an individual hypothesis cannot exceed the best single-stream power at a given deadline, but closure removes the multiplicity penalty when all are false. Gaussian, basket-trial, language-model and advertising studies illustrate both. The primary basket boundaries are 40-52% below $1/\alpha$. Across 41 simulated configurations, the median reduction in capped mean patient outcomes relative to prespecified interim-look Bonferroni tests is 31%.
A two-step log-linear procedure for graphical representation and inference of associations in cross-classified data for disease diagnosis
oai:arXiv.org:2609.27491v1
arXiv:2609.27491v1 Announce Type: new
Abstract: Biometrical sciences and disease diagnosis in particular, are often concerned with the analysis of associations for cross-classified data, for which distance association models give us a graphical interpretation for non-sparse matrices with a low number of categories. In this framework, usually binary exploratory and response variables are present, with analysis based on individual profiles being of great interest. For saturated models, we show the usual linear relationship for log-linear models is preserved in full dimension for the distance association parameterization. This enables a two-step procedure to facilitate the analysis and the interpretation of associations in terms of unfolding after the overall and main effects are removed. The proposed procedure can deal with cross-classified data for profiles by binary variables, and it is easy to implement using traditional statistical software. For disease diagnosis, the problems of a degenerate solution in the unfolding representation, and that of determining significant differences between the profile locations are addressed. A hypothesis test of independence based on odds ratio is considered. Furthermore, a procedure is proposed to determine the causes of the significance of the test, avoiding the problem of error propagation. The equivalence between a test for equality of odds ratio pairs and the test for equality of location for two profiles in the unfolding representation in the disease diagnosis is shown. The results have been applied to a real example on the diagnosis of coronary disease, relating the odds ratios with performance parameters of the diagnostic test.
Semiparametric Inference for Dynamic Causal Effects from Observational Time Series
oai:arXiv.org:2609.27531v1
arXiv:2609.27531v1 Announce Type: new
Abstract: In observational time series, statistical inference for dynamic causal effects of a one-time intervention across horizons is complicated by high-dimensional observed pre-treatment information, unmeasured confounding, and serial dependence. To address these challenges, we develop a semiparametric framework for inference from a single serially dependent time series, integrating debiased machine learning with instrumental variables through buffered block cross-fitting. Under geometric beta-mixing, we derive non-asymptotic bounds on estimation error, asymptotic normality at each fixed horizon, and feasible inference that accommodates serial dependence. We further show how learner-specific prediction guarantees under temporal dependence can be used to verify the nuisance-rate conditions required for orthogonal inference. In a monetary-policy application with 468 months and 1464 lagged FRED-MD controls, we show an instrumented policy tightening lowers housing starts at medium horizons, with sensitivity analyses that support the finding.
Robustness of Diffusion Models under Distribution Shift
oai:arXiv.org:2609.27546v1
arXiv:2609.27546v1 Announce Type: new
Abstract: Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution, yet existing theoretical guarantees largely focus on the no-shift setting. In this work, we study robust score estimation under Wasserstein perturbations of a reference distribution. For the Ornstein--Uhlenbeck diffusion, we show that robust estimation decomposes into two fundamental components: the statistical cost of learning the reference distribution and the intrinsic cost of distribution shift. The latter scales quadratically with the Wasserstein radius, and this dependence is minimax optimal. We construct an explicit finite-sample estimator achieving the resulting robust minimax rate without knowing the shift radius. When the reference distribution lies on an unknown low-dimensional subspace, the statistical term adapts to the intrinsic dimension while the shift cost remains unchanged. Finally, we show that the same decomposition governs positive-time reverse sampling and obtain matching minimax guarantees in KL divergence. Together, these results characterize how finite data, intrinsic dimension, and distribution shift affect the robustness of score-based diffusion models.
Model-agnostic noise reduction for high-dimensional time series data
oai:arXiv.org:2609.27614v1
arXiv:2609.27614v1 Announce Type: new
Abstract: We develop a model-agnostic framework for noise reduction in high-dimensional time series that explicitly targets optimal recovery of a low-dimensional latent dynamic component contaminated by observational white noise. Under the assumption that the latent dynamics live in a low-dimensional linear dynamic subspace, we characterize the optimal linear projection onto the dynamic subspace and provide a geometric description of the residual error in terms of the relative orientation of the signal and noise spaces. We propose estimators for the dynamic subspace and the optimal projection based on lagged covariance matrices, bootstrap dimension selection, and a low-rank representation of the structured noise. Under mild conditions, the resulting denoised series is shown to converge to its population target at the usual parametric rate. Simulations show that the proposed method can substantially improve subspace estimation, reconstruction error, and one-step-ahead forecast accuracy compared with both orthogonal projection-based denoising and the raw data. The approach is illustrated by empirical applications to high-dimensional stock returns and to a 20-variate time series of macroeconomic indicators.
Computer code validation via mixture model estimation
oai:arXiv.org:2609.27619v1
arXiv:2609.27619v1 Announce Type: new
Abstract: When computer codes model complex physical systems, calibration alone is insufficient; one must also assess whether a discrepancy term is needed. In this paper, we study computer code validation through a Bayesian mixture approach that compares a pure-code model with a discrepancy-corrected model. The method relies on the posterior distribution of a mixture weight, which measures the relative support of the two competing distributions. Under the assumption that the code is linear in the calibration parameters, or can be well approximated by a linear surrogate, we show that mixture component-shared parameters can be used to combine flexible modeling, even when noninformative priors are assigned to some common parameters. Inference is performed using a Metropolis-within-Gibbs algorithm. In addition, we introduce a thresholded allocation rule that complements the global mixture weight by providing a local diagnostic of where the discrepancy-corrected component is truly needed along the input domain. Beyond global model comparison, the proposed approach is also able to perform local model discrimination by identifying where the discrepancy-corrected component is truly needed along the input domain.
Rank weighting and asymmetry in Blest's rank correlation: two exact regions
oai:arXiv.org:2609.27634v1
arXiv:2609.27634v1 Announce Type: new
Abstract: Blest's rank correlation $\nu$ is a variant of Spearman's rho $\rho$ that weights the leading ranks of one variable more heavily, at the price that $\nu$ is not symmetric in its arguments. We quantify both features by determining the exact region of $(\rho,\nu)$ over all bivariate copulas, as well as that of $(\eta,\nu)$, where $\eta$ is the symmetrized Blest coefficient of Genest and Plante. The latter region is a linear image of the set of all pairs $(\nu(C),\nu(C^\top))$ formed by a copula $C$ and its transpose. Consequently, Blest's coefficient differs from Spearman's rho by at most $1/4$, and interchanging the two variables changes it by at most $27/64$, improving on the bound $1/2$ implied by the first inequality. For every given value of $\rho$ or $\eta$, each corresponding extreme value of $\nu$ is attained by exactly one copula, given in closed form and supported on finitely many line segments. Near countermonotonicity, the upper extremizers of the $(\eta,\nu)$-region are supported on the graph of a function of the second coordinate, yet their conditional laws given the first coordinate carry two atoms. The proofs rest on a rearrangement inequality with equality case and on explicit Kantorovich potentials.
FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints
oai:arXiv.org:2609.27654v1
arXiv:2609.27654v1 Announce Type: new
Abstract: Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.
The Type-II Error of Test Supermartingales: e-Power versus the Chernoff-Stein Exponent
oai:arXiv.org:2609.27765v1
arXiv:2609.27765v1 Announce Type: new
Abstract: In safe hypothesis testing with test supermartingales, Ville's inequality provides anytime-valid type-I error guarantees for every significance level $\alpha\in(0,1]$, if one rejects the null hypothesis whenever the wealth process first exceeds $1/\alpha$. Due to an inherent asymmetry, the type-II error behaves differently. We prove two things about the latter, for a simple null and alternative. First, the mean growth rate $\mathbb{E}_{P_1}[\log E]$, the e-power, that Kelly betting and growth-rate-optimal e-variables maximise, bounds nothing on its own. For every level $c>0$, every $\alpha$ and horizon $t$ we construct e-variables of conditional e-power exactly $c$ whose probability of not rejecting by $t$ is arbitrarily close to one. It forces eventual rejection, but no finite-horizon guarantee follows. Second, the quantity that does control the type-II error is the Chernoff-Stein exponent of an e-variable, $\Lambda(E)=\sup_{s\ge0}\{-\log \mathbb{E}_{P_1}[E^{-s}]\}$, whose range is exactly determined: $\sup_E \Lambda(E)=\mathrm{KL}(P_0\|P_1)$, the classical Chernoff-Stein exponent, and so the ceiling of its own per-e-variable form. One conditional application of Hoelder's inequality per step gives it, for every test supermartingale on an arbitrary filtered space, with no independence or product structure; the i.i.d. case adds that it is matched, and attained by nothing. The e-power has its own ceiling, $\mathrm{KL}(P_1\|P_0)$, and that one is attained, $P_0$-a.s. uniquely, by the likelihood ratio $R$. The two optima are the same divergence in opposite arguments, at opposite ends of the flattened family $R^{\beta}/\mathbb{E}_{P_0}[R^{\beta}]$: the ceiling as $\beta\downarrow0$, $R$ at $\beta=1$. Which $\beta$ is best is settled by the horizon, exactly: $R$ is optimal at $t=\log(1/\alpha)/\mathrm{KL}(P_1\|P_0)$ alone, beaten by sharpening $(\beta>1)$ below it and by flattening above.
Type-II Error Bounds for Test Supermartingales from Lower-Tail Hypotheses
oai:arXiv.org:2609.27766v1
arXiv:2609.27766v1 Announce Type: new
Abstract: In safe hypothesis testing with test supermartingals, Ville's inequality provides anytime-valid type-I error guarantees for every significance level $\alpha\in(0,1]$, if one rejects the null hypothesis whenever the wealth process first exceeds $\frac{1}{\alpha}$. Due to an inherent asymmetry, the type-II error does not have such guarantees: a heavy concentration of the probability on the lower tail of the log-increments can lead to one catastrophic bet that undoes any amount of accumulated evidence. This paper studies how different hypotheses on those lower-tail probabilities lead to different bounds on the type-II error of the sequential test. They all reduce to one master inequality, which bounds the type-II error at level $\alpha$, at a fixed horizon and sequentially, in terms of a one-sided Legendre transform of the (inverse-)moment generating function of the e-variables, evaluated at one number: the amount by which the lower bound of the accumulated e-powers exceeds $\log\frac{1}{\alpha}$. And, the step is lossless, in the sense, that it extracts exactly a constrained information projection. Every bound presented here is a corollary, obtained by a certain majorant of the above function. The hypotheses are: a finite negative moment; an exponentially small crash probability with a moment on the winning side; a wealth floor with a conditional variance, and its Bernstein variant, which interpolates between a Gaussian regime set by the variance and an exponential one set by the scale; a sub-Gaussian or bounded-tilt lower tail; bounded log-increments; and i.i.d. increments, where the majorant is the truth. We also provide an empirical-Bernstein variant. Each hypothesis may either be read as a condition on the e-variables one has, or as the price of betting with an approximation to the likelihood ratio rather than the ratio itself, which satisfies the weakest condition for free.
Surrogate inference for threshold-derived environmental indices depends on null placement and seasonal event concentration
oai:arXiv.org:2609.27769v1
arXiv:2609.27769v1 Announce Type: new
Abstract: Threshold-derived environmental indices convert native-resolution variables into event indicators and aggregated counts, but surrogate tests may impose their null constraints either before or after this transformation. We show that these choices define different inferential targets. A threshold--copula representation separates temporal dependence from the seasonal event-probability vector and makes seasonal concentration an explicit experimental coordinate. In a prespecified benchmark of 13,500 short-memory monthly trajectories, an index-resolution marginal-and-spectrum surrogate and a constrained native-resolution surrogate produced strongly asymmetric paired conclusions: 1084 trajectories rejected only at index resolution, compared with 16 only at native resolution. Holding the expected annual event count fixed while concentrating event probability within the seasonal cycle increased index-only discordance from 4.4\% to 16.2\%. In an independent experiment, equalizing monthly event probability reduced the corresponding rate from 15.9\% to 3.9\%; the mitigation persisted when monthly percentiles were estimated from an independent 30-year reference period. Native-resolution inference was calibrated under the benchmark null but had strongly mechanism-dependent sensitivity: the strongest prespecified history-feedback alternative had a 30.7\% detection rate under the fixed synthetic gate, with substantially lower sensitivity to a persistent-regime alternative. A Coupled Model Intercomparison Project Phase 6 Amazon dry-month case study was non-discriminating, illustrating why null placement and design-specific detectability should be reported together. The results provide a practical framework for surrogate inference on threshold-derived hydroclimatic and environmental indices.
Beyond the Mean: A Weighted k-Sample Omnibus Variance-Ratio Statistic for Covariate Balance Diagnostics
oai:arXiv.org:2609.27772v1
arXiv:2609.27772v1 Announce Type: new
Abstract: Rubin's variance ratio (VR) complements the standardized mean difference by detecting covariate imbalance in spread, but no purpose-built omnibus extension exists for more than two groups. We introduce FVR, a size-weighted quadratic combination of pairwise log-variance-ratios that generalizes VR to k groups, together with geometric-mean and maximum pairwise variance-ratio comparators. FVR has an exact relationship to Rubin's VR at k = 2 and a mathematical structure directly parallel to Cohen's f. In Monte Carlo simulations spanning four variance-driven bias mechanisms, k = 3, 4, 6, and sample sizes from 200 to 10,000, FVR and the geometric-mean statistic generally tracked downstream estimation bias at least as well as the maximum statistic. FVR retained substantially more signal when many groups shared a very small sample, while absolute-bias thresholds were unstable at N = 200. Across mechanisms, FVR < 0.10 was reasonably reassuring, values above approximately 0.30 generally indicated concern, and intermediate values were context-dependent. The statistic is implemented for arbitrary analysis weights in the Stata command varatio and illustrated using a multivalued-treatment application.
Optimal State-Space Order for Spectral Gaps of Sliding-Window Occupation Counts
oai:arXiv.org:2609.27836v1
arXiv:2609.27836v1 Announce Type: new
Abstract: Let $P$ be an irreducible reversible Markov kernel on a $m$-state space $\Omega$, and denote its right spectral gap $\gamma=1-\lambda_2(P)$. From a stationary trajectory, let $K_t$ be the occupation-count vector of the length-$n$ window beginning at time $t$. The stationary pair $(K_0,K_1)$ defines a reversible projected count kernel $\widetilde P_n$. For every $m\ge2$, let \[
c_m^\star=
\inf_{\substack{ P,\; n \ge 2}}
\frac{n\Gap(\widetilde P_n)}{\Gap(P)}. \] We prove \[
\frac1{1080m}\le c_m^\star\le q_{m-2},
\qquad
q_0=\frac14,\quad q_{r+1}=q_r(1-q_r). \] We also show $q_{m-2}=(m+\log m+O(1))^{-1}$, which implies that $c_m^\star=\Theta(m^{-1})$. Thus, the optimal comparison coefficient has order $m$, although its exact value remains open. The lower bound also holds for $n=1$ and is uniform in $P$, including sparse and periodic kernels. Its proof combines short-window decorrelation with an averaged anchor-excursion decomposition, a Green-kernel hitting estimate, and a stopped Carleson--Hardy inequality. A nested rare-state construction produces finite $m$-state witnesses whose normalized Rayleigh quotients approach $q_{m-2}$ through an ordered sequence of limits. For every fixed finite irreducible reversible aperiodic kernel on at least two states, $\Gap(\widetilde P_n)=\Theta_P(n^{-1})$.
Theoretical Study on the Evidential Learning-based Variational Autoencoder
oai:arXiv.org:2609.27853v1
arXiv:2609.27853v1 Announce Type: new
Abstract: A normal--inverse-gamma (NIG) latent hierarchy has four parameters, but its induced latent law does not identify all four. For $\sigma^2\sim\mathrm{InvGamma}(\alpha,\beta)$, $\mu\mid\sigma^2\sim\mathcal{N}(\gamma,\sigma^2/\nu)$, and $z\mid\mu,\sigma^2\sim\mathcal{N}(\mu,\sigma^2)$, the marginal law of $z$ depends on $(\nu,\beta)$ only through $c=\beta(1+1/\nu)$. Hence the reconstruction-visible parameter space is the three-dimensional quotient $(\gamma,\alpha,c)$, with a one-dimensional fiber degree of freedom. For a fixed hierarchical variational objective, exact partial minimization of the forward KL divergence to a complete NIG prior selects a unique prior-relative representative on each fiber, yielding an exact three-coordinate reduction with the same optimum as the four-coordinate objective. Writing $\rho_0=2\beta_0/\nu_0$ and $T=c/\{\alpha[(\gamma-\gamma_0)^2+\rho_0]\}$, we show that inverse canonical allocation $1/\nu_{\rm can}$ is an explicit strictly increasing function of $T$. For $\alpha>1$, the ratio $u_{\rm epi}/u_{\rm var}=1/\nu_{\rm can}$ is therefore determined by the quotient state and prior; for rank-only use under a common calibration, $T$ contains the same coordinatewise ordinal information. The residual prior gauge is characterized rather than eliminated: $(\gamma_0,\rho_0)$ govern ordinal dependence, while $(\nu_0,\alpha_0)$ determine numerical calibration and the analytic ceiling of $1/\nu_{\rm can}$.
Covariance Kernels on Unordered Pair Spaces: Theory and Applications to Network-Valued Data
oai:arXiv.org:2609.27879v1
arXiv:2609.27879v1 Announce Type: new
Abstract: Many scientific problems are relational: the quantity of interest is a connection between two objects, while information about similarity is available for the objects themselves. We develop a covariance framework for unordered relationships that transfers object-level geometry to the relations they form while preserving endpoint identity and invariance to ordering. Building on symmetric pairwise-kernel representations, we develop theory for the loop-free domains used in undirected networks. We establish spectral interlacing and trace-loss results after self-pairs are removed, connect the relational spectrum to regularization and risk, derive an exact inferential error for spectral truncation, and quantify how perturbations of the underlying geometry propagate to pair covariance and estimation. Simulations show when structured borrowing improves estimation and how geometric misspecification can erode that benefit. We apply the framework to autism neuroimaging using resting-state functional-connectivity data from the Autism Brain Imaging Data Exchange (ABIDE). With a 116-region parcellation and 6,670 unique connections, the application shows that a large connectome can have a much smaller effective covariance dimension. It also demonstrates that high explained covariance alone is insufficient for choosing a low-rank representation when inferential accuracy is the goal. The framework provides a principled foundation for covariance and regularization when the statistical units are unordered relationships.
Bayesian Tensor Regression for Neuroimaging Data
oai:arXiv.org:2609.27880v1
arXiv:2609.27880v1 Announce Type: new
Abstract: Multidimensional array data, or tensors, arise naturally in neuroimaging and other high-dimensional applications. We propose a parsimonious Bayesian tensor regression model for studies in which a brain image is the response and predictors are vector-valued covariates. The method extends Bayesian envelope dimension reduction to tensor responses, identifying material subspaces that contain regression information while removing variation that is immaterial to the predictors. This formulation leads naturally to a Tucker tensor decomposition and allows spatial dependence and multiple sources of uncertainty to be modeled jointly. We develop a computationally feasible Markov chain Monte Carlo algorithm based on Gibbs sampling and establish posterior consistency for the proposed model. Simulation studies demonstrate substantial gains in estimation accuracy and uncertainty quantification when meaningful dimension reduction is present. We apply the method to Human Connectome Project neuroimaging data to investigate associations between alcohol use and brain activity. The results illustrate the value of Bayesian tensor envelope regression for inference with high-dimensional, spatially dependent imaging responses.
A Hybrid Spatial Statistical Learning Framework for Individualized Probability Estimation: Application to Multisite Autism Neuroimaging
oai:arXiv.org:2609.27882v1
arXiv:2609.27882v1 Announce Type: new
Abstract: Autism spectrum disorder is a heterogeneous neurodevelopmental condition whose functional brain organization varies across individuals and imaging centers. Resting-state functional connectivity provides an opportunity to study this variation, but analysis is complicated by high dimensionality, strong dependence among connections, and substantial site heterogeneity. We introduce a Hybrid Spatial Statistical Learning Framework for individualized probability estimation in multisite autism neuroimaging. The framework combines edge-level connectivity, graph-based network organization, and participant characteristics as complementary representations of a common probability target. Fine-scale connectivity retains discriminatory information, while broader representations stabilize probability estimates under site shift. All preprocessing and model development are performed within complete site-held-out validation to assess transportability to unseen acquisition environments. Simulations show that the hybrid preserves the discrimination of the strongest edge model while improving probability accuracy as between-site heterogeneity increases. In 860 participants from 20 ABIDE sites, the hybrid retained ROC AUC while reducing Brier score and log loss. These results show that multiscale integration can improve the reliability and transportability of individualized probabilities from complex, spatially dependent biomedical data.
Uniform efficiency of the Ulrich-Wood sampler for the von Mises-Fisher distribution
oai:arXiv.org:2609.27884v1
arXiv:2609.27884v1 Announce Type: new
Abstract: The standard approach to stochastic simulation from the von Mises-Fisher distribution is a rejection sampler proposed by Ulrich and Wood. This note provides a theoretical justification of its efficiency, lower-bounding its acceptance probability uniformly over the dimension and concentration parameter. A novel interpretation of the proposal is given which provides some intuition for this efficiency.
A Uniformly Efficient Rejection Sampler for the Multivariate Expectile-Based Distribution
oai:arXiv.org:2609.27885v1
arXiv:2609.27885v1 Announce Type: new
Abstract: The multivariate expectile-based distribution is a `spiked' perturbation of a multivariate Gaussian distribution, proposed by Arbel et al. in 2023. The rejection sampler proposed for this distribution in the initial work is valid, but its acceptance probability can degenerate badly both in high dimension and in the strongly asymmetric limit. We give a simple alternative. After affine whitening and a polar decomposition, the sampling problem reduces exactly to a univariate distribution, and an additional hyperbolic change of variables exhibits this distribution as log-concave, so that a universal and uniformly efficient construction of Devroye applies. The resulting exact sampler therefore enjoys an acceptance probability of at least $1 - \exp \left( -1 \right) = 0.632120 \ldots$, uniformly over the dimension and all admissible asymmetry parameters.
Regression to the Mean-Adjusted Sample-Size Determination for Cutoff-Selected Single-Arm Pre-Post Studies
oai:arXiv.org:2609.27919v1
arXiv:2609.27919v1 Announce Type: new
Abstract: Studies enrolling participants on the basis of an extreme baseline value are susceptible to regression to the mean (RTM), such that some observed pre-post change is expected even without treatment. Existing methods estimate and decompose RTM retrospectively. We extend this framework to prospective study design by deriving a closed-form sample-size method for continuous, cutoff-selected, single-arm pre-post studies. The method partitions anticipated total change into that expected from RTM and the residual treatment effect, and powers the study to detect the latter. It uses the general bivariate-normal RTM expression, allowing baseline and follow-up variances to differ, and incorporates the corresponding conditional variance of the pre-post change. To our knowledge, no published method or software implements this cutoff-based RTM framework for prospective closed-form sample-size determination in this setting. The Stata command power onemean_rtm provides solutions for sample size and minimum detectable effect, normal-theory power evaluation, attrition adjustment, and sensitivity analysis. Monte Carlo simulation across 58 scenarios demonstrated accurate Type I error control and showed that achieved power rapidly approached nominal power as sample size increased, with the closed-form calculation generally conservative at very small sample sizes.
Asymptotically Lossless Exact and Pivotal E-Values
oai:arXiv.org:2609.27923v1
arXiv:2609.27923v1 Announce Type: new
Abstract: Consider testing a finite composite null $\cP={P_1,\ldots,P_L}$ against a simple alternative $Q$ using $n$ i.i.d. observations. Let $\ell_n$ denote the supremum of the expected log e-value over e-variables that are exact under every $P_i$ and pivotal across the nulls. Zhang, Ramdas, and Wang (2024) showed that $\ell_n$ is superadditive and asked whether $\ell_n/n$ converges to the upper bound $\min_i D(Q|P_i)$ under their joint atomlessness and absolute continuity assumptions. We prove that it does. The proof draws on techniques developed by Zhang, Ramdas, and Wang (2024) and by Farooq, Fritz, Haapasalo, and Tomamichel (2024).
A Prediction--Correction Analysis of Two-Way Block Splitting in Distributed Learning
oai:arXiv.org:2609.27924v1
arXiv:2609.27924v1 Announce Type: new
Abstract: This note revisits the convergence of the Parikh--Boyd two-way block-splitting algorithm for large-scale distributed learning through the He--Yuan prediction--correction framework. Simultaneous row--column partitioning is also relevant to hybrid federated learning, where data may be heterogeneous in both samples and features. We lift the reduced iteration to an equal-dimensional product space and reconstruct the primal and dual coordinates omitted by its implementation. The induced orthogonal-complement structure establishes exact iteration-by-iteration equivalence with the published updates. A mixed variational-inequality representation then yields a fundamental descent inequality, global convergence, an ergodic complexity bound, and current-iterate residual estimates under standard convexity, solvability, exact-subproblem, and invariant-initialization assumptions. The analysis also shows that the reduced state recursion is a metric proximal point iteration. No strong convexity, differentiability, or full-rank condition is imposed. The derivation clarifies which algebraic initialization conditions allow the reduced implementation to inherit the full-space convergence and complexity guarantees without modifying its local updates.
Dirichlet Process Mixtures of Trees with Gaussian Process Splits: A Bayesian Nonparametric Framework with Posterior Contraction Rate
oai:arXiv.org:2609.27930v1
arXiv:2609.27930v1 Announce Type: new
Abstract: We propose a Bayesian nonparametric mixture of regression trees with a Dirichlet process prior over tree-parameter pairs, enabling data-driven selection of ensemble size and unifying CART, BART, random forests, and boosting. A novel splitting rule driven by the posterior predictive of a Gaussian process within each terminal node generates flexible, smooth decision boundaries; remarkably, the GP density cancels exactly in the Metropolis--Hastings ratio for GROW/PRUNE moves, ensuring computational feasibility. An exact Gibbs sampler for posterior predictive inference propagates uncertainty through random tree traversal. A parallel MPI implementation distributes independent tree updates across processors, achieving adequate speedups. We prove posterior consistency at rate $n^{-1/4}$ in Hellinger distance under only continuity of the true regression function, allowing misspecification, via the identity $h(\Theta)=0$. Simulations on Friedman benchmark show near-nominal coverage (0.94 Gaussian, 0.92 Cauchy), robust to high-dimensional noise and heavy tails, outperforming BART and bagged CART. Applications to QSAR toxicity, crime, riboflavin, wheat genomics, and air quality confirm reliable credible intervals and automatic sparsity. The DP mixture offers a principled, robust, theoretically justified alternative for challenging regression with honest uncertainty quantification.
Randomized Spectral Inference for Hyperuniformity
oai:arXiv.org:2609.27961v1
arXiv:2609.27961v1 Announce Type: new
Abstract: We study inference for hyperuniformity from one large realization of a stationary point process. For point processes with an essentially free translation action of completely positive entropy, empirical Fourier statistics at Lebesgue-almost every frequency have asymptotically complex Gaussian laws, without quantitative mixing or cumulant-summability assumptions. Randomly sampled frequencies then give a tractable limiting experiment for low-frequency spectral mass. If the frequencies are uniform on the Euclidean ball $B_r$ of radius $r$, the mean of the limiting squared Fourier statistic is \[
A_r=\frac{\sigma_\eta(B_r)}{\rho_\eta\lambda_d(B_r)}. \] Hyperuniformity is characterized by $A_r\to0$ as $r\downarrow0$, so no continuity assumption on the structure factor at the origin is needed. Under \[
A_r=s+c r^\alpha+O(r^\beta),\qquad 0<\alpha<\beta, \] with a specified bound on the remainder and local square-integrability of the normalized Bartlett density, a two-radius extrapolation eliminates the $r^\alpha$ term and estimates $s$ with an explicit $O(r^\beta)$ bias bound. This yields confidence bounds and a one-sided test of $H_0:s=0$ whose asymptotic rejection probability under each fixed null process does not exceed the chosen significance level.
voigtinference: Exact likelihood calculus and conditional attribution for the Voigt profile
oai:arXiv.org:2609.27969v1
arXiv:2609.27969v1 Announce Type: new
Abstract: Fast, accurate algorithms for the Faddeeva function w(z), and hence the Voigt profile K(x,a), have existed for four decades, and analytic first derivatives are available in some implementations. What the established libraries reviewed here have not provided is the full likelihood calculus of the normalized Voigt distribution: applications still commonly resort to pseudo-Voigt approximations, finite-difference derivatives, or numerical convolution for parameter inference. Because w'(z) = -2z w(z) + 2i/sqrt(pi), every derivative of the Voigt log-likelihood is an algebraic function of K and the dispersion part L(x,a) = Im w(z), from the single complex evaluation that delivers the profile. This yields the score and Hessian in closed form, and the expected Fisher information by one-dimensional quadrature of an analytic integrand; for fixed interior widths sigma, gamma > 0, the MLE of the center and both widths is consistent and asymptotically normal at rate sqrt(n), despite the distribution having no finite mean or variance, so conventional likelihood-based standard errors apply. The conditional mean of the Gaussian component given an observation is (y - mu) - gamma L/K: a redescending function that attributes moderate deviations to the Gaussian (Doppler/resolution) component and extreme ones to the Lorentzian tail. The package voigtinference (Python, NumPy/SciPy, with a cross-validated Julia companion) supplies the toolkit: score, full parameter Hessian, expected information, Newton-based unbinned maximum likelihood with boundary diagnostics, conditional component moments, and evaluation validated against high-precision references at extreme width ratios. It applies directly to unbinned non-relativistic, constant-width Breit-Wigner x Gaussian resonance fits and supplies analytic Jacobians for line-shape refinement. Companion paper: arXiv:2605.01665.
Fourth-Moment Strong Universality for Finite-Type Transpose-Correlated Random Matrices
oai:arXiv.org:2609.27971v1
arXiv:2609.27971v1 Announce Type: new
Abstract: We prove a strong-universality theorem at the exact fourth-moment threshold for finite families of non-Hermitian random matrices assembled from independent unordered-pair vector atoms. Within one atom, the matrix colors and the two endpoint orientations may have arbitrary joint real covariance, subject to reversal consistency across ordered type pairs and endpoint exchangeability in same-type blocks; the law may also depend on finitely many endpoint types. The matrices may be adjoined to an arbitrary deterministic tuple that converges jointly strongly with the type projections. For every fixed matrix amplification and fixed noncommutative \(*\)-polynomial, the resulting tuple converges strongly to an explicit covariance-matched free Gaussian family. Cross-type blocks are described by masked circular variables, whereas same-type blocks split into independent endpoint-symmetric and endpoint-antisymmetric semicircular sectors. Only a finite radial fourth moment is assumed off the diagonal, and a finite second moment suffices on the diagonal. The mode of convergence depends on the coupling: corners of one infinite array converge almost surely on a common event, while fixed-law nonnested triangular arrays converge in probability for each fixed test. As an application, we obtain exact-fourth-moment strong limits for finite-separable continuous left/right profiles, including separately weighted literal-transpose terms. For the nested almost-sure formulation, the fourth-moment threshold is sharp already on the Wigner subfamily; at this threshold arbitrary fresh rows admit no coupling-invariant almost-sure upgrade of our in-probability conclusion.
Conformal Bayes under Continuous Label Shift: Sensitivity Analysis and the Limits of Exact Validity
oai:arXiv.org:2609.27976v1
arXiv:2609.27976v1 Announce Type: new
Abstract: Conformal Bayes combines Bayesian posterior predictive scores with conformal calibration, but under continuous label shift both the score and calibration weight depend on the unknown response-marginal density ratio. Existing methods typically estimate one shift parameter from pseudo-labels or predictive samples and plug it into calibration. We instead propose Joint Tilt-Sensitivity Conformal Bayes (JTS-CB), which performs sensitivity analysis over a prespecified set of plausible tilts; its split-conformal realization is JTS-SCB. Each tilt jointly determines the Bayesian conformal score and conformal importance weight. JTS-SCB forms a bounded sensitivity envelope over candidate tilts, but its calibration-only construction does not inherit the exact finite-sample weighted-conformal guarantee. We therefore study a separate candidate-weighted exact counterpart and show that its usefulness depends sharply on tail behavior. For scalar linear exponential tilts, any nonzero candidate tilt makes the exact set unbounded. More generally, tail-growing density ratios produce the same pathology, whereas quadratic tilts with a negative coefficient on \(y^2\) have vanishing tail weights and admit bounded exact inference on the original target. Ratio clipping provides a complementary bounded exact construction for a surrogate target when tails grow. Experiments show that strong plug-in predictive sampling can match the oracle when the shift is well identified, while sensitivity analysis is most useful for richer, weakly identified, or systematically biased shift models, at the cost of wider prediction sets.
BFI: An R Package for Bayesian Federated Inference
oai:arXiv.org:2609.27977v1
arXiv:2609.27977v1 Announce Type: new
Abstract: Bayesian Federated Inference (BFI) estimates statistical models from multicenter data when individual-level observations cannot be combined across centers. We present \pkg{BFI}, an \proglang{R} package that implements this methodology for Gaussian, binomial logistic, and survival regression models. Each center performs a Bayesian maximum a posteriori analysis and sends only parameter estimates and curvature information to a central server, where these summaries are combined to approximate the analysis of the combined data. The package supports prior specification, structured forms of between-center heterogeneity, several parametric and flexible baseline-hazard models for survival analysis, and treatment-effect estimation for observational and randomized studies. We describe the software design and the information exchanged between centers, give reproducible workflows for these analyses, and compare the federated results with pooled-data analyses for Gaussian, logistic, and survival models.
Improving Ensemble Filters with Flow Matching
oai:arXiv.org:2609.28015v1
arXiv:2609.28015v1 Announce Type: new
Abstract: Data assimilation estimates a dynamical state from partial and noisy observations. Classical ensemble filters are efficient but restrict analysis updates through finite sample covariance and affine Gaussian distribution. We introduce the Flow Ensemble Filter (FlowEF), which uses conditional flow matching to transport the forecast ensemble from a classical baseline filter to an analysis ensemble. FlowEF uses a localized Gaussian source during training, transports forecast ensemble members from a baseline filter at deployment, and conditions its velocity field on ensembles from that baseline filter and the observation. The proposed model therefore learns a nonlinear update while mapping each baseline ensemble independently. For sparsely observed dynamical systems, FlowEF improves both deterministic and probabilistic metrics over all four classical ensemble filters. It also achieves the best performance among the state-of-the-art generative data assimilation models.
The SMG-Yau-Yau Filter and Lossless Data Assimilation: Resolving Infinite Lie Algebras via Statistical Fiber Theory
oai:arXiv.org:2609.28062v1
arXiv:2609.28062v1 Announce Type: new
Abstract: Classical continuous-time non-linear filtering fails in generic non-linear state spaces due to an infinite Lie algebraic derivative explosion ($\dim(\mathcal{E}) = \infty$), leading to filter divergence. To resolve this four-decade crisis, we introduce the SMG-Yau-Yau filter and lossless data assimilation framework built on Statistical Fiber Theory over an Orlicz manifold $\mathcal{M}$. By equipping $\mathcal{M}$ with a Riemannian submersion and an Ehresmann connection, the unconstrained Duncan-Mortensen-Zakai score velocity field is orthogonally decomposed into Statistically Verifiable Directions ($\text{SVD}\chi_f$) and Structural Internal Directions ($\text{SID}_f$). System non-linearities and unclosed Lie commutators are orthogonally quarantined in $\text{SID}_f$, protecting macroscopic base parameters from spatial derivative pollution while preserving total score variance energy. We unify the asymptotics through a Dual-Axis Collapse Mechanism: proving our model identically recovers classical Yau-Yau dynamics when $\dim(\mathcal{E}) < \infty$, while large-sample limits ($N \to \infty$) induce a thermodynamic quench that flattens infinite-dimensional geometry under $\dim(\mathcal{E}) = \infty$. Finally, we formulate the Active Acausal Tension (AAT) functional to monitor accumulated model misspecification online. Exceeding a topological capacity threshold triggers Gauge Symmetry Breaking (GSB), which dynamically expands base coordinates ($d \to d+1$) to ensure non-asymptotic stability and convergence to the exact state density.
NPBoost: Neural Processes with Gradient-Boosted Fixed Effects
oai:arXiv.org:2609.28122v1
arXiv:2609.28122v1 Announce Type: new
Abstract: Neural Processes (NPs) are model-based meta-learners that implicitly learn a stochastic process and adapt to a new task from a small context set. Most extensions of NPs focus on improving the neural network architecture. We instead develop an extension motivated by the shared hierarchical interpretation of meta-learning and mixed-effects models. Specifically, we introduce Neural Process Boosting (NPBoost), which decomposes structured response variability into tree-boosted fixed effects shared across tasks and NP random effects that capture stochastic task-to-task variation. We propose to train the two components jointly using a boosting algorithm in which an NP learns residual task-specific structure and a tree ensemble estimates common patterns across tasks. Across synthetic and real-world tabular meta-learning problems, this decomposition improves over a standard NP when the shared structure contains discontinuities or other irregular patterns that boosted trees can represent effectively.
Extreme Population Selection under Multistage Sampling design With Applications
oai:arXiv.org:2609.28127v1
arXiv:2609.28127v1 Announce Type: new
Abstract: We study the problem of selecting the extreme (best or worst) population from among $K(\geq 2)$ populations, under the assumption that the extreme population is sufficiently separated from the nearest population. The selection is based on an appropriate measure, which may vary across different application domains. Since the actual value of the measure is unknown, we obtain its estimator using the generalized method of moments under a multistage sampling design. Using this estimator, we propose two sequential procedures, namely online algorithm and multi armed bandit based algorithm. Under suitable regularity conditions and without imposing parametric assumptions on the underlying distributions, both algorithms correctly identify the extreme population with a desired level of confidence. We illustrate the proposed algorithms through applications in econometrics and genetics. In the econometric application, the extreme population is selected using the Gini index as a measure of inequality and the performance of the proposed procedures is assessed through extensive Monte Carlo simulation studies conducted under various distributional settings. In the genetics application, the worst population is identified using a measure derived from the tumor mutation burden (TMB) score and the practical applicability of the proposed algorithms is demonstrated using the Memorial Sloan Kettering-IMPACT 50000 clinical sequencing cohort. Further, we use the proposed framework to identify an anomalous population, provided such a population exists.
How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
oai:arXiv.org:2609.28177v1
arXiv:2609.28177v1 Announce Type: new
Abstract: LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, pooled item correlation is 0.90, whereas composite-score correlation is 0.46 under item resampling and 0.92 when MMLU subjects are resampled. An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection. Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation. The resulting curves make these assumptions explicit without estimating the unobserved search size.
Penguin data reanalyzed via Computational Taxonomy
oai:arXiv.org:2609.28201v1
arXiv:2609.28201v1 Announce Type: new
Abstract: We employ Computational Taxonomy (CT) to reanalyze the penguin data set penguins_lter by validating and addressing two biological issues: Sexual Size Dimorphism (SSD) and mate-selection criteria. Via Scientific Data Analysis (SDA) computing, CT constructs a Taxonomic Hierarchy by splitting Species first and then Sex, without involving Island, to achieve less complexity. This Taxonomic Hierarchy validates SSD as a branch comparison: (Species, Sex = Male)-vs-(Species, Sex = Female), upon which SDA explores all potential pieces of associative information from all covariate feature-sets, including interacting effects from order-2 to order-4, and then confirms them via their idiosyncratic reliability checks. The collective of confirmed information pieces are displayed on a heatmap platform to manifest underlying dynamics of SSD with explicit block-structured heterogeneity found within males and females. SSD dynamics is explained through mechanistic dependence pertaining to one chief factor consisting of up to 8 feature-sets: Body-Mass coupled by combinations of {Culmen-length,Culmen-depth, Flipper-length}, and two minor factors consisting of low-order combinations of {Culmen-length,Culmen-depth, Flipper-length}. Such Intra-Sex heterogeneity invalidates all Logistic regression modeling on SSD in the original paper. Further, we explore potential mate-selection criteria through the data-frame of Nest-ID within-species homogeneity.
Bayesian statistical inverse problems for a coupled Fokker-Planck-Darcy system
oai:arXiv.org:2609.28242v1
arXiv:2609.28242v1 Announce Type: new
Abstract: We study the nonparametric statistical inverse problem of recovering the space-dependent permittivity in a coupled Fokker-Planck-Darcy system from discrete, noisy observations of the Fokker-Planck solution. We consider the coupled parabolic-elliptic system on a bounded domain with the physically natural no-flux boundary condition for the Fokker-Planck equation and homogeneous Dirichlet boundary conditions for the Darcy equation. The inverse problem is indirect, since the unknown coefficient enters only through the elliptic equation and is observed only through its effect on the density in the parabolic equation. We first develop the analytical theory of the forward problem required for the statistical analysis, establishing well-posedness and a priori bounds uniform over the admissible set of permittivities. We then derive two stability estimates for the inverse problem, namely a Lipschitz-type forward estimate and a generalized backward estimate. Placing a rescaled Gaussian process prior on the log-shifted permittivity, we show that the posterior contracts around the truth at an explicit polynomial rate in the number of observations, with the posterior mean converging at the same rate. Numerical experiments in one and two dimensions, using preconditioned Crank-Nicolson and ensemble Kalman filter algorithms, complement our theoretical results.
BRACE: Blockwise Rank Aggregation for Correlation Estimation
oai:arXiv.org:2609.28269v1
arXiv:2609.28269v1 Announce Type: new
Abstract: Chatterjee's estimator uses only two local comparisons per interior observation, leaving a finite-replication variance gap under fixed alternatives. We introduce BRACE, short for blockwise rank aggregation for correlation estimation, which replaces adjacent comparisons with local blocks and averages every within-block rank difference. The block size controls the variance cost of finite local replication. An L2 expansion separates the efficient first-order component of the response-rank statistic from an orthogonal finite-replication component. For an admissible diverging block size, the latter vanishes, so the direct rank estimator attains the information bound. The decomposition yields a consistent variance estimator and Wald confidence intervals under fixed alternatives. At independence, the efficient first-order term vanishes, and the proposed estimator enters a second-order regime in which the same factor controls its null variance.
Long-Term Tail Modeling in Survival Analysis via Extended Generalized Pareto Distributions
oai:arXiv.org:2609.28321v1
arXiv:2609.28321v1 Announce Type: new
Abstract: We propose a class of extended generalized Pareto models for right-censored survival data, with particular emphasis on tail inference and long-term extrapolation. Our framework integrates extreme value theory and survival analysis, combining a generalized Pareto distribution with a flexible perturbation distribution on the unit interval. We consider three perturbation specifications: a parametric Beta model, a Bernstein polynomial estimator, and a histogram-based estimator. To facilitate direct comparison, all three models are fitted using a unified iterative procedure adapted to right censoring through Kaplan-Meier-based pseudo-observations. A Monte Carlo study evaluates finite-sample performance across different tail indices, censoring levels, sample sizes, and model complexities. The results reveal a trade-off between flexibility and stability: the Beta specification generally performs best for tail-index estimation, whereas the histogram estimator performs particularly well for scale estimation under low censoring. The Bernstein estimator shows intermediate performance and greater sensitivity to sample size and censoring. Applications to bladder cancer recurrence and heart-failure survival data show that models with very similar in-sample fits can nevertheless produce markedly different tail-index estimates and long-term extrapolations. These findings emphasize the importance of perturbation specification when extended generalized Pareto models are used for survival extrapolation under censoring.
Simultaneous comparison of the predictive values of two binary diagnostic tests in the presence of categorical covariates
oai:arXiv.org:2609.28324v1
arXiv:2609.28324v1 Announce Type: new
Abstract: Comparison of predictive values of diagnostic tests is a topic of interest in Medical Statistics, and has been the subject of different studies. In clinical practice, it is frequent to observe categorical covariates when comparing diagnostic tests. In this framework, a global hypothesis test is proposed to simultaneously compare the predictive values of two diagnostic tests when in all of the individuals categorical covariates are observed. This hypothesis test is solved through regression models and also by weighted least squares method for the analysis of categorical data. Simulation experiments were carried out to study the asymptotic behavior of these methods when a binary covariate is observed and when a covariate with three categories is observed, and these were compared to the behavior of the global test when the covariate is ignored. In general, the method based on regression models has shown to have better asymptotic behavior than the other methods. Furthermore, we studied the application of the method based on the regression models when no covariate is observed, for which the individuals in the sample are randomly assigned to a binary dummy random variable. Simulation experiments carried out showed that this method has greater power than the method without the covariate. The results were applied to two examples.
Adaptive confidence intervals with missing data
oai:arXiv.org:2609.28336v1
arXiv:2609.28336v1 Announce Type: new
Abstract: We consider the construction of confidence intervals for population means when observations are subject to missingness. To accommodate more general missingness mechanisms than missing completely at random (MCAR), we adopt a reparametrised version of the realisable contamination model of Ma et al. (2026), which is a mixture of an MCAR version and a missing not at random version of the same base distribution $P$. We characterise the minimax length of confidence intervals that can adapt to potentially unknown parameters of the model, including the contamination fraction, for Gaussian base distributions and for nonparametric classes satisfying certain tail or symmetry assumptions. In all of these settings, we provide explicit constructions of simple, practical and finite-sample valid adaptive confidence intervals that attain the corresponding minimax rates. Finally, we provide implications of our results for causal inference.
Interpreting Hierarchical Composite Endpoints with Survival and Longitudinal Outcomes: Application to Amyotrophic Lateral Sclerosis Trials
oai:arXiv.org:2609.28359v1
arXiv:2609.28359v1 Announce Type: new
Abstract: Hierarchical composite endpoints combining survival and longitudinal functional outcomes are increasingly used in clinical trials, especially when death precludes subsequent functional assessment. The Finkelstein--Schoenfeld strategy analyzes such endpoints through prioritized pairwise comparisons, with survival compared before function. In amyotrophic lateral sclerosis (ALS), this strategy is implemented in the Combined Assessment of Function and Survival (CAFS), which combines survival with the ALS Functional Rating Scale Revised. Although such endpoints provide a clinically meaningful summary of overall treatment benefit, investigators may also want to understand whether the treatment effect is driven by survival, functional outcome, or both. Motivated by the estimand framework, we use ALS-informed simulations from a joint longitudinal--survival model to study settings in which treatment effects on survival and function align or conflict. The simulations show that decomposing the composite win probability into survival and survivor-based functional contributions clarifies their relative roles and separates the composite treatment-benefit question from function-focused questions, including the while-alive comparison and conceptual alternatives based on hypothetical and always-survivor estimands. The functional win probability underlying the while-alive comparison can be subject to survivor-selection bias when treatment affects survival, and inverse-probability weighting can attenuate selection induced by measured predictors under appropriate assumptions. Brief supporting analyses based on principal stratification and multiply robust estimation illustrate the always-survivor estimand. This framework provides practical guidance for reporting and interpreting hierarchical composite endpoints, with CAFS in ALS serving as a concrete motivating example.
Memory-Conditioned Diffusion Model for Generalized Langevin Dynamics
oai:arXiv.org:2609.28371v1
arXiv:2609.28371v1 Announce Type: new
Abstract: Generalized Langevin equations describe non-Markovian dynamics in which the evolution of resolved variables depends on their past. We propose a memory-conditioned diffusion method for learning stochastic flow maps of these dynamics from observed trajectories, without identifying a memory kernel or reconstructing unresolved variables. A compact, recursively updated bank of exponential filters enables the flow map to retain predictive history over multiple time scales without conditioning on long observation windows. The next-step distribution is conditioned on the current observation and this memory state, whose storage and update costs are independent of the history length for a fixed bank size. Predictive criteria guide the memory budget, with reference-assisted selection in the vector benchmark, and an optional linear projection further reduces the conditioning dimension. A kernel-based score estimator generates conditional samples without training a score network, and these samples are used to train a neural flow map for autoregressive simulation. Three numerical examples assess long-memory retention at small conditioning dimension, predictive compression in coupled vector dynamics, and non-Gaussian conditional distributions and intermittent events. The non-Gaussian example reproduces conditional asymmetry and burst statistics in a stochastic model of the plasma scrape-off layer.
Statistical Methods for Estimating Probability of Detection in Structural Health Monitoring
oai:arXiv.org:2609.28407v1
arXiv:2609.28407v1 Announce Type: new
Abstract: There is much interest in the potential to use structural health monitoring (SHM) technology to augment traditional nondestructive inspection (NDI) methods to improve safety, increase asset availability, and reduce maintenance and inspection costs. SHM has the potential to be used in many applications, including critical components in aircraft and pipelines. Probability of detection (POD) plays a critical role in aircraft structural integrity programs, leading to increased interest in developing methods to assess POD in SHM applications. In contrast to traditional NDI laboratory experiments involving specimens with cracks, SHM sensors are fixed, and SHM data are acquired over time as cracks grow or otherwise evolve. Thus, traditional statistical methods for assessing POD must be replaced or extended to properly handle repeated-measures data. The purpose of this paper is to review the basic statistical concepts of POD and show how these concepts can be extended or adapted for SHM-POD applications. The paper presents statistical methods for modifying and extending existing POD methods, including a simple size-of-damage-at-detection (SoDaD) method and a random-parameter (RP) method for repeated-measures data. The methods are compared using three case studies involving Piezoelectric Transducer (PZT), Carbon Nanotube (CNT), and Comparative Vacuum Monitoring (CVM) sensor systems. Results show that the SoDaD method provides a simple approach for POD estimation with limited data, while the RP method offers enhanced modeling fidelity by utilizing repeated measurements. These methods are applicable when a scalar damage index or similar response is used to make a detection decision.
Multivariate Continuous-Time Autoregressive Moving Average Processes for Astronomical Multiband Time Series
oai:arXiv.org:2609.28419v1
arXiv:2609.28419v1 Announce Type: new
Abstract: Large-scale astronomical surveys provide unprecedented volumes of multivariate time-series observations obtained through multiple optical filters. We develop a structured multivariate continuous-time autoregressive moving average (MCARMA) framework for multi-band time series with irregular sampling, heteroscedastic measurement errors, and partially observed bands. The framework allows band-specific stochastic dynamics while modeling cross-band dependence through correlated Brownian driving processes, with state-space and spectral representations enabling likelihood-based inference and interpretation of the fitted stochastic dynamics. We develop a two-stage estimation procedure in which a numerically stabilized preliminary fit initializes subsequent maximum likelihood estimation. Simulations show that higher-order stochastic structure can be recovered when its characteristic features are adequately resolved, but can become weakly identifiable because of limited temporal resolution or near pole--zero cancellation. Joint multivariate estimation improves parameter recovery in 23 of 27 settings and spectral recovery in 26 of 27 settings relative to separate single-band fits. Three Sloan Digital Sky Survey Stripe 82 quasars, respectively favoring MCARMA(1,0), MCARMA(2,0), and MCARMA(2,1), illustrate how joint multiband modeling uses cross-band dependence to inform marginal dynamics and can yield different model-order and spectral inference. The methodology is implemented in the Python package mcarma.
Detecting Structural Changes in High-Dimensional Multivariate Regression Models
oai:arXiv.org:2609.28462v1
arXiv:2609.28462v1 Announce Type: new
Abstract: We study structural-change testing in multivariate linear regression when the response dimension is proportional to the sample size and the number of predictors is fixed. Despite its relevance to applications across a broad range of fields, this problem remains underexplored. The alternatives of interest allow multiple unknown changes in a prescribed linear contrast of the predictor effects. We construct a least-squares-based Wald statistic standardized by the residual covariance estimator and scan it over candidate change-point segmentations. We propose flexible segment scans for both single and multiple change points, together with a discretized multiscale scan that reduces the computational cost. When the response dimension and sample size diverge proportionally, we establish weak convergence of the normalized statistic process to a centered Gaussian process, yielding implementable critical values for all three scans. We further characterize their asymptotic power under local alternatives, explicitly describing how the signal, design, and high-dimensional aspect ratio determine the limiting power. The finite-sample performance is examined through simulation studies. We apply the proposed methods to filtered and standardized returns from U.S. equity portfolios to investigate structural changes in their exposures to the Fama--French factors.
The Spectra of the Henze-Zirkler and Henze-Wagner Operators for BHEP Tests
oai:arXiv.org:2609.28464v1
arXiv:2609.28464v1 Announce Type: new
Abstract: The Baringhaus-Henze-Epps-Pulley (BHEP) tests for multivariate normality are affine-invariant goodness-of-fit tests based on a Gaussian-weighted $L^2$ distance between empirical and Gaussian characteristic functions. In 1990, Henze and Zirkler expressed the limiting null distribution through the eigenvalues of an integral operator on the standard Gaussian space. In 1997, Henze and Wagner obtained a simpler covariance kernel and raised the problem of calculating the eigenvalues of the resulting operator on a Gaussian-weighted space. Although subsequent work treated the univariate case and numerical approximations in a few low dimensions, the complete all-dimensional spectral problem remained open. This paper determines both complete spectra for every dimension $d \in \mathbb{N}$ and every smoothing parameter $\beta > 0$. The two operators are shown to have the forms $\mathcal{X}_{\beta,d}^*\mathcal{X}_{\beta,d}$ and $\mathcal{X}_{\beta,d}\mathcal{X}_{\beta,d}^*$ for the same Hilbert-Schmidt operator $\mathcal{X}_{\beta,d}$. Consequently, their nonzero eigenvalues agree, including multiplicities, while the null space of the Henze-Zirkler operator is identified exactly. The Gaussian integral operator in the Henze-Wagner decomposition is diagonalized by Mehler's formula, and rotational symmetry confines the finite-rank correction to the sectors associated with spherical harmonics of degrees $0$, $1$, and $2$. The degree-$1$ and degree-$2$ eigenvalues are characterized by scalar transcendental equations, and the radial eigenvalues by an explicit pole-safe Fredholm determinant. The paper establishes nonnegativity, multiplicities, eigenfunction reconstruction, completeness, the trace identity, and a complete characterization of all exceptional pole cases.
Outlier Detection for Multi-Network Data
oai:arXiv.org:2205.06398v2
arXiv:2205.06398v2 Announce Type: cross
Abstract: It has become routine in neuroscience studies to measure brain networks for different individuals using neuroimaging. These networks are typically expressed as adjacency matrices, with each cell containing a summary of connectivity between a pair of brain regions. There is an emerging statistical literature describing methods for the analysis of such multi-network data in which nodes are common across networks but the edges vary. However, there has been essentially no consideration of the important problem of outlier detection. In particular, for certain subjects, the neuroimaging data are so poor quality that the network cannot be reliably reconstructed. For such subjects, the resulting adjacency matrix may be mostly zero or exhibit a bizarre pattern not consistent with a functioning brain. These outlying networks may serve as influential points, contaminating subsequent statistical analyses. We propose a simple Outlier DetectIon for Networks (ODIN) method relying on an influence measure under a hierarchical generalized linear model for the adjacency matrices. An efficient computational algorithm is described, and ODIN is illustrated through simulations and an application to data from the UK Biobank. ODIN was successful in identifying moderate to extreme outliers. Removing such outliers can significantly change inferences in downstream applications.
Interpretable Causal Discovery via Causal-Effect Constraints
oai:arXiv.org:2608.12640v1
arXiv:2608.12640v1 Announce Type: cross
Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. We consider this task of conditional causal discovery and cast it as a Bayesian inference problem, in which we target the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint. Unfortunately, this poses a computational challenge: existing approaches to Bayesian causal discovery struggle when the event has small posterior mass. To address this, we adapt rare-event estimation techniques to perform inference the joint graph-parameter space. Our method gradually drives a particle population toward the constrained region while maintaining samples that approximate the conditional posterior. Empirical evaluation on synthetic graphs validates the accuracy of our approach at small and large scales, and we show in a case study on the Sachs protein dataset how our method can be used to aid scientific exploration by providing pathway-level summaries.
Unified framework for measuring segregation resolves how social and geographical space jointly shape connections
oai:arXiv.org:2609.16469v1
arXiv:2609.16469v1 Announce Type: cross
Abstract: Our understanding of how geographical and social segregation interact remains limited, as relatively few studies investigate them jointly, and existing approaches often lack a framework distinguishing geographical, social, and total segregation. Additionally, large-scale individually resolved geo-social network data are rarely publicly available. We address both. Conceptually, we develop a unified framework that measures segregation in geosocial networks by comparing network models to appropriate null models and recovers the Theil index, dissimilarity index, and network modularity as special cases. Empirically, we turn to privacy-preserving aggregated relational data (ARD): we combine the Facebook Social Connectedness Index for the US with US Census and Pew data, and introduce an ARD-compatible joint geosocial intervening-opportunities model to infer link probabilities between region--group cell pairs. Applying our segregation framework, we find that social segregation predominates over geographical segregation, with notable separation for White--Black, college-degree--no-degree, and high-income--low/middle-income across both segregation types. We find increasing social homophily with geographical distance and group-specific geographical connectivity patterns, suggesting that geographical segregation may affect cross-group connectivity not only directly but also by amplifying social segregation.
Positive weight Hermite and Legendre quadrature rules
oai:arXiv.org:2609.26840v1
arXiv:2609.26840v1 Announce Type: cross
Abstract: This paper contains new 80-digit positive-weight Gauss-Hermite and positive-weight interior-node Gauss-Legendre quadrature rules for up to five dimensions and varying polynomial degree accuracy (depending on quadrature type and dimension). Some of these rules improve on the best available rules in the literature and some offer rules where none (other than the tensor product) existed. The results were produced by combining methodology developed by previous researchers with new approaches. The full precision rules themselves are at https://doi.org/10.5281/zenodo.22881864. Software in Julia, Python, and R providing the rules is available via a package registry and/or GitHub: Quadriceps.jl (both 64-bit and 128-bit; https://github.com/NittanyLion/Quadriceps.jl), quadriceps-py (64-bit only; https://github.com/NittanyLion/quadriceps-py), and quadriceps-r (64-bit only; https://github.com/NittanyLion/quadriceps-r). The software used to create these rules is available via PositiveWeightQuadratureSolvers.jl (https://github.com/NittanyLion/PositiveWeightQuadratureSolvers.jl). The replication package is at QuadricepsReplicationPackage.jl (https://github.com/NittanyLion/QuadricepsReplicationPackage.jl). A snapshot of all five packages is archived at https://doi.org/10.5281/zenodo.22883240.
Gaussian-process surrogate indicators for residual-based adaptive GMsFEM
oai:arXiv.org:2609.26843v1
arXiv:2609.26843v1 Announce Type: cross
Abstract: Residual-based adaptive GMsFEM for high-contrast elliptic problems repeatedly evaluates local weighted $H^{-1}$ indicators on every coarse neighborhood, making indicator evaluation a recurring cost in repeated-query settings. We introduce a non-intrusive Gaussian-process (GP) surrogate for the indicator scores used in D\"orfler marking. The operational predictor is the GP posterior mean, algebraically equivalent to a kernel ridge regression (KRR) estimator under the stated convention; it uses compressed local solution and spectral features without changing the multiscale solve, local spectral construction, or basis enrichment. A nonuniform perturbed-marking result quantifies how pointwise score errors affect the exact indicator mass captured by surrogate-selected neighborhoods, while a conditional bounded-discrepancy KRR pathway identifies sufficient assumptions for such score bounds. In controlled held-out in-distribution tests, the surrogate-guided method gives error-versus-DoF trends comparable with classical $H^{-1}$-residual offline adaptivity and evaluates the online indicator component 2.0-2.1 times faster, excluding offline data generation and GP training.
WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
oai:arXiv.org:2609.27033v1
arXiv:2609.27033v1 Announce Type: cross
Abstract: Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.
Dual Boundary Condition Inference: A Two-Boundary Product Rule and Its Implications
oai:arXiv.org:2609.27120v1
arXiv:2609.27120v1 Announce Type: cross
Abstract: Many inference problems are constrained from two sides: forward information and a second constraint on outcomes. Dual Boundary Condition Inference (DBCI) multiplies them and renormalises. The rule is not new: it is the product-of-experts form and the unit-exponent member of the logarithmic-pooling family. The paper contributes a classification and reducibility characterisation. The algebra does not fix the exponent at which the inputs combine: given the statistic log p_fwd + log p_bwd, which already builds in the product form, maximum entropy supplies the family (p_fwd p_bwd)^lambda and selects no member. If the inputs are read as two equally weighted opinions, three requirements select lambda = 1/2: unanimity preservation, consistency under a common Bayesian update, and minimal symmetric Kullback-Leibler divergence. If they are read as separately applied factors, two requirements select lambda = 1: a neutral second input must leave the first unchanged, and an update to one input must pass through unchanged. For a rank-one intermediate projective measurement, the Aharonov-Bergmann-Lebowitz (ABL) rule realises this form with no extra parameter. Every positive exponent orders outcomes identically, so the choice is invisible to picking the likeliest outcome, but not to scoring a class by total mass. Under the relevant identifications, the same product form gives Bayes' rule, while a final effect proportional to the identity returns the ordinary forward Born probabilities. The second result is a Two-Boundary Reducibility Criterion: DBCI factors through the forward boundary exactly when the effective backward boundary, up to positive rescaling, does. So a fixed prior or model-fixed structural constraint adds no distinctions between cases sharing the same forward boundary, though it may encode substantial information and still materially affect the result.
The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems
oai:arXiv.org:2609.27155v1
arXiv:2609.27155v1 Announce Type: cross
Abstract: With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal account) remains underexplored. Existing studies on agent poisoning typically assume that the adversary can expose poisoned content to the agent. Although such an attack is direct and effective, it is more easily detected and mitigated. In the context of social media platforms, this leaves open whether the recommendation system itself would surface such content to the agent in a more subtle manner. Through theoretical analysis, we show that the like-score mechanism used in OASIS can be exploited, and we characterize the conditions under which a multi-stage chain of poisoned posts can steer the agent's feed. Based on these insights, we further develop an algorithm that crafts realistic poisoned posts. Experiments support our theoretical findings and demonstrate the effectiveness of the proposed algorithm. Notably, by exploiting the like-score feedback loop, the attack causes the recommendation system to select poisoned posts even when their user-post similarity falls below the retrieval threshold.
Discrete Diffusion Models via Evolving Variational Autoregressive Networks
oai:arXiv.org:2609.27306v1
arXiv:2609.27306v1 Announce Type: cross
Abstract: Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.
Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following
oai:arXiv.org:2609.27421v1
arXiv:2609.27421v1 Announce Type: cross
Abstract: Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.
Finite-Sample Binary Hypothesis Testing via R\'enyi Divergences: Strong Converse and Local Privacy
oai:arXiv.org:2609.27617v1
arXiv:2609.27617v1 Announce Type: cross
Abstract: We study asymmetric simple binary hypothesis testing between $H_0:P_0^{n}$ and $H_1:P_1^{n}$, based on $n$ independent and identically distributed observations. Leveraging a variational representation of R\'enyi divergence of order $\alpha$, we derive our main result: a finite-sample converse with $\alpha>1$. The bound uses both directions of the divergence $D_\alpha(P_1\|P_0)$ and $D_\alpha(P_0\|P_1)$, tensorises under product measures, and contains familiar data-processing converses as boundary cases. For comparison, we apply the same variational approach to general $f$-divergences and specialise it to total variation, $E_\gamma$, Hellinger, and Kullback Leibler divergences, thereby recovering familiar converses within a unified framework. Together with an achievability bound involving R\'enyi divergence with $\alpha\in (0,1)$, the main converse recovers the phase transition of the optimal Type II error under the exponentially decaying Type I error constraint $\varepsilon_n=e^{-nr}$. Under regularity conditions, the optimal Type II error vanishes exponentially when $rD(P_1\|P_0)$. We also derive sample-complexity bounds and extend both the converse and achievability analyses to locally differentially private observations, quantifying the cost of privacy and recovering the non-private achievability bound as the privacy constraint vanishes.
A weak invariance principle for triangular arrays of independent random variables in some Besov spaces
oai:arXiv.org:2609.27711v1
arXiv:2609.27711v1 Announce Type: cross
Abstract: We establish a Donsker-Prokhorov invariance principle in some Besov spaces. Specifically, we show that polygonal line processes associated with partial sum processes of a triangular array of row-wise independent random variables converge in distribution to Brownian motion, extending earlier results in the literature.
Financial Tail Risk Beyond Lipschitz Continuity via Semi-Discrete Optimal Transport
oai:arXiv.org:2609.27785v1
arXiv:2609.27785v1 Announce Type: cross
Abstract: Financial returns are heavy-tailed, and accurate tail risk estimation is central to portfolio risk management. Modern neural generators sample by pushing a simple base distribution through a learned map, and for training stability that map is built from Lipschitz components. This is the binding constraint: a Lipschitz map of a Gaussian is sub-Gaussian, so heavier-tailed targets admit no exact match at any finite Lipschitz constant. The Monge--Amp\`ere equation ties the Brenier map's local distortion to the density ratio $f/(g\circ T)$, so a deeper trough in the target density requires a higher-gain map and yields a higher-variance estimator. The argument needs only bounded distortion, so it covers normalizing flows, flow matching, GANs, and diffusion samplers alike.
Semi-Discrete Optimal Transport (SDOT) relaxes the map's regularity rather than the source's tail class. Its power diagram gives every training observation a cell holding exactly $1/N$ of the source measure, and tail observations are reached by crossing a cell boundary rather than by stretching. Our primary experiment sweeps severity over a calibrated Merton jump-diffusion spanning kurtosis 94 to 1,679. SDOT holds tail ratios at $0.85$--$0.94$ with cross-seed standard deviations below $0.025$, while every learned generator either compresses the tails or inflates them with a variance that grows alongside. Further experiments carry the result to real S\&P~500 returns and to a 21-year backtest, where SDOT gives the best risk-adjusted market-neutral strategy under CVaR optimization (Sharpe $0.70$, max drawdown $-2.60\%$, against $0.40$ for the next-best generator).
Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel
oai:arXiv.org:2609.27867v1
arXiv:2609.27867v1 Announce Type: cross
Abstract: A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time series foundation models, on a production marketplace panel of 1,887 business customers over 67 months. We hold the data, the horizon and the period fixed, and vary only the evaluation design. Three choices each reverse or dissolve a headline conclusion. Changing the unit of analysis from the market total to the individual customer moves our production baseline from second of nineteen, beaten by nothing, to twenty-third of twenty-five. Nineteen of its twenty-four challengers beat it there. Changing how much error is pooled decides whether a Diebold-Mariano test finds anything at all. Scoring prediction intervals rather than point forecasts reorders the field almost completely, with a rank correlation of 0.02 on intermittent demand. We then measure what the deployed system gets from this. Its selection rule captures 55% of the distance between doing nothing and choosing with hindsight. The reversal is not a quirk of our data. We ran the released protocol, unchanged, on the public M5 retail panel. The same baseline shape places first at the market total and last per series, beaten by everything, and a replayed selection rule closes 64.7% of the same floor-to-ceiling distance there. Adding five zero-shot foundation models to that roster changes who wins at the total, not the shape. The bands' blind spot travels too: conformal bands under-cover most on the spikiest items. Splitting our own panel into ever smaller groups turns the contrast into a curve: the baseline's rank worsens at every level of disaggregation. We release the evaluation protocol and report an error of our own that inverted a result before we caught it.
Global tree forecasters collapse at the hierarchical aggregate: a five-panel failure characterization
oai:arXiv.org:2609.27912v1
arXiv:2609.27912v1 Announce Type: cross
Abstract: Global forecasting models pool many series and learn one shared function. Gradient-boosted trees are their most common form. We measure a failure of this design that has not, to our knowledge, been documented. Train a global tree on the individual series of a hierarchy, then ask it for the hierarchical aggregate. The aggregate sits far outside the model's training range, and the forecast collapses. The model under-predicts the total by 30-50x in our production deployment, and by up to 496x in a public M5 reconstruction. The mechanism is known: beyond its training range, a tree predicts a constant. It surfaces at the aggregate because the total dwarfs every training series. The cure is not new. Per-series scaling, the preprocessing step that Montero-Manso and Hyndman (2021) recommend, prevents the collapse. So do a weighted aggregate-level training row and seasonal differencing. Our contribution is the characterization. The collapse reproduces on five panels: a production business-to-business marketplace, a synthetic hierarchy, M5, Australian Tourism, and a public business-buyer panel. It holds on three tree libraries, is invariant across training seeds, and is statistically significant. Its onset is immediate and tracks a simple support bound: a scale gap of only 1.15x already costs a third of the total. No standard configuration change prevents it: pooling every hierarchy level into training fails at scale, and the one knob that fits linear models in the leaves softens it without curing it. Rolling the forecasts forward recursively separates the cures: the aggregate-row cure re-collapses, per-series scaling degrades but stays low, and only seasonal differencing keeps its one-step accuracy unchanged. We close with a three-step procedure for diagnosing and preventing the failure in deployed systems.
Distributional Difference-in-Differences: Aggregation Before or After Quantile Inversion?
oai:arXiv.org:2609.27944v1
arXiv:2609.27944v1 Announce Type: cross
Abstract: Staggered distributional difference-in-differences produces cohort-specific potential-outcome distributions, but applied work typically wants one overall quantile treatment effect. Two natural summaries---averaging cohort quantile treatment effects (QTTs) and mixing cohort distributions before inversion---use the same policy weights yet answer different target-population questions and can disagree even in sign. We derive the exact sharp interval for their gap conditional on cohort quantiles and weights, a globally sharp range-only envelope, and a locally sharp density-tilt representation. We develop joint smooth and mass-point-safe inference and show that neither estimator is uniformly more precise, even when the estimands coincide. In a same-object reconstruction of a public staggered-QTT application, holding data, identification, distributions, and weights fixed while changing only aggregation order reverses reported signs at several quantiles. Aggregation order is therefore part of the estimand and must be chosen before inversion.
Rank-One Signal Recovery in Sparse Wishart Noise
oai:arXiv.org:2609.28163v1
arXiv:2609.28163v1 Announce Type: cross
Abstract: We study the high-dimensional recovery of a signal vector $\mathbf{x}$ in the presence of sparse Wishart-like noise. We define an $N \times N$ matrix $A = J+(\theta/N)\mathbf{xx}^{\top}$, where $\mathbf{xx}^{\top}$ is the rank-one deformation of the random noise matrix $J$. We consider a Wishart-like matrix $J={X}^{\top} X$, where $X$ is a sparse $M \times N$ random matrix with entries $X_{ij} = c_{ij}W_{ij}$, with $c_{ij}$ regulating the density of non-zero elements, and $W_{ij}$ the bond weights. Using the replica method, we compute analytically the top eigenpair statistics of $A$, and their dependence on the signal strength $\theta$, the rectangularity ratio $\alpha=\sqrt{M/N}$, and the average connectivity of the noise. The spectral observables are expressed in terms of a system of Recursive Distributional Equations, which are efficiently solved via a Population Dynamics algorithm. They allow us to compute the average largest eigenvalue $\langle\lambda_1\rangle_{A}$, the average top eigenvector component density, and the average overlap between the top eigenvector of $A$ and $\mathbf{x}$. We identify a critical threshold $\theta_{\mathrm{crit}}$--depending on the average connectivity of the noise--that marks a BBP-like phase transition: below this value, $\langle\lambda_1\rangle_{A}$ is unaffected by the signal, and the overlap vanishes. Thus, the signal is not recoverable from the top eigenvector of $A$. For $\theta>\theta_{\mathrm{crit}}$, the signal-related outlier eigenvalue becomes $\langle\lambda_1\rangle_{A}$ and the overlap is nonzero, allowing for recovery of the signal. The results are in excellent agreement with numerical diagonalisation. We show that in the dense limit, the recovery threshold and eigen-statistics converge to the results predicted by the classical BBP transition for additive rank-one deformations of dense Wishart matrices.
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions
oai:arXiv.org:2609.28297v1
arXiv:2609.28297v1 Announce Type: cross
Abstract: We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.
Local Geometric Mixing via Dobrushin Contraction with Applications to Diffusion Path Monte Carlo and the Proximal Sampler
oai:arXiv.org:2609.28338v1
arXiv:2609.28338v1 Announce Type: cross
Abstract: Local geometric mixing localizes geometric mixing by requiring geometric convergence to equilibrium in total variation only over finitely many transitions. It accommodates local convergence rates and captures rapid local equilibration, even when global mixing is much slower. We establish and discuss local geometric mixing bounds through Dobrushin contraction. We then apply this approach to Diffusion Path Monte Carlo, a recently proposed Markov chain Monte Carlo method, aimed at leveraging advances in score-based modeling, whose ideal transitions coincide with those of the Proximal Sampler. Our analysis covers both the ideal method and its implementable Metropolis-adjusted counterpart, providing mixing guarantees under minimal assumptions. For the ideal method, these guarantees complement recent spectral gap estimates, which we develop into mixing time bounds.
Even Sharper Bounds for Transductive Learning and Its Applications
oai:arXiv.org:2609.28459v1
arXiv:2609.28459v1 Announce Type: cross
Abstract: We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed-point and confidence terms as the classical inductive local Rademacher-complexity bounds, without the additional logarithmic confidence factor in earlier transductive results. For realizable learning over a binary class of VC dimension $\dVC$, with training size $m$, test size $u$, and $u\ge m\ge\dVC$, STLC yields $\cO\{\dVC\log(me/\dVC)/m\}$. This matches the standard inductive rate and, when $m\ge9$, is within a logarithmic factor of the transductive minimax lower bound of order $\dVC/m$. For transductive kernel learning, STLC gives a spectrum-adaptive excess-risk bound without the multiplicative imbalance factors appearing in the earlier local-complexity bound.
Delicatessen: Automated Estimating Equations in Python
oai:arXiv.org:2203.11300v4
arXiv:2203.11300v4 Announce Type: replace
Abstract: Estimating equation theory provides a unified framework for statistical modeling and inference. Applying estimating equations has historically involved extensive matrix algebra and by-hand differentiation of complex functions. Here, we introduce delicatessen, a Python library that automates those tedious calculations, lowering the barrier to adoption in quantitative biological and life science research. To highlight the utility of delicatessen for the acceleration of scientific research, we provide illustrative examples of linear regression with outliers, estimation of a dose-response curve, and standardization of the mean. Across these and other settings, delicatessen streamlines transparent and reproducible advanced statistical procedures for modern data analysis.
A Bayesian Framework for Multivariate Differential Analysis
oai:arXiv.org:2307.08975v4
arXiv:2307.08975v4 Announce Type: replace
Abstract: Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values.
Estimation of the invariant measure of a multidimensional diffusion from noisy observations
oai:arXiv.org:2404.12181v2
arXiv:2404.12181v2 Announce Type: replace
Abstract: We introduce a new approach for estimating the invariant density of a multidimensional diffusion when dealing with high-frequency observations blurred by independent noise. We consider the intermediate regime, where observations occur at discrete time instances $k\Delta_n$ for $k=0,\dots,n$, under the conditions $\Delta_n\to 0$ and $n\Delta_n\to\infty$. We construct a kernel density estimator based on preaveraged observations to reduce the effect of the noise and involves a two-step bias-correction procedure to appropriately account for the bias introduced by the pre-averaging. The rate of convergence of our estimator depends on both the anisotropic regularity of the density and the intensity of the noise. We establish conditions on the intensity of the noise that ensure the recovery of convergence rates similar to those achievable without any noise. Furthermore, we prove a Bernstein concentration inequality for our estimator, from which we derive an adaptive procedure for the kernel bandwidth selection.
Variance Reduction for Independent Metropolis
oai:arXiv.org:2406.17699v3
arXiv:2406.17699v3 Announce Type: replace
Abstract: Assume that we would like to estimate the expected value of a function $F$ with respect to an intractable density $\pi$, which is specified up to some unknown normalising constant. We prove that if $\pi$ is close enough under KL divergence to another density $q$, an independent Metropolis sampler estimator that obtains samples from $\pi$ with proposal density $q$, enriched with a variance reduction computational strategy based on control variates, achieves smaller asymptotic variance than i.i.d. sampling from $\pi$. The control variates construction requires no extra computational effort but assumes that the expected value of $F$ under $q$ is analytically available. We illustrate this result by calculating the marginal likelihood in a linear regression model with prior-likelihood conflict and a non-conjugate prior. Furthermore, we propose an adaptive independent Metropolis algorithm that adapts the proposal density such that its KL divergence with the target is being reduced. We demonstrate its applicability in a Bayesian logistic and Gaussian process regression problems and we rigorously justify our asymptotic arguments under easily verifiable and essentially minimal conditions.
Asymptotic confidence intervals for the difference and the ratio of the weighted kappa coefficients of two diagnostic tests subject to a paired design
oai:arXiv.org:2407.21387v2
arXiv:2407.21387v2 Announce Type: replace
Abstract: The weighted kappa coefficient of a binary diagnostic test is a measure of the beyond-chance agreement between the diagnostic test and the gold standard, and depends on the sensitivity and specificity of the diagnostic test, on the disease prevalence and on the relative importance between the false positives and the false negatives. This article studies the comparison of the weighted kappa coefficients of two binary diagnostic tests subject to a paired design through confidence intervals. Three asymptotic confidence intervals are studied for the difference between the parameters and five other intervals for the ratio. Simulation experiments were carried out to study the coverage probabilities and the average lengths of the intervals, giving some general rules for application. A method is also proposed to calculate the sample size necessary to compare the two weighted kappa coefficients through a confidence interval. A program in R has been written to solve the problem studied and it is available as supplementary material. The results were applied to a real example of the diagnosis of malaria.
FastManly: An EM-Gradient Algorithm for Manly Mixture Models
oai:arXiv.org:2410.00848v2
arXiv:2410.00848v2 Announce Type: replace
Abstract: A faster implementation of mixtures of Manly transformations is proposed. This method, called FastManly, uses Newton's method for optimization in an EM gradient algorithm instead of Nelder-Mead in a traditional EM. A gradient and full Hessian are derived. Simulations show improved performance with noticeable speedups.
Statistical Properties of Deep Neural Networks with Dependent Data
oai:arXiv.org:2410.11113v4
arXiv:2410.11113v4 Announce Type: replace
Abstract: This paper develops theory for deep neural network (DNN) estimators under dependent data. To provide theory applicable to a variety of DNN-based estimators, I first establish nonasymptotic probability bounds on the theoretical and empirical $\mathcal{L}^{2}$-errors of nonparametric sieve estimators for a general class of estimation problems under possibly nonstationary $\beta$-mixing data taking values in unbounded sets. I then apply the theory to fully connected and convolutional DNN estimators without bounds or sparsity restrictions on the DNN weights. For both DNN classes, I derive general results when the function to be estimated is H\"older smooth and the data are nonstationary, subgaussian, and $\beta$-mixing with either exponential or polynomial decay. I then specialize these to nonparametric regression, logistic regression, and quantile regression settings. Under exponential $\beta$-mixing, the resulting estimators attain the nonparametric minimax rate of Stone (1982) up to logarithmic factors.
Interpretable Deep Neural Network for Modeling Functional Surrogates
oai:arXiv.org:2503.20528v5
arXiv:2503.20528v5 Announce Type: replace
Abstract: Developing surrogates for computer models has become increasingly important for addressing complex problems in science and engineering. This article introduces an artificial intelligent (AI) surrogate, referred to as the DeepSurrogate, for analyzing functional outputs with vector-valued inputs. The relationship between the functional output and vector-valued input is modeled as an infinite sequence of unknown functions, each representing the relationship at a specific location within the functional domain. These spatially indexed functions are expressed through a combination of basis functions and their corresponding coefficient functions, both of which are modeled using deep neural networks (DNN). The proposed framework accounts for spatial dependencies across locations, while capturing the relationship between the functional output and scalar predictors. It also integrates a Monte Carlo (MC) dropout strategy to quantify prediction uncertainty, enhancing explainability in the deep neural network architecture. The proposed method enables efficient inference on datasets with approximately 50,000 spatial locations and 20 simulations, achieving results in under 10 minutes using standard hardware. The approach is validated on extensive synthetic datasets and a large-scale simulation from the Sea Lake and Overland Surge from Hurricanes (SLOSH) simulator. An open-source Python package implementing the method is made available.
Bayesian inference for the learning rate in Generalised Bayesian inference
oai:arXiv.org:2506.12532v3
arXiv:2506.12532v3 Announce Type: replace
Abstract: In Generalised Bayesian Inference (GBI), the learning rate and hyperparameters of the loss must be estimated. These inference-hyperparameters can't be estimated jointly with the other parameters, from the data, by giving them a prior. However, in some settings there exist unknown ``true'' hyperparameter-values about which it is meaningful to have prior belief. It is then possible to use Bayesian inference with held-out data to get hyperparameter-posteriors. We define two hyperparameter posteriors, one based on an Expected Log Pointwise Predictive Density (ELPPD)-utility and one aiming to cover the pseudo-true parameter. The new framework supports estimation and uncertainty quantification for multiple hyperparameters jointly. Experiments show that the resulting GBI-posteriors outperform Bayesian inference on simulated test data and select optimal or near-optimal hyperparameter values in a large real problem of text analysis. Generalised Bayesian inference is particularly useful for combining multiple data sets and most of our examples belong to that setting. We also give asymptotic results for some of the special ``multi-modular'' Generalised Bayes posteriors which we use in our examples.
Design-Life Levels for Flood Extremes: A Dependence-Aware Block-Maxima Workflow for Severity and Persistence
oai:arXiv.org:2506.14556v3
arXiv:2506.14556v3 Announce Type: replace
Abstract: Flood-risk assessment concerns the magnitude of maximum discharge over a design life and the clustering of extreme flows in time. Limited record length, extremal clustering, and pre-asymptotic behavior complicate inference from daily streamflow and flood-impact records, yet severity estimation, persistence assessment, and design-life levels are often handled separately. We develop a dependence-aware block-maxima workflow for approximately stationary records with regularly varying upper tails. The severity branch estimates the extreme value index (EVI) from sliding block-maximum quantile scaling using data-adaptive plateau selection and covariance-aware feasible generalized least squares (FGLS). The persistence branch pools native block-maxima extremal-index paths over a stable block-size window to characterize extremal clustering. Design-life levels are then derived on the relevant observation clock, with the extremal index retained as a complementary persistence descriptor. In synthetic short-record benchmarks, median-sliding-FGLS achieves the lowest grid-average Winkler score for nominal 95% EVI confidence intervals among the methods compared. For the extremal index, FGLS pooling of sliding-block Berghaus--B{\"u}cher estimates lowers mean Winkler scores relative to pooled ordinary least squares in all 84 scenarios and to the native estimator in 81. Applications to Texas and Florida streamflow and National Flood Insurance Program building-claim records show stronger extremal clustering in streamflow and steeper block-maximum scaling in insured losses on the active-day clock.
Upgrading survival models with CARE
oai:arXiv.org:2506.23870v3
arXiv:2506.23870v3 Announce Type: replace
Abstract: Clinical risk prediction models are regularly updated as new data, often with additional covariates, become available. We propose CARE (Convex Aggregation of relative Risk Estimators) as a general approach for combining existing "external" estimates with a new data set in a time-to-event survival analysis setting. Our method initially employs the new data to fit a flexible family of reproducing kernel estimators via penalised partial likelihood maximisation. The final relative risk estimator is then constructed as a convex combination of the kernel and external estimators, with the regularisation parameters and convex combination coefficients selected using cross-validation. We establish high-probability bounds for the $L_2$-error of our proposed aggregated estimator, showing that it achieves a rate of convergence that is at least as good as both the optimal kernel estimator and the best existing external model. Empirical results from simulation studies align with the theoretical results, and we illustrate the improvements our methods provide for cardiovascular disease risk modelling. Our methodology is implemented in the Python package care-survival.
Edgeworth corrections for the spiked eigenvalues of non-Gaussian sample covariance matrices with applications
oai:arXiv.org:2507.09584v3
arXiv:2507.09584v3 Announce Type: replace
Abstract: Yang and Johnstone (2018) established an Edgeworth correction for the largest sample eigenvalue in a spiked covariance model under the assumption of Gaussian observations, leaving the extension to non-Gaussian settings as an open problem. In this paper, we address this issue by establishing first-order Edgeworth expansions for spiked eigenvalues in both single-spike and multi-spike scenarios with non-Gaussian data. Leveraging these expansions, we construct more accurate confidence intervals for the population spiked eigenvalues and propose a novel estimator for the number of spikes. Simulation studies demonstrate that our proposed methodology outperforms existing approaches in both robustness and accuracy across a wide range of settings, particularly in low-dimensional cases.
Measures of Dependence based on Wasserstein distances
oai:arXiv.org:2510.06034v2
arXiv:2510.06034v2 Announce Type: replace
Abstract: Measuring dependence between random variables is a fundamental problem in Statistics, with applications across diverse fields. While classical measures such as Pearson's correlation have been widely used for over a century, they have notable limitations, particularly in capturing nonlinear relationships and extending to general metric spaces. In recent years, the theory of Optimal Transport and Wasserstein distances has provided new tools to define measures of dependence that generalize beyond Euclidean settings. This survey explores recent proposals, outlining two main approaches: one based on the distance between the joint distribution and the product of marginals, and another leveraging conditional distributions. We discuss key properties, including characterization of independence, normalization, invariances, robustness, sample, and computational complexity. Additionally, we propose an alternative perspective that measures deviation from maximal dependence rather than independence, leading to new insights and potential extensions. Our work highlights recent advances in the field and suggests directions for further research in the measurement of dependence using Optimal Transport.
Regular Fourier Features for Nonstationary Gaussian Processes
oai:arXiv.org:2602.23006v3
arXiv:2602.23006v3 Announce Type: replace
Abstract: Simulating a Gaussian process requires sampling from a high-dimensional Gaussian distribution, which scales cubically with the number of sample locations. Spectral methods address this challenge by exploiting the Fourier representation and treating the spectral density as a probability distribution suitable for Monte Carlo approximation. Although this probabilistic interpretation is valid for stationary processes, it is overly restrictive for the nonstationary case, where spectral densities are generally not probability measures. To avoid this limitation, we propose regular Fourier features for harmonizable processes with one-dimensional inputs. Our method discretizes the spectral representation directly, preserving the correlation structure among spectral weights without requiring probability assumptions. Assuming finite spectral support, this yields an efficient low-rank approximation that is positive semi-definite by construction and consistent under mild regularity conditions. When the spectral density is unknown, the framework also extends to kernel learning from data, which we explore as a proof of concept. We demonstrate the approximation on locally stationary and harmonizable mixture kernels, the latter with a complex-valued spectral density. As a feasibility study, we then apply the kernel-learning extension to real and synthetic data, where it matches competitive baselines.
Engaging students with statistics through choice of real data context on homework
oai:arXiv.org:2603.04541v2
arXiv:2603.04541v2 Announce Type: replace
Abstract: Statistics educators recommend teaching with real data with relevant contexts, but defining relevancy is challenging and varies by student. We investigated whether providing student choice of data context increases engagement through a quasi-experiment in two sections of an introductory probability and statistics course at a large public university (n=65 consenting students). Sections alternated as treatment and control: during their treatment, students chose weekly homework from three similar instructor-provided options varying by data context; during control weeks, they received randomly assigned contexts. We found no significant difference in homework grades between treatment and control conditions. However, thematic analysis revealed students with choice reported enhanced engagement and motivation, greater appreciation for statistics' real-world value, and increased autonomy. Students overwhelmingly preferred contexts relevant to their interests, experiences, daily lives, and career paths-though preferences varied considerably across individuals. Based on these findings, we provide four recommendations for statistics educators: (1) use real data with authentic contexts, (2) select contexts students care about, (3) incorporate variety across data contexts, and (4) consider choice as a pedagogical tool.
BLOC: A Global Optimization Framework for Sparse Covariance Estimation with Non-Convex Penalties
oai:arXiv.org:2603.29169v2
arXiv:2603.29169v2 Announce Type: replace
Abstract: We introduce BLOC (Black-box Optimization over Correlation matrices), a general framework for sparse covariance estimation with non-convex penalties. BLOC operates on the manifold of correlation matrices and reparameterizes it via an angular Cholesky mapping, transforming the positive-definite, unit-diagonal constraint into an unconstrained search over a Euclidean hyperrectangle. This enables gradient-free global optimization of diverse objectives, including non-differentiable or black-box losses, using a pattern search routine with adaptive coordinate polling, run-wise restarts to escape local minima, and leveraging up to $d(d-1)$ parallel threads when optimizing a $d$-dimensional correlation matrix. The method is penalty-agnostic and ensures that every iterate is a valid correlation matrix, from which covariance estimates are obtained. We establish convergence guarantees, including stationarity, probabilistic escape from poor local minima, and sublinear rates under smooth convex losses. From a statistical perspective, we prove consistency, convergence rates, and sparsistency for penalized correlation estimators under general conditions, extending sparse covariance theory beyond the Gaussian setting. Empirically, BLOC with nonconvex penalties such as SCAD and MCP outperforms leading estimators in both low- and high-dimensional regimes, achieving lower estimation error and improved sparsity recovery. A parallel implementation enhances scalability, and a proteomic network application demonstrates robust, positive-definite sparse covariance estimation.
Fitting Large Nonlinear Mixed Effects Models Using Variational Expectation Maximization
oai:arXiv.org:2604.26160v2
arXiv:2604.26160v2 Announce Type: replace
Abstract: Nonlinear Mixed Effects (NLME) models are widely used in pharmacometrics and related fields to analyze hierarchical and longitudinal data. However, as the number of parameters and random effects increases, traditional methods for maximizing the marginal likelihood become computationally expensive. This paper explores the Variational Expectation Maximization (VEM) algorithm, a scalable alternative for fitting NLME models. Originally introduced in the context of probabilistic graphical models and later popularized through variational autoencoders, VEM has not been extensively applied to NLME modeling. By leveraging flexible variational families and reverse-mode automatic differentiation, VEM can efficiently maximize the marginal likelihood, scaling to NLME models with over 15,000 population parameters. This work provides a detailed description of VEM, compares it to other NLME fitting algorithms, and highlights its scalability through computational experiments. Using the Pumas statistical software, we fit two test models: 1) a standard warfarin model, and 2) an unnecessarily over-parameterized DeepNLME Friberg model with 15,410 population parameters and 16 random effects. The warfarin model was fitted to completion to demonstrate the correctness of VEM, while the DeepNLME Friberg model instead demonstrates VEM's scalability on a toy but large model. VEM improves the log likelihood steadily over hundreds of iterations at a practical per-iteration cost, while FOCE fails to complete even one iteration within a day. The model is deliberately over-parameterized for its small dataset and over-fits it, so what this experiment establishes is that VEM optimizes the objective of a model of this size at a practical cost. Applying VEM to large models that are genuinely useful is left to future work.
Simultaneous Latent Budget Trees for Stratified Classification
oai:arXiv.org:2606.13295v3
arXiv:2606.13295v3 Announce Type: replace
Abstract: In the era of Explainable Artificial Intelligence, there is a renewed focus on single trees for their ease of interpretation. This paper introduces Simultaneous Latent Budget Trees, a probabilistic machine learning framework for classification trees in the presence of a stratification factor such as a temporal, spatial, or demographic variable, acting as a control variable or potential confounder. Standard tree growth procedures are not designed to optimize a conditional split rule. A model-based split rule is proposed in which child nodes are interpreted as latent components of a simultaneous mixture model, such as the Simultaneous Latent Budget Model and its constrained versions, fitted to the parent node. Mixing parameters drive the observations, differently for each group, to the child nodes whereas latent budgets parameters update the response classes profile of each level of the control variable. Parameters are estimated by least squares considering a neural network perspective of the model. An informative tree structure can be interactively visualized with interpretation aids on the node and the paths, including visual pruning and decision tree selection procedure. Suitable measures are proposed to handle an unbalanced response class distribution. The proposed methodology is applied to investigate gender-related differences in disease progression of Amyotrophic Lateral Sclerosis. The SLBT library with the various tree-based algorithms is available in the linked GitHub repository.
Bayesian spatial modelling framework for assessing residential flood risk in property insurance
oai:arXiv.org:2607.07609v2
arXiv:2607.07609v2 Announce Type: replace
Abstract: Spatial heterogeneity in insurance risk modelling is often represented using coarse areal structures, which can obscure fine-scale patterns critical for accurate risk assessment. This study introduces a point-referenced Bayesian framework to model claim occurrence and severity at the policyholder level, avoiding reliance on predefined geographic aggregation. Drawing on a large French insurance portfolio combined with high-resolution environmental variables, rainfall records, and institutional hazard maps, we compare a benchmark GLM with several discrete Bayesian specifications, including independent random effects, intrinsic conditional autoregressive (iCAR) and Besag-York-Mollie (BYM) models, and a continuously indexed Gaussian random field constructed using the stochastic partial differential equation (SPDE) approach. Inference is performed using Integrated Nested Laplace Approximation (INLA), enabling efficient estimation of latent spatial fields and non-linear covariate effects. Our results show that accounting for spatial dependence substantially improves occurrence modelling, while gains in severity prediction are more limited. The SPDE formulation further outperforms areal models by capturing sub-municipal risk gradients and reducing artefacts induced by arbitrary geographic partitioning. By conditioning on detailed building-level attributes, we isolate the contribution of latent spatial effects, refine the interpretation of observed covariates, and improve the allocation of risk premiums across the portfolio. In addition to enhanced predictive performance, the framework provides coherent uncertainty quantification and supports tail-risk assessment. To our knowledge, this is the first application of point-referenced SPDE models to flood insurance, offering a scalable statistical alternative for pricing and managing risks with strong spatial structure.
Conformal Prediction Through the Lens of Hypothesis Testing: Universality, Impossibility, and Optimality
oai:arXiv.org:2608.27310v2
arXiv:2608.27310v2 Announce Type: replace
Abstract: The connections between conformal prediction and permutation tests are already widely-known in the literature. Some authors motivate conformal prediction by saying that it computes a permutation p-value for the hypothesis $H_0 : Y_{n+1} = y$, and then inverts this to form a prediction set for $Y_{n+1}$ (i.e., accepts all values $y$ into the prediction set for which the p-value is large). In this paper, we examine an alternative view, which is less well-known: we again cast conformal prediction via the inversion of a permutation test, but for the null of exchangeability of the joint distribution of the $n+1$ samples. This change in perspective, while simple, adheres more closely to traditional formalization in hypothesis testing, which offers several benefits. First, we use the duality between conformal sets and testing to show that foundational universality and impossibility results in the conformal prediction literature can be reproduced directly using classical hypothesis testing theory (due to Neyman, Lehmann, Scheff{\'e}, Kraft, Le Cam, and others). Furthermore, we show that an optimality result for conformal prediction can be derived using standard Neyman-Pearson theory: for any joint distribution of the covariates and response $X,Y$, and any sample size, the optimal method for prediction sets---delivering the most efficient set among all methods with valid coverage for exchangeable distributions---is a conformal predictor whose score is the inverse conditional density of $Y|X$.
Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition
oai:arXiv.org:2609.20529v2
arXiv:2609.20529v2 Announce Type: replace
Abstract: We study simultaneous inference for maxima of canonical order-two $U$-statistics in high dimension. Degeneracy makes quadratic fluctuations leading, so ordinary Gaussian calibration can fail even after exact variance normalization. We show that the appropriate general target is a joint signed Gaussian quadratic chaos and establish a general approximation result that permits indefinite kernels. The general anti-concentration bound is too crude for high-dimensional inference, and we obtain sharper bounds under additional spectral structure. We also identify a phase transition from a non-Gaussian signed-chaos maximum to its covariance-matched Gaussian counterpart driven by the effective rank. For feasible inference, we propose a Gaussian multiplier bootstrap that avoid estimating eigensystems, and establish its validity. Two applications and extensive numerical simulations further illustrate the scope and practical performance of the proposed framework.
Riemannian Simultaneous Inference for Tangent Vector Field Regression
oai:arXiv.org:2609.21910v2
arXiv:2609.21910v2 Announce Type: replace
Abstract: We consider nonparametric tangent vector field regression on a Riemannian manifold without boundary. Because responses at different points lie in different tangent spaces, the proposed kernel estimator first parallel transports nearby responses to the target tangent space and then forms a volume-corrected local average. We first derive its uniform second-order bias, finite-bandwidth covariance, and stochastic rate. For simultaneous inference, the tangent norm is written as a supremum over the unit tangent bundle. Exact covariance whitening gives a unit-variance Gaussian field whose correlation length is of order $h$ along the base manifold and of order one along the fibre. Its local covariance geometry leads to a Gumbel limit with an explicit intrinsic constant. Combining this limit with Gaussian approximation and cross-fitted covariance estimation yields a feasible simultaneous confidence tube for the regression field. We further discuss improved finite-sample inference with bandwidth selection and high-order bias corrections. Simulations on various manifolds support the proposed inference procedure. A randomized reconstruction of global wind data illustrates how the tube's cross-sections describe spatially varying uncertainty.
Locally Private Inference for Riemannian Stochastic Optimization
oai:arXiv.org:2609.22642v2
arXiv:2609.22642v2 Announce Type: replace
Abstract: We develop inference for manifold-valued population minimizers when each observation belongs to a different participant and only locally private messages reach the analyst. The method releases randomized tangent gradients and combines them through Riemannian stochastic approximation and Polyak-Ruppert averaging. Directly inserting a private data surrogate into a nonlinear loss can shift its population target, whereas conditional centring of the released gradient preserves the first-order equation. We introduce symmetric-pair regression (SPR) to estimate the asymptotic variance from the same private messages used for point estimation, without holding out participants or requesting a second release. We prove the central limit theorem and consistency of the fully transcript-based sandwich covariance and intrinsic Wald region under local differential privacy. Simulations across various statistical problems and manifolds support the predicted decrease in estimation error and near-nominal coverage under moderate privacy. An application to NHANES anthropometric data illustrates private estimation of a leading body-size direction and its uncertainty.
Fear Moves Markets: Sentiment-Augmented POMP for Volatility Modeling of Bitcoin Returns
oai:arXiv.org:2609.23250v2
arXiv:2609.23250v2 Announce Type: replace
Abstract: Cryptocurrency markets exhibit extreme price swings and sentiment-driven regime shifts, which traditional volatility models often fail to capture. To address this, we develop a partially observed Markov processes (POMP) model augmented with sentiment and heavy-tailed distributions. Specifically, we extend Breto's framework by including the Fear and Greed Index (FGI) as an exogenous regressor in the latent volatility dynamics and replacing Gaussian measurement noise with a Student's t distribution. We fit the model via simulation-based inference using daily Bitcoin returns and FGI data from January 2020 to April 2025. Our approach significantly outperforms three benchmark models in log-likelihood and filter stability. Moreover, our model yields interpretable parameters consistent with known features of financial volatility. These results demonstrate that incorporating sentiment signals and heavy-tailed noise improves the modeling of volatility and regime shifts in cryptocurrency markets.
On Basis Function Selection for Sparse Gaussian Process Regression
oai:arXiv.org:2609.26624v2
arXiv:2609.26624v2 Announce Type: replace
Abstract: Sparse Gaussian processes achieve $O(N)$ inference by replacing the kernel with an appropriate expansion in a fixed basis $\{\phi_j\}$ on the input space. Given a compute budget $M \ll N$, practitioners conventionally truncate the basis to its first $M$ entries. Nothing in the formalism, however, prevents one from selecting only those $M$ basis functions that matter for the data at hand. This would avoid spending budget on basis functions where there is no signal, but it requires a criterion for ranking the candidates. We propose three such criteria derived from an information-theoretic view of the basis-function selection problem. Each criterion matches a different state of knowledge at selection time: a no-data state, a no-prior state, and an in-between state. We then study the performance of truncation versus selection strategies on six UCI regression benchmarks across three basis families: Hilbert-space Gaussian processes (HSGP), variational Fourier features (VFF), and variational inducing spherical harmonics (VISH). We observe that the no-data criterion is a safe default, matching or improving on truncation for HSGP, VFF and VISH, with substantial gains for VISH and improvements over a recently developed selection heuristic for that basis family. The data-aware no-prior and in-between criteria provide substantial gains over truncation specifically for HSGP, which is the most broadly used of the three families in practice.
Random Polytope Descriptors
oai:arXiv.org:2009.13987v3
arXiv:2009.13987v3 Announce Type: replace-cross
Abstract: We introduce a class of random polytopes which simultaneously generalizes several known constructions. While being fairly general, these polytopes are also computationally exceptionally benign. We indicate how these properties can be exploited for classification and clustering tasks in data analysis. Crucially, our construction lets users smoothly trade off between a tighter description of the data and faster computation.
Distribution-uniform strong laws of large numbers
oai:arXiv.org:2402.00713v3
arXiv:2402.00713v3 Announce Type: replace-cross
Abstract: We revisit the question of whether the strong law of large numbers (SLLN) holds uniformly in a rich family of distributions, culminating in a distribution-uniform generalization of the Marcinkiewicz-Zygmund SLLN. These results can be viewed as extensions of Chung's distribution-uniform SLLN to random variables with uniformly integrable $q^\text{th}$ absolute central moments for $0 < q < 2$. Furthermore, we show that uniform integrability of the $q^\text{th}$ moment is both sufficient and necessary for the SLLN to hold uniformly at the Marcinkiewicz-Zygmund rate of $n^{1/q - 1}$. These proofs centrally rely on novel distribution-uniform analogues of some familiar almost sure convergence results including the Khintchine-Kolmogorov convergence theorem, Kolmogorov's three-series theorem, a stochastic generalization of Kronecker's lemma, and the Borel-Cantelli lemmas. We also consider the non-identically distributed case.
Nonasymptotic and distribution-uniform Koml\'os-Major-Tusn\'ady approximation
oai:arXiv.org:2502.06188v3
arXiv:2502.06188v3 Announce Type: replace-cross
Abstract: We present nonasymptotic concentration inequalities for sums of independent and identically distributed random variables that yield asymptotic strong Gaussian approximations of Koml\'os, Major, and Tusn\'ady (KMT) [1975,1976]. The constants appearing in our inequalities are either universal or explicit, and thus as corollaries, they imply distribution-uniform generalizations of the aforementioned KMT approximations. In particular, it is shown that uniform integrability of a random variable's $q^{\text{th}}$ moment is both necessary and sufficient for the KMT approximations to hold uniformly at the rate of $o(n^{1/q})$ for $q > 2$ and that having a uniformly lower bounded Sakhanenko parameter -- equivalently, a uniformly upper-bounded Bernstein parameter -- is both necessary and sufficient for the KMT approximations to hold uniformly at the rate of $O(\log n)$. Instantiating these uniform results for a single probability space yields the analogous results of KMT exactly.
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
oai:arXiv.org:2503.02129v3
arXiv:2503.02129v3 Announce Type: replace-cross
Abstract: Path regularization has shown to be a very effective regularization to train neural networks, leading to a better generalization property than common regularizations i.e. weight decay, etc. We propose a first near-complete (as will be made explicit in the main text) nonasymptotic generalization theory for multilayer neural networks with path regularizations for general learning problems. In particular, it does not require the boundedness of the loss function, as is commonly assumed in the literature. Our theory goes beyond the bias-variance tradeoff and aligns with phenomena typically encountered in deep learning. It is therefore sharply different from other existing nonasymptotic generalization error bounds. More explicitly, we propose an explicit generalization error upper bound for multilayer neural networks with $\sigma(0)=0$ and sufficiently broad Lipschitz loss functions, without requiring the width, depth, or other hyperparameters of the neural network to approach infinity, a specific neural network architecture (e.g., sparsity), or boundedness of the loss function, while also taking approximation error into consideration. In particular, we solve an open problem proposed by Weinan E et. al. in 2020 regarding the approximation rates in generalized Barron spaces. Furthermore, we show the near-minimax optimality of our theory for regression problems with ReLU activations. Notably, our upper bound exhibits the famous double descent phenomenon for such networks, which is the most distinguished characteristic compared with other existing results. Our subsequent work will prove the matching lower bounds in the minimax sense, meaning that it is highly possible that our theory reveals the true underlying mechanism of the double descent phenomenon. We can also explain scaling law from this theory.
Localized Diffusion Models
oai:arXiv.org:2505.04417v3
arXiv:2505.04417v3 Announce Type: replace-cross
Abstract: Diffusion models are state-of-the-art tools for various generative tasks. Yet training these models involves estimating high-dimensional score functions, a task that in principle suffers from the curse of dimensionality. It is therefore important to understand how low-dimensional structure in the target distribution can be exploited in these models. Here we consider locality structure, which describes certain sparse conditional dependencies among the target random variables. Given some locality structure, the score function is effectively low-dimensional, so that it can be estimated by a localized neural network with significantly reduced sample complexity. This observation motivates the localized diffusion model, where a localized score matching loss is used to train the score function within a localized hypothesis space. We prove that such localization enables diffusion models to circumvent the curse of dimensionality with dimension-independent error bounds, at the price of additional localization error. Under realistic sample size scaling, we then show both theoretically and numerically that a moderate localization radius can balance the statistical and localization errors, yielding better overall performance. Locality structure also facilitates parallel training, making localized diffusion models potentially more efficient for large-scale applications.
Role of scrambling and noise in temporal information processing with quantum systems
oai:arXiv.org:2505.10080v3
arXiv:2505.10080v3 Announce Type: replace-cross
Abstract: Scrambling quantum systems have attracted attention as effective substrates for temporal information processing. Here we consider a quantum reservoir processing framework that captures a broad range of physical computing models with quantum systems. We examine the scalability and memory retention of the model with scrambling reservoirs modelled by high-order unitary designs in both noiseless and noisy settings. In the former regime, we show that measurement readouts become exponentially concentrated with increasing reservoir size, yet strikingly do not worsen with the reservoir iterations. Thus, while repeatedly reusing a small scrambling reservoir with quantum data might be viable, scaling up the problem size deteriorates generalization unless one can afford an exponential shot overhead. In contrast, the memory of early inputs and initial states decays exponentially in both reservoir size and reservoir iterations. In the noisy regime, we also prove that memory decays exponentially in time for local noisy channels. These results required us to introduce new proof techniques for bounding concentration in temporal quantum models. Beyond this extreme scrambling regime, we numerically demonstrate that exponential concentration can still exist even with a physical reservoir such as an Ising model whenever the reservoir operates in a quantum-chaotic phase. In contrast, physical reservoirs in a many-body localized phase and at the edge of chaos appear to not suffer from such phenomena
Randomization Inference with Sample Attrition
oai:arXiv.org:2507.00795v3
arXiv:2507.00795v3 Announce Type: replace-cross
Abstract: Randomization inference is a widely-used and appealing approach for analyzing treatment effects in randomized experiments, as it is finite-sample valid and does not require any distributional assumptions. However, naive application of randomization inference may suffer from severe size distortion in the presence of sample attrition, where outcome data are missing for some units. In this paper, we propose new, computationally efficient methods for randomization inference that remain valid under a broad class of potentially informative missingness mechanisms, allowing a unit's missingness to depend on its (unobserved) potential outcomes. Specifically, we construct valid p-values for testing both sharp and bounded null hypotheses on treatment effects via a worst-case consideration of the classical Fisher randomization test. Leveraging distribution-free test statistics, these worst-case p-values admit closed-form solutions. Importantly, by incorporating both potential outcomes and potential missingness indicators into the test statistic, our methods can exploit structural assumptions such as monotone missingness, which are commonly adopted in applications due to their plausibility and ability to substantially improve inferential power. Moreover, our approach connects to a range of partial identification bounds in the literature, which in some sense suggests the sharpness of our tests. We illustrate the proposed methods through both simulation studies and an empirical application. An R package implementing the proposed methods is publicly available.
InsurTech innovation using natural language processing
oai:arXiv.org:2507.21112v4
arXiv:2507.21112v4 Announce Type: replace-cross
Abstract: With the rapid rise of InsurTech, traditional insurance companies are increasingly exploring alternative data sources and advanced technologies to sustain their competitive edge. This paper provides both a conceptual overview and practical case studies of natural language processing (NLP) and its emerging applications within insurance operations, focusing on transforming raw, unstructured text into structured data suitable for actuarial analysis and decision-making. Leveraging real-world alternative data provided by an InsurTech industry partner that enriches traditional insurance data sources, we apply various NLP techniques to demonstrate feature de-biasing, feature compression, and industry classification in the commercial insurance context. These enriched, text-derived insights not only add to and refine traditional rating factors for commercial insurance pricing but also offer novel perspectives for assessing underlying risk by introducing novel industry classification techniques. Through these demonstrations, we show that NLP is not merely a supplementary tool but a foundational element of modern, data-driven insurance analytics.
BOCO: Bayesian Online Contextual Optimization for Decision-Focused Online Learning
oai:arXiv.org:2511.20413v2
arXiv:2511.20413v2 Announce Type: replace-cross
Abstract: \emph{Decision-focused learning} (DFL) trains predictive models to optimize downstream decisions rather than prediction accuracy alone. While recent studies have extended this paradigm to online settings with streaming data, existing online DFL methods generally maintain a point estimate, while their gradient-based updates require either a differentiable optimization layer or a problem-specific surrogate loss. Consequently, they can be unstable under limited data and difficult to apply across heterogeneous optimization problems. We introduce Bayesian Online Contextual Optimization (\texttt{BOCO}), a framework that maintains a decision-focused posterior over model parameters. \texttt{BOCO} aggregates the resulting predictions when prescribing decisions, thereby accounting for parameter uncertainty. To track this posterior in evolving environments, we develop two particle-based inference algorithms: a sequential Monte Carlo sampler for nondifferentiable problems and a function-space Stein variational gradient descent algorithm for differentiable problems. Across both real-world tasks, \texttt{BOCO} reduces running mean regret and temporal regret variability relative to two frequentist online DFL baselines. At the full horizon, relative to the best frequentist baseline, BOCO reduces running mean regret and temporal regret variability by 6.9\% and 6.3\% on knapsack and by 49.3\% and 38.1\% on energy scheduling, respectively. The gains are larger early in the data stream, a pattern consistent with a benefit from accounting for parameter uncertainty when data are limited.
Advances in Diffusion-Based Generative Compression
oai:arXiv.org:2601.18932v2
arXiv:2601.18932v2 Announce Type: replace-cross
Abstract: Popularized by their strong image generation performance, diffusion and related methods for generative modeling have found widespread success in visual media applications. In particular, diffusion methods have enabled new approaches to data compression, where realistic reconstructions can be generated at extremely low bit-rates. This article provides a unifying review of recent diffusion-based methods for generative lossy compression, with a focus on image compression. These methods generally encode the source into an embedding and use a diffusion model to iteratively refine it during decoding, so that the reconstruction approximately follows the true data distribution. The embedding can take various forms and is typically transmitted via an auxiliary entropy model, and recent methods also explore the use of diffusion models themselves for information transmission via channel simulation. We review representative approaches through the lens of rate-distortion-perception theory, highlighting the role of common randomness and connections to inverse problems, and identify open challenges.
Learning to Approximate Uniform Facility Location via Graph Neural Networks
oai:arXiv.org:2602.13155v3
arXiv:2602.13155v3 Announce Type: replace-cross
Abstract: Neural networks, particularly message-passing neural networks (MPNNs), are increasingly used as heuristics for hard combinatorial optimization problems. Yet many learning-based methods rely on supervision, reinforcement learning, or gradient estimators, causing high computational cost, unstable training, or limited guarantees. Classical approximation algorithms provide worst-case guarantees but are non-differentiable and cannot adapt to structure in natural input distributions. We study this tradeoff through Uniform Facility Location (UniFL), a problem with applications in clustering, summarization, logistics, and supply chains. We propose a fully differentiable MPNN that incorporates approximation-algorithmic principles without solver supervision or discrete relaxations. The model has provable approximation guarantees and empirically improves on standard approximation algorithms, narrowing the gap to integer linear programming.
The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
oai:arXiv.org:2603.18482v4
arXiv:2603.18482v4 Announce Type: replace-cross
Abstract: Why does machine-generated text remain detectable? We investigate a mechanistic explanation at the decoding stage: standard strategies such as top-$k$ and nucleus sampling restrict generation to high-probability tokens, while human writers routinely choose contextually appropriate words from deeper in the model's probability distribution. Truncation makes a measurable share of these choices unreachable; we call this the \emph{truncation blind spot}. Across five open models and three domains, 8--18\% of human-selected tokens fall outside common truncation boundaries. Linguistic analysis further reveals disproportionate exclusion of content-word tokens. In a benchmark comprising 1.8 million machine generations, classifiers using only predictability and lexical diversity achieve mean AUC-ROC near 0.97, with substantial variation across decoding settings and strong transfer across generators. Probability-floor samplers substantially narrow the blind spot, demonstrating that the choice of truncation criterion matters for retaining human-used tokens. Together, these findings characterize a source of human--machine distributional mismatch and motivate decoding methods that preserve contextually appropriate low-probability choices while maintaining generation quality. Code and data are available at https://github.com/EstebanGarces/human_vs_machine.
ProteinJEPA: Latent prediction improves protein language model pretraining
oai:arXiv.org:2605.07554v2
arXiv:2605.07554v2 Announce Type: replace-cross
Abstract: Protein language models are trained primarily with masked language modeling (MLM), which predicts masked amino-acid identities. Joint-embedding predictive architectures (JEPA) instead predict latent representations, but have not been applied to proteins.
ProteinJEPA supplements MLM with a cosine loss for predicting the half-depth hidden states of a teacher given the unmasked sequence. On 19 tasks, with ESM2 at 35M and 150M parameters and three pretraining seeds, MLM+JEPA outperforms compute-matched and step-matched MLM-only continued training in 78 and 76 of 114 comparisons (14 losses, 22 ties). The median compute-matched gain is $+0.0106$ on structure- and homology-sensitive tasks versus $+0.0041$ elsewhere, led by SCOPe-40 retrieval and remote homology with improvements of 6.1 percentage points in Recall@1 and 2.7 points in accuracy, respectively. Gains on these tasks increase with model size from 8M to 150M. Against the off-the-shelf checkpoint, MLM+JEPA wins 81 of 114 comparisons (median $+0.0068$) without improving MLM loss.
In random initialization the gain is smaller and replicates inconsistently across seeds ($p{=}0.059$). The same recipe improves the causal ProGen3 model, beating a compute-matched next-token-prediction control on 12 of 16 tasks. Ablations show that cosine loss beats mean squared error, while adding shallower targets removes most of the task gain.
JEPA-only training collapses downstream performance: latent prediction complements MLM rather than replacing it. Code: https://anonymous.4open.science/r/protJepa-FF24
A lift for input-convex neural net training
oai:arXiv.org:2605.24274v2
arXiv:2605.24274v2 Announce Type: replace-cross
Abstract: Input-convex neural nets parametrize the convex potentials of density models and transport maps, and their convexity requires the inter-layer weights to be non-negative. Projected gradient descent enforces this by projecting after each step, and due to mini-batch noise the boundary is re-crossed indefinitely, which leads to an active set the projection never identifies. The differentiable alternative, direct softplus, optimizes a free latent weight through a softplus positivity map whose derivative attenuates the gradient exponentially where the weight is negative---the shoulder---so a coordinate that reaches it stays for an exponentially long time. To keep this unconstrained parametrization without its slow escape, we propose the lift, which replaces the free latent weight by a learnable slack plus an unconstrained network---the body---that takes a permutation-invariant summary of the training batch as input. The latent weight thus varies with the batch before the positivity map, and couples to the gradient formed on it. We show that this coupling enters the variance of the update to the latent weight at first order in the fluctuation, and that the slack, the batch dependence and the shared batch are each needed for it to act. Where the coupling aligns positively with the loss curvature, that variance is larger under the lift than under direct softplus, and a coordinate leaves the shoulder sooner. We compare the lift with the two existing methods on several applications. Where a constrained weight of direct softplus reaches the shoulder and does not leave, the lift fits the target more closely and reaches the same reconstruction about three times sooner. Where almost none reaches it, the methods agree.
Empirical Trajectory Sparsity Biases Mobility-Informed Epidemic Modeling
oai:arXiv.org:2605.31282v2
arXiv:2605.31282v2 Announce Type: replace-cross
Abstract: GPS mobility data are increasingly used in epidemic modeling, allowing the construction of co-location networks or population flows. These trajectories typically exhibit high temporal sparsity because data collection is opportunistic and tied to phone use. Despite growing awareness of this limitation, the analysis and treatment of biases derived from it have been largely overlooked in existing epidemic modeling studies, raising concerns about the robustness of downstream inferences. We introduce a principled framework to quantify the impact of trajectory sparsity on key epidemic modeling outcomes across different levels of data missingness. Our approach leverages a highly complete dataset that exhibits both near-complete and sparse GPS trajectories. Near-complete trajectories provide baseline epidemic outcomes, while sparse trajectories provide realistic missingness patterns that we impose on the baseline to measure bias. In this way, we show how missing records can result in substantial underestimation of key measures of epidemic intensity, explained not only by the amount of missing data, but by more complex features of data missingness that should be taken into account when designing correction methods. Finally, we propose and evaluate a correction based on inverse probability weighting of the contact network before epidemic model calibration, which is shown to reduce bias and parameter misspecification. We also demonstrate this correction on a separate anonymized sample from a commercial GPS mobility dataset and report on its effect. Together, our findings provide a first rigorous quantification of trajectory-sparsity bias in epidemic modeling, offering initial guidance on the treatment of this issue.
Energy-guided Recursive Model
oai:arXiv.org:2607.10128v3
arXiv:2607.10128v3 Announce Type: replace-cross
Abstract: Recursive models show promise on reasoning and language tasks, yet their test-time scaling lacks a principled criterion for selecting trajectories or determining recurrent depth. We introduce \textbf{Energy-guided Recursive Model (ERM)}, which uses Hopfield-type memories of valid local and global structures to assign intrinsic energies to candidate trajectories. These energies guide candidate selection and suggest an effective range of recurrent depths, implying that deeper recurrence does not necessarily improve reasoning accuracy. They also enable sampling methods such as parallel tempering to improve exploration. For reasoning tasks, ERM achieves optimal solutions on Sudoku ($98.97\%$), Pencil Puzzle Bench (PPBench, $88.04\%$) and Maze ($99.30\%$), reaching the best accuracy in recursive modeling. On language modeling, ERM reduces RedPajama-V2 perplexity by $1.74\%$ with marginal inference overhead. The results support energy guidance as a practical framework for improving test-time scaling in recursive models.
Bootstrapping autoregressive duration models
oai:arXiv.org:2607.28294v2
arXiv:2607.28294v2 Announce Type: replace-cross
Abstract: This paper develops bootstrap methods for likelihood-based inference in autoregressive conditional duration (ACD) models, where the sample size is endogenously determined by durations observed over a fixed time span. This feature fundamentally shapes the asymptotic framework, particularly so when the durations do not have finite expectation. Building on recent limit theory for heavy-tailed and integrated ACD processes, we analyze recursive bootstrap schemes that either fix the time span (yielding a random sample size) or fix the number of durations (yielding a random time span). We establish a bootstrap theory for ACD models that links naturally to renewal theory with random sample sizes. For the fixedcount bootstrap, we prove first-order validity in the finite-mean and boundary cases and characterize the random limiting bootstrap distribution in the infinite-mean case. Although classical bootstrap consistency can fail when the durations have infinite expectation, we argue that the bootstrap remains valid and yields asymptotically normal t-statistics. Monte Carlo evidence shows that the proposed methods have good finite-sample properties in both finite- and infinite-mean settings, and are robust to distributional misspecification relative to the exponential likelihood. We conclude with an empirical application to cryptocurrency ETFs.
TabPFN-3.5: Technical Report
oai:arXiv.org:2609.17895v2
arXiv:2609.17895v2 Announce Type: replace-cross
Abstract: We introduce TabPFN-3.5, our new flagship Tabular Foundation Model. It significantly outperforms its predecessor, TabPFN-3, and all existing baselines across a broad range of tabular problems. TabPFN-3.5 sets a new state of the art on standard tabular prediction in TabArena, and extends it to the data practitioners encounter in practice: non-i.i.d. data with temporal or grouped splits, tables with strings, text and images, high-cardinality categorical features, and wide tables with many features. These gains carry over to our task-specific harnesses: state of the art on relational data and stronger time-series forecasting. For faster inference, our variant TabPFN-3.5-Fast runs up to 3x faster than TabPFN-3 while keeping most of the accuracy gains. In addition, we upgrade TabPFN-3.5-Plus, expanding our multimodal capabilities with advanced text and date handling alongside proprietary inference optimizations. Finally, we release a new version of our Thinking mode, TabPFN-3.5-Thinking, which scales inference-time computation to push the state of the art further. It benefits from our stronger base model and from inference-time improvements that make it up to 12x faster than TabPFN-3-Thinking.
Learn Your Own Thoughts: Abstract Token Curriculum
oai:arXiv.org:2609.19717v2
arXiv:2609.19717v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token Curriculum (ATC), a novel curriculum learning framework that elicits effective continuous intermediate representations without direct supervision or manual scratchpad design. ATC gradually increases problem complexity through a sequence of distributions, training the model to develop internal abstract ``thoughts'' in the continuous representation space. This paper provides both theoretical and experimental evidence for the benefits of ATC and its advantages over previous methods for training continuous thoughts. Theoretically, we show that for learning parity functions with single-layer softmax attention using ATC, attention naturally focuses on the CoT tokens in the context that provide the ``easiest path'' to predicting the next token. Experimentally, we show ATC's effectiveness on graph reachability and arithmetic learning tasks.