• raseliarison
  • nirinA
  • adrien
  • blog
  • code
  • FAQ
  •  home  
  •  news  
    • arXiv
      • astro-ph
      • cond-mat
      • cs
      • eess
      • gr-qc
      • hep-ex
      • hep-lat
      • hep-ph
      • hep-th
      • math
      • math-ph
      • nlin
      • nucl-ex
      • nucl-th
      • physics
      • q-bio
      • quant-ph
      • stat
    • physics
      • phys.org
      • physics world
    • linux
      • kernel
      • slackware
    • nature
      • natcomputsci
      • natastron
      • natbiomedeng
      • nenergy
      • nnano
      • natmachintell
      • nbt
      • nmeth
      • natecolevol
      • nmicrobiol
      • ng
      • nchembio
      • natelectron
      • micronano
      • nphoton
    • bioRxiv
    • plos one
    • world
      • BBC
      • Al Jazeera
    • earth
      • earth observatory
      • weather
      • weather forecast
    • universe
      • apod
      • hubble
      • atel
      • nasa
  •  wiki  
  •  gemini  
  •  python  
  • eess updates on arXiv.org

    eess updates on the arXiv.org e-print archive.

    Impact-Time Guidance via Normal Contraction to a Time-to-Go Isochron

    oai:arXiv.org:2609.26906v1

    arXiv:2609.26906v1 Announce Type: new Abstract: We develop a contraction-based perspective on impact-time guidance that augments a baseline homing command with a timing bias. The proposed perspective treats the prescribed schedule as a moving time-to-go isochron and regulates motion normal to that set through velocity-normal lateral acceleration while the interceptor's speed remains constant. We derive a transport equation that characterizes homing-compatible time-to-go coordinates and define a predictor defect that quantifies the mismatch of approximate maps. We show that the scalar timing channel induces a coordinate-invariant rank-one metric on the normal quotient. To account for bounded lateral acceleration, we formulate a robust scalar filter and derive a necessary and sufficient condition for pointwise feasibility. We then show that terminal calibration and funnel invariance establish first interception at the prescribed time under the stated assumptions. We also develop a preterminal alignment and homing handover that avoids singular inversion as lateral timing authority vanishes near collision-course alignment. The proposed perspective accommodates analytic, numerical, and learned time-to-go maps that satisfy the required calibration and regularity conditions.

    https://arxiv.org/abs/2609.26906


    A Temporal-Envelope Frontend with Learnable Per-Channel Energy Normalization for Whisper-Based Children's ASR

    oai:arXiv.org:2609.26937v1

    arXiv:2609.26937v1 Announce Type: new Abstract: Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children's speech, where high acoustic variability demands robust feature representations. We propose a modular time-domain frontend that decomposes speech into sub-band envelopes using mel-spaced windowed-sinc filters and the Hilbert transform, with learnable per-channel energy normalization (PCEN) jointly optimized with the Whisper model. On the MyST children's speech corpus, systematic ablations identify full-band windowed-sinc filters, Hilbert envelopes, a 25 Hz smoothing cutoff, and learnable PCEN as the best configuration. Under the same Whisper-small fine-tuning setup, the frontend reduces WER from 13.16% to 11.08%, a 15.8% relative reduction over the log-mel baseline, and outperforms the evaluated Kid-Whisper checkpoint on the same cleaned test split. These results show that temporal-envelope representations and learnable frontend normalization are effective complements to backend adaptation for children's ASR.

    https://arxiv.org/abs/2609.26937


    Untangling the Geometry and Speed for RF Sensing Spectrograms

    oai:arXiv.org:2609.26960v1

    arXiv:2609.26960v1 Announce Type: new Abstract: A fundamental challenge in RF sensing is that Doppler signatures observed by a link entangle the target's motion with the sensing geometry, resulting in limited applicability to unconstrained real-world settings. In this paper, we establish a new foundation for physically interpretable RF sensing that disentangles reflector speed from geometry, jointly recovering the speed, geometry factor, relative amplitude, and width of each dominant Doppler ridge. More specifically, we first develop a compact parametric representation of WiFi spectrograms and establish its low-dimensional structure through a systematic computer-vision analysis of a large and diverse human-activity dataset, thereby providing a tractable foundation for learning. Building on this representation, we then design a physics-informed autoencoder whose structured bottleneck and differentiable RF forward model enforce physically meaningful estimates of reflector speed and geometry. We further introduce a synthetic-to-real training framework, eliminating the need for real WiFi training data. We extensively validate the proposed framework under both known and time-varying geometries, using both independently generated synthetic test sets and 31 real WiFi experiments. The results demonstrate the superior performance in speed and geometry extraction, robustly recovering the underlying geometry, speeds, Doppler-ridge amplitudes, and ridge widths across all settings, while substantially outperforming the strongest baselines.

    https://arxiv.org/abs/2609.26960


    Three High Performance Global Tracking Composite Adaptive Controllers for Fully Actuated Euler-Lagrange Systems: Experimental Validation

    oai:arXiv.org:2609.27010v1

    arXiv:2609.27010v1 Announce Type: new Abstract: Three adaptive global tracking controllers for fully actuated Euler-Lagrange systems, with verifiable performance improvement over existing designs, are reported in this letter. Two of these controllers ensure global exponential convergence under a weak interval excitation condition. Besides, one of the proposed controllers features a simple adaptive PID-like structure that-unlike classical solutions-avoids the need for additional filtering. We adopt a composite adaptation architecture, invoke a novel parameterization of the system dynamics and use a high performance estimation scheme recently introduced in the literature. Real-time experiments and a comparative study with a learning-based adaptive controller on a two-degrees-of-freedom manipulator arm illustrate the effectiveness of the proposed controllers.

    https://arxiv.org/abs/2609.27010


    Digital Twin Enhanced Channel Twin for AI-Native CSI Inference: Generalizability and Scalability

    oai:arXiv.org:2609.27017v1

    arXiv:2609.27017v1 Announce Type: new Abstract: Accurate channel state information (CSI) is critical for advanced multi-antenna wireless networks. While high-fidelity and site-specific ray-tracing (RT) equipped wireless digital twins can overcome overhead for CSI acquisition. However, computing deterministic, calibrated RT based CSI from a wireless digital twin for every orthogonal frequency-division multiplexing (OFDM) symbol violates the strict microsecond latency budgets of the 5G NR numerology. To overcome this computational bottleneck, this paper investigates three approaches within a calibrated 3D digital twin framework, namely (i) the generalization of the channel twin, (ii) the performance enhancement using a neural receiver, and (iii) a data-driven interpolation framework for scalability. We compute high-precision CSI for a sparse subset of temporal anchors and employ an attention-based Transformer to predict the remaining symbols. Unlike polynomial splines or sequential long short-term memory (LSTM) networks, the Transformer exploits a global receptive field to capture the non-linear multipath dynamics while enabling parallelizable, real-time inference. To achieve spatial scalability, we introduce channel twin generalization. By fine-tuning the 3D RT models alongside the Transformer, the framework leverages the CSI dataset of one location to infer the channel behavior of an unseen environment. Transfer learning further adapts the model to a new environment using only a small fraction of locally collected data. Simulation results demonstrate that the proposed architecture substantially outperforms the baseline interpolators, achieves robust spatial transferability, and lowers the bit error rate through the unified neural receiver. These results establish a scalable, environment-agnostic foundation of distributed channel twins for high-quality, low-overhead CSI acquisition in next-generation cellular networks.

    https://arxiv.org/abs/2609.27017


    Dual-Microphone Steerable High-Order Neural Differential Beamformer

    oai:arXiv.org:2609.27021v1

    arXiv:2609.27021v1 Announce Type: new Abstract: Linear arrays of omnidirectional microphones produce beampatterns that are symmetric about the array axis. For these arrays, a beamformer is considered steerable if its beampattern maintains the same shape in the semicircular plane across all look directions from 0{\deg} to 180{\deg}. For dual-microphone arrays, conventional differential beamformers are generally non-steerable and restricted to first-order, which significantly limits spatial selectivity. To address these limitations, this study presents a neural differential beamformer (NDBF) with a dual-microphone array. The contributions are as follows: (i) NDBF is steerable; (ii) NDBF achieves high-order frequency-invariant beampatterns; and (iii) NDBF enables stereo recording using only two closely spaced omnidirectional microphones. Experimental results demonstrate that NDBF outperforms existing methods while overcoming the limitations of classical differential beamforming.

    https://arxiv.org/abs/2609.27021


    Merging Large Language Models and Battery Physics for User-Aware Electric Vehicle Driving Management

    oai:arXiv.org:2609.27050v1

    arXiv:2609.27050v1 Announce Type: new Abstract: Electric vehicle (EV) battery performance is strongly coupled with driver behavior, yet human intent is typically expressed semantically rather than numerically. This paper proposes a hybrid physics-artificial intelligence framework that integrates a Large Language Model (LLM) as a high-level behavioral reasoning layer within a physics-driven supervisory architecture. The LLM interprets textual user intent and structured battery feedback to generate bounded behavioral parameters that shape a discharge current envelope. A physics-driven safety filter then enforces physical safety constraints before computing feasible velocity recommendations. Lyapunov-based analysis establishes bounded recommendation error under battery model and prompt inaccuracies. Simulation results demonstrate adaptive, user-aware operation without compromising physical safety. The proposed reasoning-enforcement architecture provides a principled pathway for safe AI integration in EV energy management.

    https://arxiv.org/abs/2609.27050


    Evaluating Grid Strength with Rising Penetration of Inverter Based Resources

    oai:arXiv.org:2609.27066v1

    arXiv:2609.27066v1 Announce Type: new Abstract: Over the years, Dominion Energy Virginia (DEV) has experienced recurring power quality disturbances, including voltage and power oscillations, particularly in areas with significant inverter-based generation. While these phenomena suggest a potential relationship to system strength, no definitive correlation has yet been established. This paper presents methodologies for assessing system strength across the DEV network and identifying regions that may be vulnerable to such disturbances. To this end, three grid-strength metrics are evaluated: the Simple Short-Circuit Ratio (SSCR), the Weighted Short-Circuit Ratio (WSCR), and the Composite Short-Circuit Ratio (CSCR). Historical event data will be analyzed using steady-state snapshots from DEV's proprietary Analysis on Demand (ANODE) platform, which provides 10-minute interval load-flow data. These snapshots will be processed using a developed software tool to compute system strength metrics. The analysis will investigate potential correlations between system strength and observed disturbances. Where such correlations are identified, the study will further evaluate the impact of deploying synchronous condensers (SynCons) as a mitigation strategy.

    https://arxiv.org/abs/2609.27066


    Automated feature-region selection for soft-sensor development from spectral-like measurements

    oai:arXiv.org:2609.27082v1

    arXiv:2609.27082v1 Announce Type: new Abstract: Sustainable production increasingly relies on process analytical chemistry and process analytical technology to support monitoring, control, and automation. Techniques used in these contexts, including Raman spectroscopy, infrared spectroscopy, and electrochemical voltammetry, generate high-dimensional signals ordered along physical measurement axes, referred to here as spectral-like measurements. These signals can contain redundant, weakly informative, and noisy regions, complicating the development of data-driven soft sensors to map them to process variables. Selecting informative regions is nontrivial, as visually prominent regions are not necessarily the most predictive, while synergistic effects among regions cannot be readily inferred. Here, we introduce an automated feature-region selection framework that identifies a parsimonious set of contiguous regions while preserving channel ordering. The framework combines channel-level target correlation and a signal-to-noise indicator into a latent information fingerprint. Variable-width candidate intervals are derived from peaks in this fingerprint, and their combinations are ranked using held-out validation data. Final selection favours models using fewer channels among candidates with comparable predictive performance. The framework is demonstrated using cyclic voltammetric measurements of glucose acquired with a gold electrode sensor across four progressively broader nominal concentration ranges up to 350 g/L. Gaussian process regression is used within the framework to accommodate possible nonlinear relationships, account for observation noise, and quantify epistemic uncertainty. The selected models retained only 11-42 of the original 400 channels and reduced held-out test root mean squared error by 88.3-94.6% relative to full-feature Gaussian process regression benchmarks.

    https://arxiv.org/abs/2609.27082


    AI-Enabled Wireless Propagation Modeling and Radio Environment Maps for 5G Aerial Wireless Networks

    oai:arXiv.org:2609.27083v1

    arXiv:2609.27083v1 Announce Type: new Abstract: With the gaining prominence of aerial mobility applications, their success depends on the seamless integration of terrestrial and non-terrestrial network connectivity. However, providing reliable connectivity from terrestrial telecommunication networks remains challenging due to multi-cell interference from base stations (BSs) under line-of-sight (LoS) conditions to unmanned aerial vehicles (UAVs), coverage holes caused by antenna sidelobe degradation, localized multipath fading effects, and the high-speed dynamics of aerial users. To model such complexities, often exacerbated by sparse real-world data, this work proposes a dual-stage radio environment map (REM) framework. Our approach physically decouples the channel modeling, where a spatial Transformer first anchors the deterministic, large-scale path loss geometry, while a gated recurrent unit (GRU) subsequently extrapolates the stochastic, localized fast- fading deviations. By reformulating 3D spatial interpolation as a 1D radial sequence prediction task, the framework inherently aligns with the physics of propagation. We evaluate the proposed framework against state-of-the-art baselines, including 3D Kriging, UNet, Mamba, and Inception, using empirical 5G datasets. The results demonstrate improved intra-site generalization across diverse altitudes, user dynamics, and reference signal received power (RSRP) datasets, achieving signal-strength predictions with errors near 3 dB and REM spatial-similarity indices exceeding 0.75. Finally, we examine the influence of REMs and channel rank conditions on UAV channel quality, underscoring the necessity of reliable channel modeling for robust aerial connectivity.

    https://arxiv.org/abs/2609.27083


    GCN-Based Model for Fault Classification in Distribution Networks under Harmonic Distortion

    oai:arXiv.org:2609.27169v1

    arXiv:2609.27169v1 Announce Type: new Abstract: The increasing integration of distributed energy resources introduces new challenges for accurate fault detection and classification in distribution networks. This paper presents a topology-aware graph convolutional network-based model for fault classification under inverter-induced harmonic distortion. The model uses pre-fault and post-fault voltage and current phasors, including harmonic components, mapped to the IEEE 34-bus system topology. The simulation data generated in OpenDSS encompass varying levels of penetration of solar photovoltaics (30%, 50%, 70%) and current total harmonic distortion (0%, 1%, 3%, 5%) across multiple fault types: single line-to-ground, line-to-line, double line-to-ground, and three-phase-to-ground. A differential evolution algorithm is employed to optimize the hyperparameters of the network, resulting in a three-layer model that achieves an F1-score of 0.98. Stress testing across scenarios confirms high robustness and minimal class confusion. The results demonstrate that a graph convolutional network-based model can effectively classify distribution faults under varying photovoltaic penetration levels and harmonically distorted conditions. Comparative analyses with convolutional neural network, long short-term memory, and multi-layer perceptron architectures demonstrate that the proposed model delivers superior classification performance, particularly under the most stressed condition with 70% photovoltaic penetration and 5% harmonic distortion.

    https://arxiv.org/abs/2609.27169


    Reshaping Converter-Network Interactions in Microgrids: From Virtual Impedance to Virtual Two-Port Control

    oai:arXiv.org:2609.27223v1

    arXiv:2609.27223v1 Announce Type: new Abstract: Converter terminal characteristics are central to both dynamic interactions with the network and steady-state power sharing in inverter-based microgrids. Conventional additive virtual impedance (VI) shapes these characteristics through a single virtual branch. This paper proposes virtual two-port control, which uses four coordinated transfer channels to reconstruct the terminal impedance. The connected network consequently observes the original converter through a virtually inserted twoport interface, with conventional virtual impedance recovered as a degenerate case. The proposed control transforms the original impedance through a matrix linear-fractional map. We also derive a necessary and sufficient condition for the reconstruction to be stable and proper. Among its various potential applications in microgrids, we develop three representative ones in detail: passivation with reduced control effort, uncertainty compression, and seriesshunt power-flow regulation. Simulations of VSG-controlled converters demonstrate the advantages of the proposed method over conventional VI in these applications.

    https://arxiv.org/abs/2609.27223


    Positioning a Movable Antenna Without CSI: A Kernelized Bandit Under Costly Movement

    oai:arXiv.org:2609.27226v1

    arXiv:2609.27226v1 Announce Type: new Abstract: Movable antenna (MA) systems reconfigure the wireless channel by repositioning the antenna within a confined region. The position optimization behind this capability has been studied under the premise that the channel at every candidate position is known. However, acquiring that knowledge is costly: a position can be measured only after the antenna has moved there, which takes many slots at the speed of the actuator, and the channel drifts meanwhile. We formulate MA positioning as a non-stationary kernelized bandit with a reachability constraint, in which exploration, tracking, and the reward lost in transit are coupled through a single physical action. We propose MoveUCB, which addresses the reachability constraint through a global target search, a persistence rule that holds the target across slots, and a transit-aware movement cost term. We prove that MoveUCB attains sublinear dynamic regret under the standard variation-budget condition of non-stationary bandits alone, with the physical parameters entering only through constants. Simulations against a dynamic-programming oracle and movement-unaware baselines show that MoveUCB attains the lowest regret with about half the travel of the closest baseline that moves.

    https://arxiv.org/abs/2609.27226


    One-Step Voice Conversion by Learning kNN Transport in WavLM Space

    oai:arXiv.org:2609.27230v1

    arXiv:2609.27230v1 Announce Type: new Abstract: Voice conversion (VC) systems fall into two families: non-parametric embedding-space methods, which need no trained model but degrade on short target utterances, and spectrogram-based neural architectures, which achieve strong quality via multi-module pipelines with tens of millions of parameters. We propose kNN-FM-VC, a single conditional flow-matching network that learns to approximate the kNN-VC mapping between WavLM embedding distributions of source and target speakers, replacing explicit pointwise kNN matching with a neural regressor trained on kNN-generated pairs. The model is conditioned on the target speaker via cross-attention and FiLM, and trained under three Gaussian conditional paths (Schr\"odinger bridge, straight line, and constant-variance Gaussian tube), enabling few-step sampling. Unlike Phoneme Hallucinator, which uses an upsampling stage followed by kNN matching, our 13M-parameter model performs conversion with a single learned network and supports one-step inference. On LibriSpeech, the one-step Gaussian Bridge achieves lower WER and higher estimated speech quality than FreeVC and Phoneme Hallucinator. Relative to kNN and kDOT, it substantially reduces WER.

    https://arxiv.org/abs/2609.27230


    JASPER: Joint Audio and Speech Pre-trained Encoder Representations

    oai:arXiv.org:2609.27260v1

    arXiv:2609.27260v1 Announce Type: new Abstract: Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.

    https://arxiv.org/abs/2609.27260


    Beyond DER: Speaker Counting in Crowded End-to-End Diarization

    oai:arXiv.org:2609.27315v1

    arXiv:2609.27315v1 Announce Type: new Abstract: Speaker diarization must solve two problems: counting how many speakers are present in a given conversation, and assigning speech to each one. This becomes harder in crowded conversations with five or more speakers. We evaluate how end-to-end neural diarization models count speakers, and find systematic under-counting in crowded recordings. The standard diarization error rate (DER) hides this failure, because it is duration-weighted and barely penalizes the dropped, low-activity speakers. We therefore also report the Jaccard error rate (JER) and explicit counting metrics. We propose a gated loss that couples speaker existence with frame activity. This loss can be computed in two ways, duration-weighted or speaker-weighted, and the speaker-weighted variant, used as a regularizer, reduces the under-count. On real-world crowded recordings our method clearly improves the counting metrics, lowering JER by about 6% relative and DER by about 9% relative, while leaving sparse recordings unharmed.

    https://arxiv.org/abs/2609.27315


    Evolving Inspectable O-RAN Slicing xApps with LLMs

    oai:arXiv.org:2609.27337v1

    arXiv:2609.27337v1 Announce Type: new Abstract: Open RAN (O-RAN) slicing xApps must adapt resource allocations to changing channel conditions and traffic demands while meeting service-level agreements (SLAs). Deep reinforcement learning can produce adaptive policies, but their allocation rules remain encoded in neural-network parameters. Our goal is to retain this adaptability while making the controller's decision logic directly inspectable and editable by operators. We use a large language model (LLM) to evolve slicing controllers as compact Python programs whose decision logic remains readable and editable after optimization. The LLM proposes and revises candidates offline, while a calibrated simulator scores them, and the selected decision module runs unchanged in the O-RAN control path. On the NSF POWDER 5G testbed, the evolved controller releases resources from a guaranteed slice whose throughput target becomes unattainable under a sustained channel fade, improving best-effort throughput from 158.2 to 228.6 Mbps, a 44.5% gain over the best static allocation. Since the controllers are readable source code, their behavior can be predicted from their equations, defects can be diagnosed by reading the code, and calibration errors can be corrected with one-line edits, reducing SLA misses from 79.9% to 2.2% in one case and more than doubling fitness in another. In a four-slice trace-driven simulation calibrated to the same testbed, evolutionary search achieves higher average evaluation scores than independent prompting at a matched proposal budget, with mean normalized gains on held-out traces of 16.3% for prompting alone, 32.1% for evolution from scratch, and 51.0% for evolution from a starting program.

    https://arxiv.org/abs/2609.27337


    Dynamic Sensor Pairing for TDOA-Based Target Tracking via mixed-integer Second-Order Cone Programming

    oai:arXiv.org:2609.27343v1

    arXiv:2609.27343v1 Announce Type: new Abstract: We propose a sensor pairing method for online target tracking based on time-difference-of-arrival (TDOA) measurements through a sensor network. Existing pairing designs fix active sensor pairs in advance. In online tracking, however, adapting the pairing to the current target estimate maintains a favorable sensor-target geometry around it. We select $K$ pairs at each time step to maximize a Fisher information matrix (FIM) criterion under a communication-degree constraint. The original combinatorial problem is cast as a mixed-integer second-order cone program. We solve it at every time step to obtain a globally optimal pairing. Experiments under various noise settings show that the proposed method obtains the lowest tracking error among the methods under the same communication-degree constraint.

    https://arxiv.org/abs/2609.27343


    Single-RF-Chain Multiuser Semantic Communications via Metasurface Modulation

    oai:arXiv.org:2609.27361v1

    arXiv:2609.27361v1 Announce Type: new Abstract: Multiuser transmission with multiple radio-frequency (RF) chains enables flexible spatial multiplexing but incurs increased hardware complexity and power consumption. A programmable metasurface (MTS) provides an alternative means of introducing spatial degrees of freedom with a single-RF-chain access point (AP). This paper proposes \emph{Semantic Prism}, an MTS-enabled framework for concurrently transmitting independent semantic messages to multiple users. Semantic Prism jointly maps multiuser semantic representations to a common AP-symbol sequence and a corresponding sequence of discrete MTS configurations, enabling each user to recover its intended message from the received signal sequence. This mapping is realized by a learning-based modulator that takes both the multiuser semantic representations and channel state information (CSI) as inputs, while accounting for the discrete phase-control constraint of the MTS. To facilitate practical deployment, we further develop a field-channel fine-tuning method that enhances robustness to CSI uncertainty. We successfully implement Semantic Prism on a 400-element MTS prototype with 2-bit phase control operating at 3.5-GHz. Field tests show that Semantic Prism substantially improves multiuser semantic reconstruction over the considered baselines and enables a single-RF-chain AP to concurrently serve up to seven users.

    https://arxiv.org/abs/2609.27361


    A Harmonic-Regression ARMA model for rotating machinery dynamics identification and fault detection under varying conditions

    oai:arXiv.org:2609.27436v1

    arXiv:2609.27436v1 Announce Type: new Abstract: Fault detection in rotating machinery under varying operating conditions remains a challenging problem. Existing approaches are mainly based either on non-parametric vibration signal representations, parametric deep learning models, or on parametric ARMA-type models. While deep learning approaches often lack transparency and require large training datasets, ARMA-type models rely on the assumption of a purely rational spectrum, which can be restrictive for rotating machinery vibration signals exhibiting mixed spectral characteristics. To address these limitations, a novel Harmonic Regression ARMA (HR-ARMA) model is presented for rotating machinery dynamics identification, which combines an explicit harmonic regression part for representing the dominant deterministic cyclical components with an ARMA part for the remaining broadband stochastic dynamics. Based on this model and within a Multiple Model framework, fault detection is achieved under varying operating conditions. The HR-ARMA modelling performance is validated through simulation and experimental studies on a single-stage gearbox, demonstrating accurate representation of both cyclical and broadband dynamics, improved parsimony and significantly higher fault detection performance compared with conventional ARMA models, achieving at least 98% True Positive Rate at 5% False Positive Rate in the experimental study.

    https://arxiv.org/abs/2609.27436


    Hidden Markov Model-Based Remaining Useful Life Estimation of Rolling Bearings Using Vibration Signals: A feasibility study

    oai:arXiv.org:2609.27481v1

    arXiv:2609.27481v1 Announce Type: new Abstract: This feasibility study presents a Hidden Markov Model (HMM)-based methodology for the Remaining Useful Life (RUL) estimation of rolling element bearings when the training vibration signals differ significantly from those of the target bearing. An AutoRegressive (AR) model is first employed to capture the machinery dynamics using vibration acceleration measurements from an initial operating period where all components including the considered bearing are still under healthy condition. The AR model is subsequently used to filter a newly acquired vibration signal from the faulty machinery, and typical Envelope Analysis is performed on the residual signal to identify fault-related repetition frequencies. The amplitudes of these frequencies, combined with statistical features of the AR residuals are fused through Principal Component Analysis to construct a sensitive to bearing degradation Condition Indicator (CI). Based on training data from the healthy state, a threshold is established, and a fault is declared when the CI exceeds this threshold. Once a fault is detected, its progression is modelled as a sequence of consecutive, distinct states, which are identified using K-Means clustering. A left-right HMM with continuous observation densities is estimated to represent the fault progression and thus predict the RUL. The HMM-based RUL estimation methodology is trained using vibration measurements from a limited number of two run-to-failure experiments with two nominally identical bearings, while RUL estimation performance is assessed on a third bearing of the same type. Despite the significant differences between the vibration data used for training and the life cycles of the training bearings with respect to the target bearing, the RUL estimation results indicate an adequate but conservative performance of the postulated HMM-based methodology.

    https://arxiv.org/abs/2609.27481


    Active Learning for Low-Altitude Radio Map Construction via Plug-and-Play Flow Matching

    oai:arXiv.org:2609.27486v1

    arXiv:2609.27486v1 Announce Type: new Abstract: The deployment of unmanned aerial vehicles (UAVs) in low-altitude airspace requires accurate and timely radio maps for reliable communication and safe navigation. However, constructing such radio maps is challenging due to the prohibitive overhead of exhaustive measurements and the limited flight endurance of UAVs. To address this challenge, we propose an active learning framework based on flow matching for efficient low-altitude radio map construction from sparse measurements. We first analyze a plug-and-play (PnP) inference scheme with a flow-matching prior. By characterizing the late-stage refinement behavior through an ordinary differential equation (ODE), we theoretically show how the inference steps smoothly align with a continuous ODE flow to refine the map details. Recognizing that the early generative stages are largely noise-dominated, this insight motivates our proposed truncated flow matching plug-and-play (TFM-PnP) approach. TFM-PnP utilizes a spatial interpolation-based initialization to start the reconstruction from an intermediate flow time, thereby bypassing the inefficient early stages. We further use the generative diversity of flow matching to derive an uncertainty map to guide the UAV trajectory design. Specifically, we propose a weighted sampling approach to select a target location, followed by a Utility-Aware Path Search (UAPS) algorithm to design the corresponding UAV trajectories. Simulation results based on Sionna ray-tracing datasets show that the proposed framework outperforms the considered baselines, achieving more than 50% reduction in normalized mean squared error (NMSE).

    https://arxiv.org/abs/2609.27486


    Codebook-Based Effective CSI Feedback for Precoding Design in MIMO Systems

    oai:arXiv.org:2609.27505v1

    arXiv:2609.27505v1 Announce Type: new Abstract: Downlink multi-user multiple-input multiple-output (MIMO) precoding with limited channel state information (CSI) feedback in multi-cell systems is studied. Conventional codebook-based CSI feedback compresses the UE-specific physical channels, which leaves an inherent mismatch between the base-station (BS) precoders and the interference-aware UE combiners. To address this limitation, an effective CSI (ECSI) feedback framework is proposed in which UEs compute their linear combiners from precoded downlink pilots, construct corresponding post-combining effective channels, and report compact quantized representations using the same payload structure as in conventional schemes. The BS reconstructs the effective CSI and iteratively refines its precoders through over-the-air signaling without incurring additional downlink overhead. Analytical expressions are derived to characterize the effective-channel estimation error under both conventional CSI and ECSI feedback, explicitly capturing the impact of beam selection and rank truncation. Simulations demonstrate fast convergence within a few iterations and up to 30\% average sum-rate improvement over conventional CSI-based precoding, with gains reaching 100\% at the 10th percentile of per-stream rates in interference-limited conditions.

    https://arxiv.org/abs/2609.27505


    Conformalized Kalman Filters for State Estimation with Trustworthy Confidence Regions

    oai:arXiv.org:2609.27506v1

    arXiv:2609.27506v1 Announce Type: new Abstract: Kalman-type filters are widely used for tracking dynamic systems, yet the confidence regions commonly derived from their estimated covariances can become unreliable under nonlinearities, non-Gaussian disturbances, and model mismatch. In this work, we develop a conformal prediction (CP) framework for equipping Kalman-type filters with statistically reliable confidence regions. Unlike CP applied to black-box estimators, our approach exploits the recursively estimated first- and second-order moments of the state posterior, which capture the time-varying uncertainty induced by the underlying dynamics. Based on these statistical features, we propose three complementary constructions. The first conformally calibrates Gaussian confidence regions induced by the filter moments. The second employs quantile regression to map the estimated moments into the boundaries of adaptive convex regions, which are subsequently calibrated. The third learns a Gaussian-mixture representation of the posterior and conformalizes the resulting density-based regions, enabling the characterization of multimodal and non-convex uncertainty sets. We establish finite- sample coverage guarantees for both sample-wise confidence, which controls miscoverage at individual time instances, and trajectory- wise confidence, which jointly covers the state sequence over a prescribed horizon. Numerical experiments across diverse linear and nonlinear dynamic systems demonstrate that the proposed methods attain the prescribed coverage while producing tight and informative confidence regions, and highlight the relative merits of the three constructions under different posterior characteristics.

    https://arxiv.org/abs/2609.27506


    The Second MLC-SLM Challenge: Multilingual Conversational Speech Diarization, Recognition, and Understanding

    oai:arXiv.org:2609.27514v1

    arXiv:2609.27514v1 Announce Type: new Abstract: This paper summarizes the Interspeech2026 second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which aims to advance the development of effective multilingual conversational speech language models. We describe the two challenge tasks: multilingual conversational speech diarization and recognition, and multilingual conversational speech understanding, together with the released real-world conversational speech dataset, evaluation protocols, and baseline systems. The challenge attracted 91 teams worldwide, with 704 valid leaderboard results and 14 technical reports across the two tasks. Based on the participating systems, we summarize representative approaches and distill practical insights into multilingual conversational speech recognition and understanding to support future research in the community.

    https://arxiv.org/abs/2609.27514


    An Operational Leakage Metric for Eavesdropping in Joint Sensing and Communications Systems

    oai:arXiv.org:2609.27527v1

    arXiv:2609.27527v1 Announce Type: new Abstract: In upcoming integrated sensing and communications (ISAC) systems, security depends on whether an eavesdropper (Eve) can recover data symbols or infer sensing information from a common observation with a legitimate receiver. However, existing information and estimation measures do not quantify this dual risk under finite observation records. This letter introduces the novel metric of operational leakage, which signifies Eve's optimal probability, in the Bayesian sense, of reconstructing communication symbols or estimating a sensing parameter within prescribed distortions. By showing that prior adversarial success probability separates the effect of side information from that of the observation, we propose a conditional Laplace approximation for efficiently evaluating the proposed metric and jointly quantifying the eavesdropping risk to the sensing and communication functionalities. Numerical results showcase that the risk enabled by adversarial observation increases with Eve's signal-to-noise ratio (SNR), conventional metrics do not uniquely determine finite-record adversarial success probability, and that the proposed conditional Gaussian approximation closely follows numerical Bayes evaluation.

    https://arxiv.org/abs/2609.27527


    Degrees of Freedom Sampling

    oai:arXiv.org:2609.27541v1

    arXiv:2609.27541v1 Announce Type: new Abstract: A geometry-based framework for sampling electromagnetic fields over surfaces is presented. The approach extends the shadow-area formulation of electromagnetic degrees of freedom (DoF) from a global mode count to a spatially resolved DoF density that determines the sampling distribution. The observation surface is partitioned into cells containing approximately one DoF, with the cell area determined by the DoF density and the cell shape determined by the projection of the DoF density vector onto the observation surface. In the far field, the source shadow area defines a directional DoF density over the observation sphere, while in the near field, the mutual-shadow density and a DoF density direction determine the local sampling geometry. The resulting method requires only the source-observation geometry and does not require an explicit singular-value decomposition of the propagation operator. Numerical results for several source and observation geometries demonstrate reconstruction performance close to that of operator-based sampling methods. A beamforming interpretation further shows that the sampling cells correspond to geometry-dependent beam regions, establishing a unified connection between DoF, sampling, and beamforming.

    https://arxiv.org/abs/2609.27541


    Robust High-Dimensional MVDR Beamforming under Heavy-Tailed Noise via Spiked Covariance Modeling

    oai:arXiv.org:2609.27552v1

    arXiv:2609.27552v1 Announce Type: new Abstract: This paper proposes a robust high-dimensional minimum variance distortionless response (MVDR) beamforming method for array observations corrupted by heavy-tailed noise. The proposed approach constructs an MVDR-oriented precision matrix estimator by combining Maronna's robust scatter estimator with spiked covariance modeling. Using tools from random matrix theory, we derive a deterministic equivalent of the MVDR output power and obtain asymptotically optimal shrinkage weights for the dominant signal subspace. A fully sample-based implementation is then developed for practical beamformer design. Numerical simulations under elliptically distributed noise demonstrate that the proposed beamformer achieves stronger interference suppression and higher output SINR than competing methods in high-dimensional and implusive noise regimes.

    https://arxiv.org/abs/2609.27552


    Optimal Trajectory Generation for Improved Magnetic Navigation

    oai:arXiv.org:2609.27553v1

    arXiv:2609.27553v1 Announce Type: new Abstract: Magnetic navigation has emerged as a promising alternative for navigation in Global Positioning System (GPS)-denied environments, leveraging geomagnetic field maps in conjunction with onboard magnetometer measurements. However, its performance is highly sensitive to trajectory-dependent observability, which limits its practical effectiveness under conventional flight paths. This paper proposes an optimal trajectory design framework for magnetic navigation that maximizes information content along the flight path. The trajectory generation problem is formulated as an optimal control problem that minimizes the posterior Cram\'{e}r--Rao lower bound on the position estimation error, subject to a penalty on path length. The resulting trajectories are non-intuitive and significantly enhance the observability of the navigation system. Simulation results demonstrate that the proposed optimal trajectories yield substantial reductions in estimation error compared to conventional straight-line trajectories, highlighting the critical role of trajectory design in enabling high-accuracy magnetic navigation. These findings suggest that trajectory optimization can substantially improve the viability of magnetic navigation as a robust alternative for aerospace applications in GPS-denied environments.

    https://arxiv.org/abs/2609.27553


    Action-Directed Information for Distributed Control and Agentic Interaction

    oai:arXiv.org:2609.27580v1

    arXiv:2609.27580v1 Announce Type: new Abstract: Distributed intelligence concerns systems in which semi-autonomous components with local dynamics and partial observations coordinate through information exchange to maintain a shared function. This paper proposes an operational way to study such systems: measure information at the interface where a message changes a receiving action, then connect that measure to function by intervention and disturbance evaluation. We instantiate this proposal in DI-Walker, a two-dimensional four-limb embodied plant controlled by frozen Cross-Entropy-Method policies. We compare a controller using each limb's own realized-force sensor with one using the realized-force sensors of peer limbs. Under limb loss, limb slip, and weak central-control dropout, Peer-Sensor has lower late tracking error in several conditions. A corrected finite-history action-predictive estimator shows a substantially larger peer-message gain under compound failure. A future scalar functional-prediction estimator does not show the same stable advantage. We interpret this discrepancy as a methodological result: information useful for an intermediate control action can be hidden by later plant dynamics, redundancy, and context. The paper relates this result to Predictive Information, Transfer Entropy, Directed Information, information-to-go/IT-PAC ideas, empowerment, and the robust control data-rate perspective, while explicitly distinguishing operational predictive gains from exact Directed Information, channel capacity, and a formal data-rate theorem.

    https://arxiv.org/abs/2609.27580


    DriftAudio: Marginal Drifting for Distributional Post-Training of One-Step Text-to-Audio Generators

    oai:arXiv.org:2609.27598v1

    arXiv:2609.27598v1 Announce Type: new Abstract: Recent one-step text-to-audio (TTA) models substantially reduce inference cost, yet their generated distributions can still be improved through post-training. We propose DriftAudio, a distributional post-training method that adapts Drifting to pretrained one-step TTA generators. Applying Drifting condition-wise is challenging under free-form text conditioning, where only one or a few real samples are typically available for a particular condition. DriftAudio instead performs Drifting on the marginal audio distribution while retaining text-conditioned generation. The drifting field is estimated in a frozen audio feature space using resampled real samples, a rolling bank of generated samples, and the current batch of generated samples. The resulting Drifting field provides a detached training target for updating only the generator, keeping the original one-step inference procedure. On AudioCaps, starting from MeanAudio, DriftAudio reduces FAD and FD by 33.9% and 17.6%, respectively, while also improving KL and CLAP. Starting from FdAudio, it further reduces FAD, FD, and KL, with trade-offs in IS and CLAP. These results demonstrate the effectiveness of marginal distributional post-training for one-step TTA generation.

    https://arxiv.org/abs/2609.27598


    EmphTTS: an emphasis-control TTS with reinforcement learning

    oai:arXiv.org:2609.27599v1

    arXiv:2609.27599v1 Announce Type: new Abstract: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.

    https://arxiv.org/abs/2609.27599


    Forced Oscillations in Power Systems Induced by Data Centers Hosting AI Workloads

    oai:arXiv.org:2609.27698v1

    arXiv:2609.27698v1 Announce Type: new Abstract: Power swings in large Data Centers (DTCs) running Artificial Intelligence (AI) workloads can excite poorly damped modes in power systems. The resulting forced oscillations can lead to flicker, equipment disconnection, or blackouts. This paper investigates the risks of such load fluctuations for different system strengths and damping conditions. Using an analytical approach based on transfer functions, we identify critical DTC locations in the power grid at which load fluctuations could induce the largest forced oscillations and further characterize the harmonic spectrum of the resulting system response. Depending on the frequency and magnitude of the DTC load fluctuations, forced oscillations can become unbounded. The underlying instabilities are classified into saddle-node bifurcations of the forced periodic response and impasse-surface encounters, using Floquet multipliers and the minimum singular value of the algebraic Jacobian. Furthermore, the impact of different duty cycles and harmonic components beyond the fundamental oscillation frequency in square-wave load profiles is analyzed. Finally, the interaction of two oscillating DTCs is investigated for different locations and forcing frequencies, considering both synchronized and unsynchronized operation. The findings can help system operators to define new regulations on the maximum load fluctuations permitted for DTC facilities at specific grid locations, without negatively affecting the stability and operation of the system.

    https://arxiv.org/abs/2609.27698


    GA-Agent: Large Language Models as Hyperparameter Optimizers for Evolutionary Controller Synthesis

    oai:arXiv.org:2609.27725v1

    arXiv:2609.27725v1 Announce Type: new Abstract: Tuning PID controllers to satisfy competing objectives - low tracking error, fast settling, limited overshoot, and moderate control effort - is labor-intensive and requires expertise. Genetic algorithms (GAs) offer gradient-free optimization of controller gains against a weighted fitness function, but success depends on meta-level choices: population size, generation budget, gain bounds, and fitness weights. These are usually set by manual trial-and-error or costly bilevel optimization, exposing a tension: GAs excel at dense numerical search, but configuring them needs high-level, context-dependent semantic reasoning. We propose GA-Agent, which decouples these modes. A standard GA handles low-level PID gain optimization. A large language model (LLM) agent operates at the meta-level: it observes completed GA runs, diagnoses gaps versus user control objectives, and proposes updated GA configurations. The architecture uses structured memory, quantitative goal translation, resource-aware termination, and outcome-driven routing. We evaluate GA-Agent on eight control case studies with diverse dynamics (DC motor, inverted pendulum, aircraft pitch, autonomous underwater vehicle, and others). GA-Agent achieves 100% success on all benchmarks, outperforming a Regular GA with fixed hyperparameters in solution quality and sample efficiency. It matches or surpasses a Cascade-GA baseline while reducing function evaluations by one to two orders of magnitude, typically converging in one to three optimization attempts. Sensitivity analysis shows robustness across LLM backbones and memory configurations. A compact memory buffer (size 2-3) and cost-effective models (DeepSeek-V4-Flash at about $0.002 per run) achieve superior performance.

    https://arxiv.org/abs/2609.27725


    Analytical Framework of Radial Resolution for Near-Field Communications

    oai:arXiv.org:2609.27733v1

    arXiv:2609.27733v1 Announce Type: new Abstract: As extremely large antenna arrays (ELAAs) become central to next-generation wireless systems, the transition into the near-field propagation regime enables the exploitation of spherical wavefronts for radial-domain beamfocusing. This capability is pivotal for emerging applications requiring precise spatial isolation, such as Space Division Multiple Access (SDMA), hierarchical localization and advanced sensing. However, fully realizing these technologies requires specific design rules to dimension multi-user systems without relying on computational expensive full-wave simulations. To address the gap in modeling contiguous focal regions with controllable radial resolution, this paper expands the Angular Spectrum Representation (ASR) approach to propose a comprehensive analytical framework. Through the introduction of a tunable inter-beam overlap parameter $\rho$, we derive closed-form expressions to synthesize multiple focal regions, providing the flexibility to tailor their radial resolution. Furthermore, the resolution capabilities are analyzed to characterize the interplay between key operational variables, such as the transmitter size, beam radius and operation frequency. System-level assessment of per-user and sum-rate spectral efficiencies across varying signal-to-noise (SNR) regimes reveals how the inter-beam overlap dictates a fundamental trade-off between user capacity and inter-user interference, delivering design guidelines for future near-field communications.

    https://arxiv.org/abs/2609.27733


    Holonic Graceful Transitions Across Centralized, Decentralized, Distributed, and Local Control in DER-rich Cyber-Power Distribution System

    oai:arXiv.org:2609.27759v1

    arXiv:2609.27759v1 Announce Type: new Abstract: The increasing penetration of distributed energy resources (DERs) in distribution system necessitates adaptive coordination frameworks. These frameworks must remain optimal during normal operation and resilient under cyber physical disturbances. Existing coordinated control approaches are typically deployed as static architectures with limited ability to adapt when communication degrades, local instability emerges, and operating conditions become spatially heterogeneous. This work addresses the gap by proposing a DER service-independent edge autonomous holonic adaptive coordination framework. In this framework, each DER controller executes its own control actions and transitions among centralized, distributed, decentralized, and local autonomous coordination modes without always relying on static coordination commands from the grid operator. The framework preserves coordination continuity across modes by retaining local controller states while reconfiguring only the coordination topology, information exchange pattern, and fallback action associated with the active DER service. Graceful transitions are enabled through dwell timers, rate limiting, and safety overrides to prevent dynamic instability during mode changes. Volt VAR control is used as a representative distribution automation application to validate the proposed architecture in a cyber-physical Hardware-in-the-loop (HIL) testbed. The proposed approach is evaluated under diverse cyber-physical event-based scenarios, showing region-confined adaptation under localized disturbance, reduced coordination traffic during coordination switching, and node-confined mitigation under cyber attack through edge anomaly detection and neighbor-corroborated impact estimation.

    https://arxiv.org/abs/2609.27759


    Formation Keeping Control for Deorbiting an Uncooperative Satellite by Laser Ablation

    oai:arXiv.org:2609.27838v1

    arXiv:2609.27838v1 Announce Type: new Abstract: This paper proposes the formation keeping control law for deorbiting debris by a laser ablation. Laser ablation is vital technology for contactless active debris removal, where a chaser satellite with a laser system irradiates laser pulses to a target object to generate the ablation force for deorbiting. The deorbiting force decelerates the target, and the chaser must maintain its relative position and continue irradiating. In other words, both the chaser and the target are supposed to be deorbited simultaneously, where both have accelerations. Although conventional formation flying missions assume that only a chaser maneuvers, the formation flying in this paper considers that both a chaser and a target have accelerations. Thus, this paper derives the relative equations of motion between the chaser and the target in powered flight and their analytical solution using relative orbital elements. A control law based on the analytical solution is proposed, which determines the timings and directions of the laser ablation and the electrical thrust so that the formation periodically returns to a desired formation. Numerical simulations first examine the control law in two cases with different maneuver timings. Then, a Monte Carlo simulation is performed to verify the effectiveness of the control law for a variety of desired formations.

    https://arxiv.org/abs/2609.27838


    Optimization of Fault-Tolerant Thruster Configurations for Satellite Control

    oai:arXiv.org:2609.27839v1

    arXiv:2609.27839v1 Announce Type: new Abstract: The fault tolerance of spacecraft actuators significantly affects the reliability of satellites and the likelihood of successful missions. To enhance the fault tolerance of the actuators, this study derives optimal fault-tolerant configurations of fixed thrusters that maximize the controllability of a fully-actuated or underactuated satellite. The proposed method optimizes thrust and torque directions generated by the thrusters. Thus a cost function in terms of the thruster locations and directions is defined as the summation of the generated control forces and torques with respect to the body-fixed frame. The optimal configuration is obtained by the successive use of an energy potential method that is motivated by Thomson's problem. Some numerical examples are provided that show the effectiveness of the proposed formulation and optimization method.

    https://arxiv.org/abs/2609.27839


    Suboptimal Formation Reconfiguration of Satellites Under Input Directional Constraints

    oai:arXiv.org:2609.27841v1

    arXiv:2609.27841v1 Announce Type: new Abstract: Proximity operations of satellites such as formation flying and on-orbit servicing offer more advanced missions than missions achieved by a single satellite. In a practical situation of formation flying, thrust directions for keeping and controlling a relative orbit is limited, e.g., for astronomical observation and plume impingement avoidance. The aim of this paper is to provide an energy efficient control method for a formation reconfiguration under input directional constraints with respect to both an inertial and a leader-fixed frames. The proposed controller is designed consisting of two parts: 1) guaranteeing a formation reconfiguration to a desirable formation on the basis of an energy optimal controller and 2) satisfying the input directional constraints by superimposing additional inputs. The analytical form of the control input shows that the input direction forms an ellipse in the leader-fixed frame when a particular boundary condition is satisfied, which is exploited as a nominal controller to take into account the input directional constraints. Due to the singular avoidance of the nominal controller, the additional inputs can be analytically obtained. The effect on the follower trajectory due to the additional inputs is compensated by setting a virtual target orbit, and thus the successful formation reconfiguration is still guaranteed. Some numerical simulation results verify the effectiveness of the proposed method and compare the energy efficiency.

    https://arxiv.org/abs/2609.27841


    Semantic-Guided Fusion Network for Multi-Source Remote Sensing Image Classification

    oai:arXiv.org:2609.27854v1

    arXiv:2609.27854v1 Announce Type: new Abstract: Multi-source remote sensing image classification has attracted increasing attention due to the complementary spectral, structural, and geometric information. However, existing methods still suffer from two limitations: insufficient semantic contextual modeling and unreliable feature fusion caused by slight spatial misalignment. To address these issues, we propose a Semantic-Guided Fusion Network (SGFNet) for multi-source remote sensing image classification. Specifically, the Semantic Mixing Convolution Block (SMCB) is designed to dynamically generate semantic-aware convolution kernels according to contextual relationships among feature representations. In addition, the Frequency Modulated Fusion Block (FMFB) is introduced to perform cross-modal interaction in the frequency domain, which effectively alleviates the influence of slight spatial misalignment and improves complementary information fusion. Extensive experiments conducted on the Augsburg and Houston 2018 datasets demonstrate that the proposed SGFNet consistently outperforms several state-of-the-art methods. The codes are publicly available at https://github.com/oucailab/SGFNet .

    https://arxiv.org/abs/2609.27854


    Image Denoising Using Lower Semi-Frames

    oai:arXiv.org:2609.27893v1

    arXiv:2609.27893v1 Announce Type: new Abstract: A blind image denoising framework based on an infinite directional lower semi-frame (DLSF) is proposed for additive white Gaussian noise. The model employs scale-dependent directional analysis with resolvent regularization of the unbounded semi-frame operator. Noise variance is estimated directly in the DLSF domain by modeling the joint covariance of four directional difference channels and applying covariance whitening to obtain a chi-square statistic. A lower-tail moment estimator provides blind noise estimation without median absolute deviation. The estimated noise level is incorporated into channel-wise Wiener-type shrinkage and canonical-dual synthesis, followed by a data-consistent iterative reconstruction with automatic stopping. Experiments on three standard grayscale images at noise levels 15--30 yield a mean relative noise-estimation error of 3.28\%, with average improvements of 7.45 dB in PSNR and 0.367 in SSIM. At 30/255 noise, the estimation error decreases to 1.73\%, with a mean PSNR gain of 8.31 dB. Results demonstrate effective noise suppression and structural preservation, with the strongest performance on smooth and edge-dominated images.

    https://arxiv.org/abs/2609.27893


    Electromagnetic-Twin Channel Estimation: From Sparse Pilots to Persistent CSI

    oai:arXiv.org:2609.27894v1

    arXiv:2609.27894v1 Announce Type: new Abstract: Each repeated pilot block is typically treated as a separate channel-estimation problem, although much of the angle--delay state persists across updates. This paper defines an electromagnetic twin (ET) as a recursively synchronized signal-processing state containing channel state information (CSI), coefficientwise uncertainty, support confidence, and age. A channel knowledge map may initialize this state; subsequent pilots predict, correct, validate, and store the state. We formulate an anchored-window inverse problem that yields online uncertainty-weighted correction and fixed-lag smoothing as two convex operating modes of a persistent-state estimator. Uncertainty-aware ET FISTA (UET-FISTA) gives a zero-delay update, while an alternating direction method of multipliers (ET-ADMM) jointly refines a short window when delayed CSI is acceptable. Held-out pilots choose update, complete refresh, or hold under a prescribed family-wise false-acceptance probability, preventing an inaccurate stored state from being reused blindly. Conditional Bayesian Cram\'er--Rao bounds quantify the information supplied by the stored state and by future pilot blocks. With 20\% pilots, UET-FISTA reaches $-11.55$ dB normalized mean-square error, compared with $-3.42$ dB for static FISTA. Over 200 independent 50-epoch trajectories, three-epoch smoothing improves matched-state NMSE from $-10.28$ to $-13.41$ dB. On 40 off-grid QuaDRiGa trajectories, online UET-FISTA improves static recovery from $-1.37$ to $-1.88$ dB and effective rate from 1.688 to 1.919 bit/s/Hz. These results identify when persistent CSI reduces pilot demand and when the receiver should instead refresh or preserve its state.

    https://arxiv.org/abs/2609.27894


    Electromagnetic Twin: From Sparse Measurements to Persistent Wireless Intelligence

    oai:arXiv.org:2609.27895v1

    arXiv:2609.27895v1 Announce Type: new Abstract: Radio maps and channel knowledge maps provide reusable propagation knowledge, but a stored map can become locally outdated after persistent changes to doors, partitions, furniture, large equipment, or infrastructure. Motivated by digital-twin state synchronization, we develop an electromagnetic twin that recurrently converts sparse channel measurements into a persistent wireless state, exposes that state to communication queries, and uses its uncertainty to request subsequent measurements. The twin tracks persistent or semi-persistent propagation changes rather than transient human motion or fast fading. Each update is a constrained inverse problem that combines the previous map, a scene graph, and an imperfect physics prior. Measurement consistency with a confidence-calibrated radius and physical gain bounds enforce feasibility, while scene-aware spatial regularization and selective temporal memory preserve propagation boundaries and unchanged regions. Successive convex approximation (SCA), majorization--minimization alternating direction method of multipliers (MM-ADMM), and a low-complexity primal--dual hybrid gradient (LC-PDHG) mode realize the framework at different computational scales, and a local perturbation analysis gives an explicit multi-update tracking recursion. The updated state supports access-point association and codebook beam selection, while change-weighted A-optimal design closes the measurement--update--query loop. Over 500 realizations, the recurrent update remains stable through eight persistent scene events and raises best-beam accuracy from $76.64\%$ to $90.83\%$ as measured-location density grows from $2\%$ to $16\%$.

    https://arxiv.org/abs/2609.27895


    Resilient Monitoring of Social Dynamical Systems through Collaborative Multi-Agent Networks under Latency

    oai:arXiv.org:2609.27902v1

    arXiv:2609.27902v1 Announce Type: new Abstract: Social dynamical networks significantly influence contemporary digital landscapes, affecting realms from social activism to public policy formulation. This paper investigates the use of multi-agent systems (MAS) to monitor and analyze these networks. Firstly, we propose a single-time-scale distributed inference model designed to effectively manage challenges such as latency and agent failure. Secondly, we provide sufficient conditions that ensure the stability of the proposed scheme. Notably, the observer gain design remains effective regardless of time delays. Thirdly, we develop a computationally efficient recovery mechanism for agent failures that relies on employing graph-theoretic approaches to restore network observability by replacing failed agents (due to sensing failure or unbounded delays, i.e., packet drops) by implementing computationally efficient graph-theoretic methods to assign observationally equivalent agent counterparts. Lastly, we illustrate the proposed scheme through a pedagogical example and real-world network applications.

    https://arxiv.org/abs/2609.27902


    Local SVD-Entropy Maps as a Complementary Structural Representation for Full-Reference and No-Reference Image Quality Assessment

    oai:arXiv.org:2609.27959v1

    arXiv:2609.27959v1 Announce Type: new Abstract: We investigate a local spectral-complexity representation for perceptual image quality assessment (IQA) based on Shannon entropy of singular values computed directly from two-dimensional image patches. For each $3\times3$-pixel grayscale patch, SVD is applied directly and the normalized singular-value entropy defines one HSVD-map value. The construction requires neither flattening nor delay embedding, uses no boundary padding, and is invariant to $90^{\circ}$ rotations and mirror reflections at the local-descriptor level. A nested salt-and-pepper experiment on Lena separates absolute similarity to a clean reference from sensitivity to an additional degradation step. HSVD-SSIM responds more strongly to local corruption and retains a larger neighboring-state response at severe noise levels. Validation on all 10,125 distorted KADID-10k images shows that HSVD-SSIM is weaker than conventional SSIM as a standalone full-reference metric (SRCC $0.450$ vs. $0.619$), but complementary when combined with it: grouped cross-validation increases SRCC from $0.618$ to $0.659$, with a bootstrap 95\% confidence interval of $[0.036,0.046]$ for the gain. In a no-reference experiment, adding HSVD-derived single-image descriptors improves the best nonlinear model from SRCC $0.528$ to $0.575$ (95\% CI $[0.033,0.061]$) and also improves prediction of quality changes between neighboring distortion states. These results support direct local SVD entropy as an interpretable structural channel that complements conventional image-domain similarity and remains informative without a pristine reference.

    https://arxiv.org/abs/2609.27959


    When Gigawatts of Computational Load Disappear: Cycle-Space Certificates for Grid Synchronization and Transient Stability

    oai:arXiv.org:2609.27989v1

    arXiv:2609.27989v1 Announce Type: new Abstract: Rapid growth of data centers and artificial-intelligence services is producing computational loads at scales once associated mainly with largest power plants. Recent grid events show that a routine transmission disturbance can cause several gigawatts of data-center demand to disconnect or transfer to backup nearly at once. This article revisits the classical synchronization and transient-stability theory needed to reason about such events. We organize four lines of work---graph-based synchronization conditions, winding-number descriptions of nonlinear power flow, separable convex network optimization, and direct energy methods---into a single cycle-space certificate framework for the lossless fixed-voltage model. The static layer gives an exact strict-cohesion test within a prescribed winding cell and reveals the widely used D\"orfler--Chertkov--Bullo test as a quadratic surrogate of the same convex problem. The dynamic layer converts the critical-energy calculation into a finite family of convex boundary problems. Standard MATPOWER benchmarks illustrate both what the stronger static test gains and where it gains nothing: the 118-bus case admits $16.2\%$ more loading than the sufficient screen, while the 39-bus case is bridge-limited and the thresholds coincide. A stylized $2.7$-GW 39-bus event further shows that transient margin can change by about a factor of two depending on where balancing power is supplied, even when every final balanced operating point remains statically feasible. The result is a tutorial synthesis and an extensible deterministic certificate for emerging gigawatt-scale computational-load contingencies.

    https://arxiv.org/abs/2609.27989


    Subjective Evaluation of DNN AND Auditory-Model-Based Hearing-Loss Compensation

    oai:arXiv.org:2609.28033v1

    arXiv:2609.28033v1 Announce Type: new Abstract: Outer-hair-cell (OHC) loss is a primary deficit of sensorineural hearing loss (SNHL), impairing cochlear amplification and frequency selectivity and thereby elevating hearing thresholds. Biophysically-inspired DNN-based hearing-aid (HA) algorithms have been proposed to compensate for OHC deficits and have shown clear benefits in objective speech intelligibility and quality metrics (e.g. HASPI, HASQI). However, comprehensive subjective validation of these benefits in human listeners is still missing. In this work, we present a subjective evaluation of a biophysically-inspired HA model targeting OHC deficits. The cochlear module of an auditory model was individualized based on each listener's pure-tone audiogram and integrated into a trainable system, which includes the personalized model and a normal-hearing reference model, and the resulting trained HA was evaluated using a Matrix test comparing intelligibility scores for unprocessed and HA-processed noisy speech. The results revealed a significant benefit of the HA model over the unprocessed condition in the range of +1 to +27%, providing behavioral confirmation of the efficacy of the HA model. This study closes the gap between objective and perceptual evidence for this new generation of DNN-based HA algorithms, paving the way for their integration into next-generation DNN-accelerated chips for hearables and hearing aids.

    https://arxiv.org/abs/2609.28033


    Towards Energy-Neutral IoT Sensors: Low-Cost Chirp-Based Backscattering Node Using a COTS Microcontroller and Multi-Load Front-End

    oai:arXiv.org:2609.28052v1

    arXiv:2609.28052v1 Announce Type: new Abstract: Backscatter communication has proven to be a key enabler for energy-neutral sensor nodes in the IoT. It allows ultra-low-power nodes to transmit sensor data by reflecting incident RF signals. Practical large-scale deployment requires hardware solutions that are both cost-efficient and widely available. This paper investigates the practical limits of implementing backscatter transmitters exclusively using COTS components. We compare a two-load and multi-load architecture in terms of hardware requirements, complexity and spectral efficiency highlighting the most important trade-offs. A multi-load backscatter architecture is developed using a resource constrained low-power microcontroller, levering DMA to achieve high switching speeds while minimizing processing overhead. Furthermore, we present a layout agnostic load design methodology based on scattering parameters, allowing precise control of the reflection coefficients. A prototype implementation demonstrates a successful backscatter transmission of a single-sideband chirp with reduced harmonic distortion. Measurement results show suppression of the backscattered carrier and the second sideband by 10 dBm. This demonstrates that low-cost sensor nodes can be built using COTS components enabling remote monitoring at extended lifetimes.

    https://arxiv.org/abs/2609.28052


    Event-driven signal reconstruction through neuromorphic compressive sensing

    oai:arXiv.org:2609.28063v1

    arXiv:2609.28063v1 Announce Type: new Abstract: Compressive sensing (CS) exploits intrinsic signal sparsity for efficient representation of high-dimensional data. However, conventional CS relies on fixed-length measurement representations and dense reconstruction operations, leaving communication and computational costs largely determined by system dimensions rather than signal sparsity. Here, we propose a neuromorphic CS framework centred on a spike-driven learned iterative shrinkage-thresholding algorithm (S-LISTA). By representing compressed measurements and reconstruction updates as sparse events, the framework links intrinsic signal sparsity to the spatiotemporal sparsity of spiking neural networks (SNNs), extending the benefits of sparsity across sensing, transmission and reconstruction. Communication and computational costs consequently depend on event activity, allowing sparse representations to translate into resource savings. Extensive experiments demonstrate substantial reductions in transmitted data volume and estimated reconstruction energy while maintaining competitive reconstruction and downstream task performance. The framework also exhibits robustness under challenging channel conditions. These findings establish neuromorphic CS as a promising paradigm for signal reconstruction in resource-constrained scenarios.

    https://arxiv.org/abs/2609.28063


    Echo Detection in Spatial Room Impulse Responses Measured with Spherical Microphone Arrays Using the Herglotz Wavefunction

    oai:arXiv.org:2609.28068v1

    arXiv:2609.28068v1 Announce Type: new Abstract: Early reflections in spatial room impulse responses (SRIRs) measured using spherical microphone arrays (SMAs) play an important role in spatial audio analysis, rendering, and reverberation modeling. Their accurate localization and characterization can facilitate the analysis, processing, and manipulation of measured reverberation fields. This paper proposes a method for detecting and characterizing early reflections based on the Herglotz wavefunction formalism. The measured sound field is represented as a continuous superposition of incident plane waves, from which a localization function is derived in the spherical harmonic domain. An adaptive radial Gaussian fitting procedure is then used to estimate both the directions of arrival (DoAs) and the number of incident reflections within a single analysis frame. The proposed framework is evaluated using simulated and measured SRIRs and compared with conventional steered response power (SRP) and multiple signal classification (MUSIC) localization methods. The results demonstrate improved localization accuracy and more reliable detection of multiple simultaneous reflections, highlighting the potential of the proposed Herglotz-based formulation for the analysis of early reflections in SMA-measured SRIRs.

    https://arxiv.org/abs/2609.28068


    Recursive Uncertainty-Gated Image Registration for Learning-based Algorithms

    oai:arXiv.org:2609.28081v1

    arXiv:2609.28081v1 Announce Type: new Abstract: Conventional image registration algorithms are robust to domain shifts and achieve low errors, but they are slow and computationally expensive. Deep-learning methods are efficient at inference-time, but face challenges in out-of-domain samples. We propose Recursive Uncertainty-Gated Image Registration (RUGI), an algorithm for iteratively refining deformation fields predicted by learning-based registration models. At each iteration, the registration model predicts an incremental deformation, and a gating map modulates the update. Refinements are hence concentrated in regions that remain difficult to register. We explore two gating strategies: a learned uncertainty-based approach and an image residual error approach. We evaluate RUGI on cardiac MRI and echocardiography datasets and show consistent improvements over single-step inference. Ablation experiments demonstrate that iterative refinement alone improves registration, but informative spatial gating provides a significant additional benefit. The error-gated variant of RUGI can also be applied directly to existing pretrained models; applied to VoxelMorph, TransMorph, and CycleMorph, it yields MSE reductions of 27-37% with no modification to the original training procedure. The improvements in registration performance are reflected in decreased errors in ejection fraction estimation relative to ground truths. These results demonstrate that spatially selective iterative refinement provides an effective strategy to improve registration accuracy at inference-time.

    https://arxiv.org/abs/2609.28081


    Exact Average Consensus under Noisy Communication Links: A Decentralized Gradient Perspective

    oai:arXiv.org:2609.28082v1

    arXiv:2609.28082v1 Announce Type: new Abstract: We study the distributed average consensus problem under persistent link-level disturbances modeled as a martingale difference sequence with uniformly bounded conditional second moments. Under such disturbances, the standard stochastic-approximation-based linear iteration with diminishing stepsizes drives the network to consensus on an unbiased random variable with non-vanishing variance instead of the exact initial average. To understand and resolve this limitation, we develop an anchoring-based mechanism derived from a decentralized gradient descent formulation and study the effect of incorporating a decaying anchoring term that continuously pulls each agent state toward its initial value. This perspective provides an intuitive interpretation of how state anchoring counteracts disturbance accumulation. Under standard summability conditions, we prove that the resulting algorithm achieves exact average consensus almost surely. Furthermore, this decentralized gradient perspective offers a unifying framework for several related methods and an interpretable design principle for exact average consensus under persistent disturbances.

    https://arxiv.org/abs/2609.28082


    EvEMTBench: An Open Benchmark for Machine Learning in Power System Protection

    oai:arXiv.org:2609.28149v1

    arXiv:2609.28149v1 Announce Type: new Abstract: Studies of machine-learning-based power system protection are difficult to compare because task definitions, measurement access, data partitions, metrics, and generalization conditions often differ. EvEMTBench addresses this gap with an open, executable, and versioned benchmark that fixes these evaluation choices while leaving model design open. Across four grids spanning 20-345 kV, it defines 12 protection and event-analysis functions instantiated as 24 scored tasks and supports structured evaluation across observability conditions, predefined distribution shifts, and zero-shot and fine-tuned cross-grid transfer. Committed partitions, leakage controls, and reproducible reporting provide a common basis for comparing future methods. A reference evaluation spanning trivial, conventional, feature-based, and deep-learning baselines shows that wider observability is not uniformly beneficial, shifted conditions can reveal failures not apparent in-distribution, and cross-grid transfer is substantially stronger for fault detection than for fault localization. Protection-relevant diagnostics identify failure modes not apparent from primary metrics alone. EvEMTBench therefore makes generalization in machine-learning-based protection an explicit and reproducible evaluation problem.

    https://arxiv.org/abs/2609.28149


    Low-Rank Prior-Guided Rank-One Sensing for Efficient CSI Feedback

    oai:arXiv.org:2609.28156v1

    arXiv:2609.28156v1 Announce Type: new Abstract: Downlink channel state information (CSI) feedback is essential for frequency-division duplex massive MIMO, yet the feedback overhead grows rapidly with the numbers of antennas and subcarriers. To reduce this overhead, most deep learning approaches compress CSI by treating it as a generic image, leaving its low-rank multipath structure unexploited. In contrast, model-driven alternatives explicitly embed this structure in iterative recovery, but at the cost of high computational complexity and latency. To overcome these limitations, we propose a low-rank prior-guided (LRP) framework that performs learnable rank-one sensing at the user equipment to compress CSI into low-dimensional codewords, and reconstructs CSI by direct rank-one synthesis at the base station. Both compression and reconstruction are jointly optimized, and fully exploit the low-rank structure of CSI. We further develop DCRNetV2, which preserves LRP as its backbone and uses gated dilated-convolutional residual paths to compensate for finite-rank errors. Experimental results show that LRP outperforms iterative model-based methods with lower complexity, and DCRNetV2 achieves a better accuracy-complexity tradeoff than existing learning-based methods.

    https://arxiv.org/abs/2609.28156


    UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation

    oai:arXiv.org:2609.28206v1

    arXiv:2609.28206v1 Announce Type: new Abstract: Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. However, reconstruction-oriented representations may be suboptimal for generation, motivating joint representation and generative learning. To this end, we introduce \textbf{Unite-Audio}, to our knowledge, is the \textbf{first} to jointly learn continuous audio representations and latent flow matching for TTA. By coupling reconstruction with self-supervised generative prediction, Unite-Audio allows the generative objective to directly shape the latent space rather than treating it as a fixed intermediate representation. We further employ Flow-GRPO post-training to improve text-conditioned generation. Experiments show competitive TTA performance with a compact latent flow model, while ablation studies confirm the benefit of jointly learning the audio representation and generative model. Audio samples are available at https://runwushi.github.io/Unite-Audio.

    https://arxiv.org/abs/2609.28206


    Reward-Rate Congestion Games and Replicator--Dinkelbach Dynamics

    oai:arXiv.org:2609.28240v1

    arXiv:2609.28240v1 Announce Type: new Abstract: Reward rate is a key performance criterion in cyber-physical and robotic systems where time, workload, and coordination costs are limiting resources. We introduce reward-rate congestion games, where agents seek to maximize reward per unit execution time. The direct reward-rate game is generally not an exact potential game. We develop a Dinkelbach-based framework in which, for every fixed Dinkelbach parameter, the transformed game is an exact potential game. This yields a potential-level Dinkelbach iteration that terminates finitely at the optimal potential reward rate when the inner potential maximization problem is solved globally. We also provide a sufficient condition under which an equilibrium of the transformed game is an equilibrium of the original reward-rate game. To optimize aggregate performance, we introduce marginal externality corrections that make the corrected potential coincide with the Dinkelbach-transformed social reward-rate objective, thereby enabling optimization of the social reward rate. Finally, we develop a continuous-time replicator--Dinkelbach dynamics for reward-rate population games coupling fast replicator dynamics with a slow reward-rate update. We establish convergence of the fixed-parameter replicator dynamics, global asymptotic and local exponential stability of the reduced Dinkelbach dynamics, and local exponential stability of the coupled system for sufficiently slow Dinkelbach updates. The framework is illustrated on a continuous task-allocation problem.

    https://arxiv.org/abs/2609.28240


    Modularity is Not Enough: Demonstration of a Solderless 400 V DC, 2.5 kW Three-Phase Inverter

    oai:arXiv.org:2609.28244v1

    arXiv:2609.28244v1 Announce Type: new Abstract: This paper presents the design and experimental evaluation of a fully solderless realization of a 400 V, 2.5 kW GaN-based variable speed drive (VSD), using screw-clamped resin molds and rubber compression pads instead of soldered interconnections. The power stage uses 650 V GaN power transistors and is operated at a switching frequency of 200 kHz. The solderless demonstrator is compared to a soldered reference realization using an identical printed circuit board (PCB). Over 120 thermal cycles with heatsink temperatures up to 90 {\deg}C, the solderless contacts show no degradation in effective on-state resistances (including contact resistances). Separately, open-loop vibration sweeps from 5 Hz to 2 kHz with acceleration amplitudes above 10 g were performed on the solderless assembly and left the continuously powered demonstrator electrically intact; subsequent resistance and nominal-power checks likewise indicate no contact degradation. An initial life-cycle assessment (LCA) indicated a higher embodied carbon footprint for the solderless realization due to 3D-printed resin molds, whereas a prospectively evaluated injection-molding scenario reduces the carbon footprint to near that of the soldered reference. The solderless assembly furthermore enables non-destructive component replacement, as demonstrated after a power transistor failure, as well as component re-use. These results support the feasibility of repair-oriented, industrially relevant kilowatt-class solderless power converters.

    https://arxiv.org/abs/2609.28244


    PARAFAC-Based Low-Latency Angular Tracking for RIS-Assisted Sensing

    oai:arXiv.org:2609.28255v1

    arXiv:2609.28255v1 Announce Type: new Abstract: This paper proposes adaptive tensor tracking for angular trajectories (ATTRACT), a structure-aware alternating least squares algorithm for low-latency target tracking in reconfigurable intelligent surface (RIS)-assisted sensing. The received echo is represented as a dynamic third-order PARAFAC tensor that separates angular, delay, and Doppler information. Unlike batch tensor estimators, ATTRACT processes one data slice at a time and carries forward the preceding factor estimates, avoiding the need to collect an entire sensing window before updating the target state. Its distinguishing feature is the integration of delay and Doppler structure into each iterative update through the known transmitter-RIS channel, pilot signals, and RIS configurations. Numerical results show that, for a large enough RIS, ATTRACT achieves accuracy comparable to the offline tensor baseline at high signal-to-noise ratios while reducing computational overhead and providing slot-wise estimates.

    https://arxiv.org/abs/2609.28255


    Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3

    oai:arXiv.org:2609.28273v1

    arXiv:2609.28273v1 Announce Type: new Abstract: State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.

    https://arxiv.org/abs/2609.28273


    A multi-resolution spectrogram approach for estimating the physical parameters of a plate reverb

    oai:arXiv.org:2609.28320v1

    arXiv:2609.28320v1 Announce Type: new Abstract: The ResNet-18 image classification model is employed to determine the physical parameters of a plate reverb from a recording of the impulse response. The model is adapted to derive parameters using normalized and downsampled multi-resolution spectrograms computed from the provided impulse responses (IRs). To refine the prediction of the output location, the spectral phase response is also included as an additional input channel to the network since multiple output locations can give the same magnitude response for high-order resonant modes. On a 5000 IR validation set, our model achieves an average normalized mean squared error (NMSE) of 0.02920 across all parameters, with the lowest average NMSE occurring for parameters yo (0.00228), Ly (0.00347), and xo (0.00574).

    https://arxiv.org/abs/2609.28320


    Neural Field-of-View for Binaural Signal Matching with Wearable Microphone Arrays

    oai:arXiv.org:2609.28343v1

    arXiv:2609.28343v1 Announce Type: new Abstract: The growing use of spatial audio in applications such as augmented and virtual reality has driven the development of binaural reproduction methods for wearable arrays with a limited number of microphones. Binaural signal matching (BSM) is one such method, producing high-quality binaural signals under a diffuse-field assumption, but degrading at high direct-to-reverberant ratios (DRR) where the direct sound dominates. Previous extensions incorporate Field-of-View (FoV) weighting, either with fixed apertures or based on explicit source localization, but these approaches are limited by coarse spatial coverage or reliance on localization estimation accuracy. This paper introduces FoV-BSM-Net, a signal-dependent FoV-BSM formulation that avoids explicit source estimation by learning the FoV parameters end-to-end from the microphone signals using a Convolutional Recurrent Neural Network. The method is evaluated in simulated rooms across varying reverberation conditions, and compared against BSM and a fixed FoV-BSM baseline. Results show that FoV-BSM-Net consistently improves over BSM, with gains that grow with DRR in both binaural NMSE and interaural cue errors, and are further supported by perceptual evaluation showing a substantial advantage over both baselines across low and high DRR conditions.

    https://arxiv.org/abs/2609.28343


    Benchmarking Curvature-Domain Signaling for Continuous-Aperture Wireless Communications: Capacity, Robustness, Detection, and Conditioning Against Legacy Modal Bases

    oai:arXiv.org:2609.28353v1

    arXiv:2609.28353v1 Announce Type: new Abstract: Continuous-aperture and holographic MIMO systems motivate signaling that operates directly on large electromagnetic apertures rather than on a few antenna ports. This paper is the empirical companion to the operator-theoretic curvature-domain framework: a reproducible benchmark of curvature-domain signaling under a common scalar aperture-channel model, stress-testing the theory's dual-budget generalized water-filling law against the modal bases used in near-field and holographic MIMO. All methods share the same apertures, quadrature, Fresnel or Green-function propagation, phase-only constraints, power, phase-energy and curvature-energy budgets, receiver noise, phase quantization and training assumptions. The compared coordinates are curvature-regularized eigenmodes, raw phase coefficients, Fourier phase modes, polynomial and Zernike-like wavefront modes, near-field matched-focus profiles, random and optimized RIS phase codebooks, and SVD water-filling upper bounds. The claim is deliberately limited: curvature-domain signaling is a gauge-invariant, physically realizable coordinate system that can approach the phase-space SVD water-filling bound with fewer stable modes where derivative noise, phase quantization, sampling density or ill conditioning limit conventional bases. It is not a claim of new electromagnetic physics, nor that curvature modes dominate every baseline. We prove the SVD upper-bound relation for the discretized phase-control tangent space, derive pairwise-error and perturbation bounds, and report capacity, retained modes, symbol error, robustness, quantization, sampling, regularization, conditioning and cost, with 95% bootstrap confidence intervals on Monte Carlo results. A verdict table locates the regimes where curvature-domain signaling is engineering-relevant: moderate-to-high phase noise, coarse phase quantization, and non-Fourier-diagonal channels.

    https://arxiv.org/abs/2609.28353


    A partitioned fluid-structure interaction solver for two-phase sloshing and flexible spacecraft dynamics

    oai:arXiv.org:2609.28355v1

    arXiv:2609.28355v1 Announce Type: new Abstract: This paper presents a high-fidelity direct numerical simulation (DNS)-fluid-structure interaction (FSI) framework for rigid-liquid-flexible spacecraft dynamics under microgravity conditions. The liquid-gas flow is simulated with the incompressible two-phase solver implemented in DIVA, validated against FLUIDICS experiments conducted aboard the International Space Station (ISS). The flexible appendages are described by a rotating assumed-mode plate model that accounts for geometric stiffening. The fluid and structural operators are coupled through a Dirichlet-Neumann fixed-point algorithm with Aitken relaxation, and a closed-system mechanical energy balance is used as an a posteriori diagnostic to assess the energy imbalance of the partitioned discretisation. The coupling strategy is validated against an experimental free-decay sloshing benchmark, and its numerical consistency is assessed through spatial sensitivity studies of the energy-balance defect. Prescribed-motion, rigid open-loop, and flexible open-loop simulations of a spin-up manoeuvre are compared to isolate the effect of structural feedback on the sloshing response. Reduced liquid models identified from the different simulation architectures exhibit different predictive capabilities when embedded in the same rigid-flexible plant. A controller synthesized from the reduced model identified from the flexible simulation is replayed in the nonlinear CFD-FSI environment. The reduced model reproduces the principal attitude and actuator responses for the considered manoeuvre but does not recover the detailed nonlinear sloshing-load history. The framework provides a high-fidelity environment for analysing coupled spacecraft dynamics, identifying control-oriented models, and assessing reduced-model-based control strategies beyond their linear design representation.

    https://arxiv.org/abs/2609.28355


    Curvature-Domain Wireless Communications: Gauge-Fixed Signal Spaces, Fredholm Capacity, and Differentiation-Limited Scaling for Continuous Apertures

    oai:arXiv.org:2609.28363v1

    arXiv:2609.28363v1 Announce Type: new Abstract: We develop a curvature-domain formulation for continuous-aperture signaling, in which the transmit phase is represented through its second spatial derivative after quotienting out affine piston-and-tilt gauge freedom. The resulting gauge-fixed synthesis operator is bounded and compact, with sharp Poincare-Wirtinger constant $C_L=L^2/\beta_1^2$, and its modal Gram spectrum is available in closed form, $\rho_m=(L/\beta_m)^4$, with $\beta_m$ the roots of $\cos\beta\cosh\beta=1$. Under a bounded-support square-integrable propagation kernel the tangent operator is Hilbert-Schmidt, so the infinite-dimensional capacity is a well-defined Fredholm-determinant supremum. The optimal signaling law is a dual-budget generalized water-filling with one Lagrange multiplier for curvature power and one for phase excursion. From the exact nonlinear phase-only aperture law we derive the coherent tangent channel with an explicit Frechet remainder bound and a multi-chart atlas for large excursions. At the receiver, curvature inferred from noisy phase samples by second differences has a pentadiagonal noise covariance with spectral norm $\Theta(\Delta x^{-4})$. A deterministic diagnostic suite measures each mechanism against its closed form, with a null run beside every claim: the computed spectrum matches $(L/\beta_m)^4$ to relative error $3.78\times10^{-15}$; the tangent remainder has fitted slope 1.0000; the dual-budget law is solved with both multipliers strictly active to a KKT residual of $3.90\times10^{-15}$; the derivative-noise bound is approached to 0.999981; and the differentiation-limited branch is observed at exponent 0.2175, then collapses to 0.0576 once the mode count saturates at the Shannon number, while a flat-propagation null run holds at 0.2019. Curvature is thus a well-posed, gauge-invariant coordinate, and below the Shannon number the gauge, rather than the medium, governs the scaling.

    https://arxiv.org/abs/2609.28363


    Optimal Guidance with Terminal Intercept-Angle Constraints and Acceleration Bounds

    oai:arXiv.org:2609.28381v1

    arXiv:2609.28381v1 Announce Type: new Abstract: Terminal intercept-angle control against a maneuvering target can substantially increase the required missile acceleration, potentially leading to saturation and interception failure unless acceleration limits are explicitly addressed. The engagement is therefore formulated as a linear-quadratic optimal-control problem with bounded acceleration commands. Polynomial approximations of the line-of-sight projection coefficients are used to better represent the nonlinear engagement geometry and estimate the time-to-go. The bounded optimal command is derived over saturated and unsaturated arcs, whose switching times are computed at each guidance step. The guidance law is derived for arbitrary linear missile dynamics and implemented for zero-order missile dynamics. For the zero-order model, the conditions under which the terminal demands can be met are derived in closed form, yielding the minimum and maximum reachable commanded terminal intercept angles. Performance is evaluated in nonlinear simulations. Compared with its unconstrained counterparts, the bounded formulation yields substantially smaller miss distances and terminal-angle errors when saturation is encountered. Unlike corresponding bounded miss-only guidance laws, the proposed law does not reduce to its unconstrained counterpart for minimum-phase missile dynamics because the acceleration command can saturate near the end of challenging engagements. The bounded law anticipates this saturation and compensates through earlier maneuvers.

    https://arxiv.org/abs/2609.28381


    Empirical Analysis of Near-Field Beam Shaping for Blockage Events

    oai:arXiv.org:2609.28400v1

    arXiv:2609.28400v1 Announce Type: new Abstract: Millimeter-wave (mmWave) to sub-terahertz (sub-THz) links can realize high data rates, yet are highly vulnerable to dynamic blockages due to high directivity, limited multipath, and dependence on line-of-sight (LoS) propagation. While structured and self-healing beams are promising solutions, it remains unclear how beam types perform under dynamic blockage. In this paper, we present an empirical analysis of structured near-field beams using two transmission policies. At one extreme, we study beam resilience without obstacle adaptation. At the other extreme, we empirically optimize beam parameters to study the limits of multiple beam types under idealized adaptation. We employ numerical analysis based on the angular spectrum method (ASM) and experimental validation using a sub-THz time-domain spectroscopy (TDS) platform to characterize beam performance over complete blockage events. For the scenarios considered, despite lacking self-healing, focused beams provide the strongest resilience when beam parameters are not adapted during a blockage event (i.e., when all beams are optimized only a priori). In contrast, when optimally adapted to obstacle position, curved beams provide the best performance. Lastly, although Bessel beams benefit from self-healing, this property alone can be insufficient to outperform optimized curved and focused beams. These findings provide guidance for blockage-aware beam selection and adaptation protocols.

    https://arxiv.org/abs/2609.28400


    Robustness Guarantees for Optimal RIS Placement in Self-Localization Under Placement Errors

    oai:arXiv.org:2609.28402v1

    arXiv:2609.28402v1 Announce Type: new Abstract: This paper studies the effect of Reconfigurable Intelligent Surface (RIS) placement errors on optimal RIS deployment for Time-of-Flight (ToF)-based self-localization. We consider a setup in which a source transmits a signal, receives the RIS-reflected echo, and estimates its own position from the corresponding ToF measurements. Under the conditions of moderate, i.i.d. Gaussian RIS position errors and i.i.d. Gaussian ToF noise, we demonstrate that the placement uncertainty manifests as a geometry-independent scaling factor in the Cram\'er-Rao Bound (CRB). Consequently, we prove that the optimal RIS placement remains unchanged by the presence of RIS position errors. This result is established for A-, D-, and E-optimality criteria. Simulation results verify the analysis and show that optimized RIS deployments retain their advantage even in the presence of placement uncertainty.

    https://arxiv.org/abs/2609.28402


    Two-impulse Rendezvous Planning about Thrusting Spacecraft on $\mathrm{SE}_2(3)$

    oai:arXiv.org:2609.28413v1

    arXiv:2609.28413v1 Announce Type: new Abstract: The classical Hill--Clohessy--Wiltshire equations assume an unforced Keplerian reference trajectory, an assumption that is violated by missions requiring continuous thrust. We address this limitation with a relative motion framework on the $\mathrm{SE}_2(3)$ Lie group that encodes position, velocity, and attitude in a unified geometric state. For computational tractability we linearize both the gravity mismatch and the body-frame control mismatch between the two vehicles, deriving tight analytic upper bounds on the neglected higher-order terms in each case. Under circular coasting Keplerian assumptions the framework recovers the Hill--Clohessy--Wiltshire equations exactly, establishing classical rendezvous theory as a special case rather than an independent linearization. For thrusting reference trajectories the state transition matrix acquires off-diagonal attitude--translation coupling blocks absent from classical formulations, and absorbing the control mismatch re-centers the linearization at the mean of the two vehicles' inputs. A two-impulse rendezvous planner derived directly from the state transition matrix accounts for both effects. Numerical simulations confirm recovery of the classical equations to machine precision, demonstrate successful rendezvous about a thrusting reference where classical planners fail, and validate the gravity and control mismatch linearization bounds throughout the transfer.

    https://arxiv.org/abs/2609.28413


    Key Reconciliation with RC-LDPC/Error Estimation for Satellite-based FSO/QKD Systems

    oai:arXiv.org:2609.25646v1

    arXiv:2609.25646v1 Announce Type: cross Abstract: Satellite-based free-space optics (FSO) quantum key distribution (QKD) systems have recently attracted significant research interest due to their potential to enable globally secured applications. However, the inherent uncertainty of FSO channels, caused by weather conditions and satellite mobility, induces severe fluctuations in quantum bit-error rate (QBER) between legitimate users. This makes designing an efficient key reconciliation, an essential step in the QKD post-processing, particularly challenging. In this work, we propose a key reconciliation scheme that combines protograph rate-compatible (RC) low-density parity-check (LDPC) codes with a syndrome-based error estimation method. The proposed error estimation method reduces the number of communication rounds without requiring additional information disclosure. Furthermore, to our best knowledge, an analytical framework is first developed to evaluate end-to-end secret-key throughput (SKT), accounting for the impact of imperfect error estimation and dynamic FSO channel conditions. Numerical results demonstrate that the proposed scheme consistently outperforms conventional blind reconciliation under diverse FSO channel conditions and provide practical guidelines for system parameter selection. Finally, we validate the proposed framework through a case study that incorporates data from a Starlink low-Earth orbit (LEO) satellite and moving ground vehicles.

    https://arxiv.org/abs/2609.25646


    Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

    oai:arXiv.org:2609.26823v1

    arXiv:2609.26823v1 Announce Type: cross Abstract: Post-training quantization of speech language models is often summarized with text-output scores and nominal bit widths. Those numbers alone do not establish behavior that depends on information missing from a transcript, or efficiency for a particular runtime. We introduce an evaluation protocol that separately tests lexical output, a transcript-insufficient endpoint, and a measured packed implementation. In a Qwen2-Audio case study, a translation-selected 6-bit allocation improves chrF by 2.36 on a frozen English-to-German replay, with paired 95% bootstrap interval [1.04, 3.62], but loses 3.91 percentage points on speaker-disjoint emotion recognition. At the same 6-bit budget, the uniform structural control reaches higher emotion accuracy than the selected allocation, and the front-layer control is also higher by point estimate on the same frozen set. At 7 bits, chrF improves by 3.28 with interval [2.08, 4.59], the emotion interval against FP16 includes zero, and a same-budget front-layer control still exceeds the selected allocation. A separate matched-budget 4.08-bit study finds roughly 10-point emotion deficits for every tested low-bit allocation and no selected-allocation advantage over frozen controls. Finally, a dequantized average-6-bit simulation retains the FP16 peak memory. This case study identifies a precision-dependent mismatch between lexical output, waveform-dependent behavior, and nominal precision. It does not establish a general failure of low-bit speech models or a deployment benefit for the selected allocation.

    https://arxiv.org/abs/2609.26823


    FINN-Tro: Exploiting Verification Gaps in Dataflow Inference Accelerators

    oai:arXiv.org:2609.26824v1

    arXiv:2609.26824v1 Announce Type: cross Abstract: The growing adoption of dataflow accelerators for neural network inference introduces new attack surfaces that existing verification methodologies fail to address. Inference ac- celeration frameworks such as FINN, which transform quantized neural networks into FPGA-deployable dataflow architectures, implicitly assume semantic equivalence between the software model and the synthesized hardware. In this work, we in- troduce the FINN-Tro attack, which identifies and exploits a critical verification gap in the FINN compilation pipeline that enables stealthy hardware Trojan insertion without modifying the original quantized model. The Trojan is placed in the last Matrix-Vector Activation Unit (MVAU) layer and supports two counter-based trigger modes, periodic and persistent, and three payload types: bias addition, logit swapping, and bias subtraction, resulting in six different configurations. FINN-Tro is evaluated on an MNIST feed-forward network and a CIFAR-10 convolutional neural network deployed on a PYNQ-Z1 board. Across the evaluated configurations, accuracy reductions range from 0.90% to 82.84%, while throughput and runtime remain close to the corresponding baseline designs. The most severe configuration, persistent Bias Addition, reduces accuracy from 92.96% to 10.12% on MNIST and from 84.19% to 10.00% on CIFAR-10. The inserted logic introduces modest implementation overhead, with maximum LUT and FF increases of 6.71% and 7.49% for MNIST, and 2.50% and 3.98% for CIFAR-10, respectively. Our findings reveal that widely used pre- and post- compilation verification flows are insufficient for detecting such temporally delayed hardware manipulations, motivating the need for stronger verification mechanisms in accelerator toolchains.

    https://arxiv.org/abs/2609.26824


    nnFoundation: 3D Foundation Models for Radiology

    oai:arXiv.org:2609.26924v1

    arXiv:2609.26924v1 Announce Type: cross Abstract: Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.

    https://arxiv.org/abs/2609.26924


    A Stem-Agnostic Approach to Hybrid AI Music Detection

    oai:arXiv.org:2609.26956v1

    arXiv:2609.26956v1 Announce Type: cross Abstract: The inclusion of generative audio in the music production process has led to an increase in hybrid music tracks that blend authentic human performances with AI-generated stems, challenging traditional AI music detectors which operate in a binary setting. In this work, we propose a stem-agnostic framework for identifying synthetic audio sources within hybrid musical mixtures. We introduce the inspectrogram, a novel time-frequency representation that maps localized probabilities of synthetic content across the audio spectrum. By combining the inspectrogram with a Wiener filter estimating target stem energy dominance, a single CNN model evaluates whether the specific stem is generated. Trained on rendered hybrid mixtures and evaluated across various stem classes, our model achieves strong performance on high-frequency sources such as vocals, drums, and guitar, but struggles on the low-frequency, narrow-band bass. We conclude that the quality of separation impacts the detection accuracy and identify source separation as a primary bottleneck and a crucial direction for future research.

    https://arxiv.org/abs/2609.26956


    Resource-Efficient Distributed Recursive Gaussian Processes

    oai:arXiv.org:2609.26979v1

    arXiv:2609.26979v1 Announce Type: cross Abstract: Gaussian processes (GPs) provide a flexible framework for learning unknown functions from noisy measurements while quantifying predictive uncertainty, making them well suited for estimation in multi-agent systems. However, when measurements are collected by multiple agents, maintaining a unified GP model without centralized processing requires efficient distributed algorithms that can operate using local measurements and communication with neighboring agents. In this work, we develop two distributed recursive GP (RGP) algorithms for multi-output GP regression: ADMM-RGP and PDMM-RGP. We analyze the stability and convergence of both algorithms and develop parameter selection strategies to accelerate convergence, thus reducing the communication burden. The proposed methods are validated on a real-world multi-output wind dataset, and their convergence behavior is examined across communication graphs with varying connectivity. Numerical experiments demonstrate that ADMM-RGP and PDMM-RGP can significantly reduce communication relative to the state of the art, while maintaining comparable estimation accuracy and network-wide consensus.

    https://arxiv.org/abs/2609.26979


    Spiderbot: An Open-Source Energy-Efficient Hexapod with Passive Gravity Compensation

    oai:arXiv.org:2609.26989v1

    arXiv:2609.26989v1 Announce Type: cross Abstract: Hexapod robots can achieve static stability with fewer actuated joints than bipeds or quadrupeds, yet many platforms still use 3-DOF legs, increasing weight and continuous torque requirement with limited gain in locomotion capability on flat, inclined and moderately rough terrains. We release Spiderbot, an open-source hexapod that uses a 4-bar linkage with a passive spring to mechanically support body weight, with a 2-DOF per-leg design that substantially reduces energy consumption. This mechanism substantially offloads gravitational torque during standing stance consuming only 1.5W (reduction of over 90\% over the unsprung version and up to 96\% over other similar hexapods). The passive spring compensation extends to payloads of up to 3.25kg with no additional torque requirements. The platform enables long-duration deployments on a modest battery budget and costs under \$400, making it suitable for large-scale multi-agent experiments. We validate the locomotion capabilities of the platform with an RL policy trained in mjlab, including successful sim-to-real transfer, despite the complexity of the mechanism. The platform is evaluated on flat and rough terrains, slope up to $15^\circ$ and step obstacles. We release all the CAD files, assembling instructions, and full training and deployment code along with the model checkpoints at https://erc-bpgc.github.io/SpiderBot/.

    https://arxiv.org/abs/2609.26989


    Fast Direction-Conditioned Reachability for Motion Prediction Under Model Uncertainty

    oai:arXiv.org:2609.27077v1

    arXiv:2609.27077v1 Announce Type: cross Abstract: To avoid collisions, a robot must repeatedly predict where nearby agents may move, usually with an imperfect model of their dynamics. Reachable sets provide such predictions, but computing them when the system matrices themselves are uncertain can become computationally expensive and conservative for frequent replanning. Moreover, a planner often needs to know only how far an agent can move in one particular direction, for example toward the robot, rather than the complete reachable set. We propose a direction-conditioned reachability method for linear systems with uncertain state and input matrices. Given a query direction $d$, the method selects one admissible model $(A^\star,B^\star)$ whose reachable set extends nearly as far along $d$ as the reachable set of the entire uncertain model family, and then computes the reachable set of only this model with a standard reachability solver. On an uncertain linearized bicycle model, the complete selection-and-computation pipeline is about three times faster than computing the reachable set of the full uncertain family in the CORA toolbox, while its extent along $d$ is within $5\%$ of the full family's in the reported directions. We also use the method in a closed-loop multi-vehicle simulation in which the robot queries, at each replanning step, how far each nearby vehicle can move toward it, and replans to avoid the resulting sets.

    https://arxiv.org/abs/2609.27077


    ZO-COSMO: Index-Free One-Hop Mixing for Decentralized Zeroth-Order Optimization

    oai:arXiv.org:2609.27199v1

    arXiv:2609.27199v1 Announce Type: cross Abstract: Sparse communication in decentralized zeroth-order learning requires compatible peer-state coordinates. We characterize this one-hop condition and develop \textsf{ZO-COSMO}, coupling two-query estimation with average-preserving masked consensus using $q$ values per active link. Global supports serve all-neighbor mixing; matching updates require agreement only within each pair. We derive a sharp contraction-per-scalar bound within the matching class and convergence guarantees for the core and sparse-momentum updates. At fixed matching, exact moment identities characterize how shared directions preserve gradient-heterogeneity cancellation and redistribute estimation error and disagreement. Mechanism experiments cover unequal curvatures, noise, and sparse momentum. Further tests span $64$ synthetic agents and eight logical Qwen LoRA workers. At matched payload budgets, Qwen2-7B QNLI gains $3.65$ accuracy points over explicit-index Rand-$k$; edge-local updates gain $3.42$ and $2.53$ points over all-neighbor mixing on eight-worker complete and ring graphs. A matched-first-step ablation gives a $3.92$-point momentum benefit. Seed-aware and same-matching controls distinguish encoding, scheduling, and query correlation.

    https://arxiv.org/abs/2609.27199


    Physiologically Informed Digital Auscultation for Pneumonia Detection in Long-term Care Residents

    oai:arXiv.org:2609.27222v1

    arXiv:2609.27222v1 Announce Type: cross Abstract: Pneumonia is difficult to diagnose in older long-term care residents; multimorbidity and atypical presentations obscure signs, motivating operationally efficient objective testing. We analyzed multi-channel digital stethoscope recordings from 185 Japanese residents (73 pneumonia, 112 symptomatic without), using radiologist-confirmed chest X-rays and clinician diagnoses as supervisory signals that train convolutional neural networks, multimodal fusion, and channel-based variants with time-domain Grad-CAM interpretability. Models were evaluated with repeated patient-level cross-validation showing models with X-ray supervision outperformed clinician supervision (F1 0.729, accuracy 0.783 vs. F1 0.637, accuracy 0.711). Additionally, a three-channel selection protocol maintained performance (F1 0.736; accuracy 0.803), with two mid-thoracic sites ranking highest and Grad-CAM attention overlapping adventitious sounds. These findings indicate automated multi-channel lung-sound analysis can aid long-term care pneumonia diagnosis, with X-ray supervision being more reliable than clinical, and fewer channels preserving performance while lowering acquisition times.

    https://arxiv.org/abs/2609.27222


    Full-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation

    oai:arXiv.org:2609.27244v1

    arXiv:2609.27244v1 Announce Type: cross Abstract: A neural network's layers can be treated as time steps of a state-space model, turning Bayesian training into a smoothing problem: a forward pass propagates Gaussian moments through the network, and a backward Rauch--Tung--Striebel pass updates the weight posteriors in closed form. Such methods learn from each observation in a single pass, in an uncertainty-aware manner, and without gradient-based iterations or replay, which makes them well suited for online adaptation and data-efficient learning. Existing smoothing-based methods, however, are restricted to diagonal covariances across activations, discarding correlations between neurons. We overcome this limitation via a cross-covariance identity that enables full-covariance propagation through a network's nonlinear activations. We derive a one-step-per-layer smoother that approximates as Gaussian only each layer's affine output, and that applies both to deterministic systems with noisy observations and to stochastic systems described by output statistics. We demonstrate this method in non-stationary classification, online dynamics learning, and policy adaptation of a vision-language-action model, and find that it is generally more accurate than other smoothing-based methods.

    https://arxiv.org/abs/2609.27244


    A note on bistability of a two-gene competitive system

    oai:arXiv.org:2609.27270v1

    arXiv:2609.27270v1 Announce Type: cross Abstract: Positive autoregulation together with mutual competition is one of the simplest mechanisms that can produce bistability in gene-regulatory models. We study a two-gene system in which each gene activates its own expression and the two genes compete through regulatory terms with Hill exponent one. We show first that the system has at least one and at most three equilibria in the positive quadrant. Exactly two positive equilibria can occur only at a degenerate nullcline tangency; thus, in the nondegenerate case, the number of positive equilibria is one or three. If there are exactly three distinct positive equilibria, then no nondegeneracy assumption is needed: all three equilibria are automatically hyperbolic, the two outer equilibria are asymptotically stable nodes, and the middle equilibrium is a saddle. Moreover, every positive solution converges to an equilibrium. Consequently, the positive quadrant is the disjoint union of the basins of attraction of the two stable nodes and the one-dimensional stable manifold of the saddle, yielding global bistability.

    https://arxiv.org/abs/2609.27270


    Graph Learning with Spectral Connectivity Priors for Scarce Data

    oai:arXiv.org:2609.27278v1

    arXiv:2609.27278v1 Announce Type: cross Abstract: Learning a sparse graph from scarce data is practically important but challenging. Motivated by the desirable combination of local sparsity and strong global connectivity exhibited by expander-like graphs, we propose spectral connectivity-regularized graph learning (SCoGL), a framework that incorporates a family of Laplacian spectral priors to explicitly promote global connectivity. Specifically, SCoGL augments a combinatorial-Laplacian-constrained graphical lasso (GLASSO) objective over a target adjacency matrix $\mathbf{W}$ with a general connectivity prior computed from Laplacian eigenvalues. We derive gradients for several representative connectivity priors and develop a projected gradient descent (PGD) algorithm with Armijo backtracking to efficiently optimize $\mathbf{W}$. Experiments show that the proposed SCoGL variants improve graph recovery and enhance downstream tasks such as graph signal denoising when signal observations are scarce.

    https://arxiv.org/abs/2609.27278


    Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation

    oai:arXiv.org:2609.27296v1

    arXiv:2609.27296v1 Announce Type: cross Abstract: Validating Autonomous Driving Systems (ADS) in simulation requires testing architectures that can discover rare, safety-critical failures while generating scenarios that are executable, diverse, and useful for downstream failure analysis. We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery. A high-reasoning Teacher LLM acts as an adaptive search controller, while a low-reasoning Student LLM emits simulator-executable scenarios in a strict JSON schema. The Teacher intervenes only when rolling collision rate and time-to-collision metrics stagnate, providing strategic guidance to redirect the search. In a CARLA case study with two experimental setups that vary the ego vehicle's speed policy, Teach-to-Crash achieves the highest Collision Hit Rate (90.79%), the shortest mean Time-to-Collision (18.31 s), and a competitive Collision Discovery Rate (136.21). PAFOT attains a higher mean CDR (179.44), but with substantially larger variance. Teach-to-Crash also yields the highest diversity (0.547) and, averaged across both setups on the CARLA Traffic Manager controller, the highest avoidability-based usefulness proxy (60.04%) among the compared methods. These results, within the evaluated CARLA scope, provide evidence that closed-loop dual-LLM reasoning can steer adversarial simulation-based testing over a constrained executable program space, generating failures that are frequent, structurally diverse, and assessed as more frequently avoidable.

    https://arxiv.org/abs/2609.27296


    Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning

    oai:arXiv.org:2609.27312v1

    arXiv:2609.27312v1 Announce Type: cross Abstract: Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.

    https://arxiv.org/abs/2609.27312


    Anomaly-Free Self-Optimization via AUC Bounds

    oai:arXiv.org:2609.27362v1

    arXiv:2609.27362v1 Announce Type: cross Abstract: Anomalies are rare, and anomalous data are often unavailable during development, making it difficult to determine which anomaly detection models and configurations will generalize to unseen anomalies. Recent approaches address this challenge by generating pseudo-anomalies and using bounds on the achievable area under the ROC curve (AUC) to select the optimal configuration from a finite set of candidates. Instead, we use the AUC bound as a differentiable, anomaly-free objective for directly optimizing continuous parameters of anomaly detection systems. We demonstrate this framework by optimizing ensemble weights and introducing a learnable score-rescaling mechanism that adapts pseudo-anomaly scores, enabling optimization beyond a predefined candidate set. Experiments across multiple datasets and embedding models show that AUC-bound optimization achieves significant performance gains over conventional model selection and prior development-set-based parameter selection. The results further show that direct optimization is less sensitive to the choice of pseudo-anomaly construction.

    https://arxiv.org/abs/2609.27362


    When Labels Are Scarce: An Oscillatory State Space Model for Vibration Diagnosis

    oai:arXiv.org:2609.27411v1

    arXiv:2609.27411v1 Announce Type: cross Abstract: Machine fault diagnosis from vibration requires learning from scarce labelled fault recordings while meeting the computational constraints of edge devices for local inference. We introduce DualRes, a compact oscillatory state-space model that combines two complementary spectral views of vibration, capturing rapid changes and fine frequency structure. Time-aligned views are processed by selective oscillatory memory, which learns how long to retain temporal patterns. The encoder contains 39,528 parameters. We evaluate supervised learning across six bearing datasets and a gearbox benchmark, with an additional gearbox pilot. Recording-level splits and explicit accounting of labelled duration distinguish data efficiency from repeated exposure to correlated samples. On the main gearbox benchmark, DualRes achieves state-of-the-art performance among the nine evaluated methods at six of seven label budgets. With about six labelled seconds per class, it improves macro-F1 by 16.1 percentage points over the next strongest comparator. On the same benchmark, DualRes achieves a 1.44-fold recording-level speedup and a 24.8-fold reduction in checkpoint storage relative to a selective state-space baseline under matched hardware and runtime conditions. Bearing results reveal task-dependent trade-offs. These findings support oscillatory memory as a compact approach to vibration diagnosis under limited labelled exposure.

    https://arxiv.org/abs/2609.27411


    Stable Neural Decoding Across Sessions via Task-Conditioned Latent Alignment for Brain-Machine Interfaces

    oai:arXiv.org:2609.27441v1

    arXiv:2609.27441v1 Announce Type: cross Abstract: Achieving stable long-term neural decoding in invasive brain-machine interfaces (BMIs) remains challenging due to variations in recorded neural populations across sessions. Current latent alignment approaches may overlook task-dependent structure during cross-session adaptation. We propose Task-Conditioned Latent Alignment (TCLA), a framework that stabilizes neural decoding by learning a shared latent space. TCLA learns a low-dimensional source representation using neural reconstruction and continuous behavioral supervision. During target-session adaptation, the shared representation is fixed, while target neural activity is mapped into the source latent space by aligning source and target distributions separately for each task condition. We evaluated TCLA on seven nonhuman primate datasets spanning multiple tasks. In long-term cross-session evaluation, TCLA achieved a mean $R^2$ of $0.476\pm0.014$ with a negative $R^2$ failure rate of only 6.8\%. Across 1,356 within-subject session pairs, TCLA achieved a mean $R^2$ of $0.371\pm0.009$ with a failure rate of 6.8\%. Across 2,134 cross-subject session pairs, TCLA achieved a mean $R^2$ of $0.218\pm0.004$ with a failure rate of 12.9\%, substantially better than those of the comparison methods. These results demonstrate that by preserving behaviorally relevant and task-dependent latent structure, TCLA improves the robustness of neural decoding across recording sessions and subjects. The source code is publicly available at \href{https://github.com/FAMD-CASIA/TCLA}{https://github.com/FAMD-CASIA/TCLA}.

    https://arxiv.org/abs/2609.27441


    SatUnreal: A High-Precision Synthetic Dataset for Satellite Stereo Matching via Unreal Engine

    oai:arXiv.org:2609.27442v1

    arXiv:2609.27442v1 Announce Type: cross Abstract: 3D reconstruction from satellite imagery is essential for large-scale topographic analysis, yet the lack of high-fidelity training datasets with accurate occlusion labels remains a primary bottleneck. Existing benchmarks, such as US3D and WHU-Stereo, face inherent challenges in spatio-temporal mismatch -- environmental changes and shadow displacements between multi-view acquisitions -- and provide ambiguous ground truth in occluded regions due to LiDAR sparsity. In this paper, we propose SatUnreal, a high-precision synthetic dataset designed to fundamentally overcome these limitations through an Unreal Engine-based simulation pipeline. SatUnreal provides 10,000 stereo pairs with high resolution (0.3m GSD) and is characterized by: (1) Physical Geometry Simulation, replicating realistic satellite orbits by systematically varying baselines and azimuths; (2) Spatio-temporal Consistency, eliminating temporal noise through fixed virtual environments; (3) Topographic Diversity, spanning dense urban canyons to low-texture natural terrains; and (4) Mathematical Label Integrity, utilizing a novel two-step linetrace algorithm to generate flawless occlusion masks. Experimental results using SOTA iterative models demonstrate that models trained exclusively on SatUnreal achieve superior zero-shot transfer performance on real-world benchmarks (US3D, WHU-Stereo) compared to those trained on real datasets. Our findings prove that physically accurate synthetic data provides a more effective supervisory signal for learning geometric features than complex real-world observations, establishing a new paradigm for Sim-to-Real transfer in Earth Observation. Code and dataset are available at https://github.com/jmp-Telepix/SatUnreal_A_High-Precision_Synthetic_Dataset_for_Satellite_Stereo_Matching_via_UnrealEngine

    https://arxiv.org/abs/2609.27442


    Safety-Filtered Distributed Koopman-MPC

    oai:arXiv.org:2609.27463v1

    arXiv:2609.27463v1 Announce Type: cross Abstract: Distributed model predictive control (DMPC) often constructs both predictions and collision constraints from neighbor trajectories, so packet loss can remove both. We separate these roles: received trajectories drive Koopman-MPC, while local sensing and shelf geometry define a hard-constrained quadratic program (QP) that projects the applied input. Its radial demand is the least constant acceleration that keeps a supporting-plane clearance nonnegative throughout one zero-order-hold interval. Complementary pair rows recover the coupled demand without exchanging safety decisions. We give an intersample separation theorem under bounded snapshot and directional plant errors, an exact max-min test for simultaneous local feasibility, and a sensing-radius condition for switching interaction graphs. Anticipatory high-order rows may be relaxed for performance, but the finite-hold rows contain no safety slack. Matched eight-robot warehouse simulations use a frozen Koopman model, nonlinear drift, bounded inputs and speed, shelf constraints, a 120 ms control period, and packet dropout. The full controller is collision-free in 20/20 matched trials and reaches 160/160 robot goals; predictive Koopman-MPC without the final projection is collision-free in 1/20 trials. All 38,400 full-method hard-row sets pass the online feasibility test, and every local QP solves. Five-stream fleet sweeps are collision-free and hard-row feasible through 16 robots; the 20-robot boundary fails only after the online margin turns negative, while the reconstructed per-agent critical path remains below the sampling period. Bounded-sensing and differential-drive tests provide additional deployment stress.

    https://arxiv.org/abs/2609.27463


    CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

    oai:arXiv.org:2609.27468v1

    arXiv:2609.27468v1 Announce Type: cross Abstract: Action-chunked vision-language-action (VLA) policies improve inference efficiency, but limited feedback within committed action chunks can lead to accumulated execution errors. Residual adaptation can correct such deviations without retraining the VLA; however, existing corrections are typically optimized for reference-action consistency without explicitly considering their downstream consequences. To address this limitation, we present Cerebellum-Inspired Consequence-Aware Residual Governance (CereVLA), a unified framework that integrates lightweight residual refinement and predictive consequence evaluation into frozen VLA execution. Corrective actions are first generated by flow-based residual refinement, and their short- and interval-horizon consequences are then evaluated by a recurrent state-space model and a history-aware classifier. Residual corrections predicted to be unfavorable are selectively suppressed by a lightweight governor. Comparisons with state-of-the-art methods on LIBERO-10 and LIBERO-GOAL demonstrate the effectiveness of CereVLA. On SO-101, CereVLA increases task success from 57.5% to 90.0% and reduces mean control steps by 19.6% among successful trials, relative to the frozen SmolVLA baseline.

    https://arxiv.org/abs/2609.27468


    Collocated Shape Regulation for Soft Robots

    oai:arXiv.org:2609.27469v1

    arXiv:2609.27469v1 Announce Type: cross Abstract: Controlling the shape of a continuum soft robot typically requires an accurate dynamic model and actuation of all degrees of freedom. We show that regulating only the actuated coordinates, through collocated shape control, achieves provably stable convergence of those coordinates and, under an explicit compatibility condition, of the entire robot shape. While collocated control is a cornerstone of high-performance motion control in rigid robotics, extending this formulation to continuum soft robots has remained challenging due to the complexity of their dynamics. We present the first general framework for collocated control of continuum soft robots and derive a unified family of controllers, including PD, PID, PsatID, and their counterparts with compensation and cancellation components. The framework unifies existing approaches while introducing new controller designs. In particular, we develop three classes of PD and PID like regulators with local, semi-global, and global stability guarantees, and provide rigorous convergence analyses for each. Extensive experimental validation demonstrates the effectiveness of the proposed methods across different model discretizations and controller parameters. The resulting framework provides practical design guidelines for selecting and implementing controllers with known stability guarantees, without requiring a complete dynamic model of the robot

    https://arxiv.org/abs/2609.27469


    Gray-Box Model Predictive Control for Articulated Dump Trucks via Gaussian Process Learning of Sideslip

    oai:arXiv.org:2609.27597v1

    arXiv:2609.27597v1 Announce Type: cross Abstract: The growing demand for automation in the mining industry, particularly for the autonomous operation of articulated dump trucks (ADTs), has drawn increased attention to accurate vehicle modeling. The importance of such models lies in their use in model predictive control (MPC), model-based estimation methods, and vehicle simulation. While dynamic modeling offers a viable solution for these purposes, it is associated with complex setup and parametrization and may require recalibration in changing operating environments. As a result, kinematic models have dominated ADT modeling, especially in MPCs, at the expense of reduced prediction accuracy. In this work, we propose an approach using Gaussian Process Regression (GPR) to learn the sideslip angle of the vehicle, which is identified as the primary contributor to the reduced accuracy of kinematic models. The learned GPR function is augmented into the kinematic model to form a gray-box model that aims to reduce the gap to dynamic models. We show that the gray-box model can predict the sideslip angle and, consequently, the vehicle's lateral velocity, thereby improving the MPC's prediction performance. The resulting gray-box MPC is compared against two white-box MPCs in a simulation environment. The results indicate an improvement in terms of maximum lateral tracking error from over 2 m to 0.56 m.

    https://arxiv.org/abs/2609.27597


    AI-Driven Neural Surrogates for In Silico Design of Cognitive-Affective Neuromodulation Targets

    oai:arXiv.org:2609.27729v1

    arXiv:2609.27729v1 Announce Type: cross Abstract: In neuropsychiatry, the primary goal is often not only to decode brain activity but to change it, for example to lessen a negative affective bias or an overly salient memory. Motivated by control theory, we develop an AI-driven neural-surrogate framework that proposes candidate representational changes and tests their predicted perceptual effects from snapshots of stimulus-evoked fMRI activity, without physical stimulation. The framework combines fMRI decoding, deep generative modeling, and constrained latent-space steering. Valence and memorability are used only as worked examples. Using more than 36,000 image-fMRI observations from four deeply sampled Natural Scenes Dataset participants, subject-specific models recovered coarse generative structure from visually responsive cortex (two-way identification, 0.79-0.88; chance, 0.5). Graded perturbations were reconstructed as images and evaluated with automated scorers and human ratings from 7,200 trials by 18 participants. In the primary VDVAE model, valence shifted from -0.61 to +1.03 SD and memorability from -1.34 to +1.45 SD; a later Versatile Diffusion refinement reduced or altered these effects. Across five perturbation levels, human valence ratings moved in the predicted direction under the linear time-correction model (mean slope, 0.038 SD per unit of alpha; 95 percent CI, 0.003-0.074; positive in 16 of 18 participants). Perceived memorability did not change reliably. Baseline agreement with the automated assessor was suggestive for valence (r = 0.30) and weak for memorability (r = 0.10). Extreme perturbations drifted from the original stimulus, so intended change must be weighed against loss of fidelity. These findings provide a falsifiable upstream method for designing and behaviorally testing candidate representational targets for future neuromodulation in psychiatry, while marking the limits of the present static approximation.

    https://arxiv.org/abs/2609.27729


    Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis

    oai:arXiv.org:2609.27816v1

    arXiv:2609.27816v1 Announce Type: cross Abstract: Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in sensing capabilities, environmental awareness, and motion execution roles. This paper presents a centralized safety-aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot team comprising a vision-capable quadruped and a camera-less robotic vehicle. The objective is to guide both platforms toward a goal region while avoiding static and dynamic obstacles and preventing unsafe inter-robot interactions. Under the principle of shared perception, the vision-capable robot provides semantic environmental awareness through a centralized server over an MQTT broker, enabling the camera-less platform to navigate using this shared scene representation alongside its own odometry, IMU, and state feedback. A vision-language model (VLM) interprets the visual stream, and the extracted semantic data is mapped into conservative metric geometric constraints, including inflated obstacle sets, safe corridors, and goal regions. A large language model (LLM) proposes high-level task allocations, while physical command authority is restricted to a robot-specific zonotope reachability gate. This verification engine propagates independent reachable tubes to evaluate obstacle avoidance, safe-corridor containment, and inter-robot separation predicates before approving commands. Online experiments across clear-path and dynamic-obstacle scenarios show that the pipeline reliably approves safe motion, triggers conservative replanning or holding maneuvers upon constraint violation, and enforces a strict architectural separation between advisory semantic reasoning and formally verified motor execution.

    https://arxiv.org/abs/2609.27816


    A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation

    oai:arXiv.org:2609.27828v1

    arXiv:2609.27828v1 Announce Type: cross Abstract: Legacy smart HVAC controllers rely on vendor-cloud TLS secured by ECDH and RSA, both broken by Shor's algorithm, and typical 10-15 year lifespans mean today's devices remain in service through the quantum-threat era. Direct on-device post-quantum cryptography is infeasible: an ESP32-S3, representative of capable HVAC hardware, has only 339 KB free heap against the 900 KB ML-KEM-768 requires, and even classical ECDH-P256 keygen (111.93 ms) dwarfs hardware AES-128 (0.032 ms). We propose a non-invasive PQC proxy, requiring no device, firmware, or vendor-cloud changes, performing ML-KEM-768 encapsulation and ML-DSA-65 authentication (NIST FIPS 203/204) with AES-256-GCM session keys via HKDF, implemented with Open Quantum Safe liboqs on a Raspberry Pi 4B gateway. Over 500 runs, the post-quantum handshake (Steps 1-6) completes in 2.48 ms, 0.38 ms slower than classical baseline, with PQC computation around 8% of handshake time at 20 ms simulated round-trip network latency. The gateway sustains 443 sessions/second, 100% success under 32 concurrent connections, extrapolating to 3546 sessions/second on a 32-core cloud instance. Five side-channel tests, including verified in-place session-key zeroization and a fixed-vs-random TVLA timing analysis, found no exploitable timing leakage or susceptibility to man-in-the-middle attacks. The architecture is vendor-agnostic and becomes unnecessary once vendors adopt NIST PQC natively.

    https://arxiv.org/abs/2609.27828


    Q-MAP: Multi-Platform Benchmarking of Distributed Quantum Computing for Coherent Controlled Islanding

    oai:arXiv.org:2609.27829v1

    arXiv:2609.27829v1 Announce Type: cross Abstract: The integration of distributed energy resources into power networks is accelerating. The resulting variability narrows operating margins, so a disturbance can cascade into a wide-area blackout. Controlled islanding arrests that propagation by splitting a compromised grid into self-sustaining islands that keep coherent generators together. Exact classical solutions become intractable as the bus count and island number grow. Gate-based quantum optimization provides a different route through this combinatorial space, although its reach is limited when one circuit carries every bus assignment, since qubit count and depth then follow grid size. In this study, a round-synchronous distributed quantum computing framework is developed for coherent controlled islanding under a fixed per-circuit qubit budget. Every round derives all regional subproblems from one frozen grid-wide snapshot, dispatches them to independent quantum backends at the same time, and merges the returned candidates classically into one globally evaluated update. Circuits executed in parallel therefore keep a constant size as the grid grows, and a round costs the slowest region rather than the sum of all of them. Benchmarking spans IEEE systems from 9 to 300 buses on simulation and on quantum processors of different architectures. The framework attains optimal and operationally feasible partitions on every platform under noise, even where compilation cost differs by nearly an order of magnitude. Bounding width in this way places grids beyond the reach of monolithic circuits within range of present devices and establishes a multi-backend baseline for quantum computing in large-scale power-system optimization.

    https://arxiv.org/abs/2609.27829


    Anchor-Free Hidden-Target Seeking via Certified Self-Calibration under Correlated Odometry

    oai:arXiv.org:2609.27905v1

    arXiv:2609.27905v1 Announce Type: cross Abstract: We study hidden-target seeking in an anchor-free regime where neither the vehicle, the relay, nor the target has an accessible global pose. The vehicle never senses the target directly; it receives only range-bearing observations through a single relay of unknown position and orientation, with motion known only via integrated body-frame odometry. Absolute localization is fundamentally impossible: the joint configuration retains an exact three-dimensional SE(2) gauge no estimator can resolve. Yet the quantities needed for control remain fully recoverable: motion collapses calibration to a task-relevant quotient, the relay-to-odometry yaw and target displacement, identifiable in closed form from two distinct vehicle views. We introduce an O(K) multi-view self-calibration estimator with an exact first-order yaw-uncertainty certificate that propagates the cross-view correlations integrated odometry induces: the full certificate attains 95.0% pooled coverage at the nominal 95% level, versus 82.1% when correlated poses are treated as independent. The certificate drives a hybrid policy that excites until calibration is trustworthy, refuses uncertified estimates, seeks using continuously re-measured geometry, and detects relay-frame changes via persistent certified inconsistency, avoiding unbounded dead-reckoning drift. Across 200 randomized closed-loop trials, the method attains 0.064 m median station error versus 0.065 m for an oracle given the true relay yaw, despite 27 m median dead-reckoning drift over long horizons; 90 physics-based ROS 2/Gazebo trials retain 30/30 success under nominal operation, communication degradation, and relay-frame disturbances. These results show global localization is unnecessary for reliable hidden-target seeking under this single-relay model, even when both the sensing infrastructure and the vehicle's own reference frame are uncalibrated.

    https://arxiv.org/abs/2609.27905


    A comparative assessment of global building and settlement datasets across geographic and settlement contexts

    oai:arXiv.org:2609.28154v1

    arXiv:2609.28154v1 Announce Type: cross Abstract: Global building and settlement datasets increasingly support population mapping, exposure assessment, urban monitoring, and other analyses of the built environment, yet comparative evidence remains fragmented across products, geographic regions, reference datasets, spatial scales, and evaluation methods. We benchmark seven global or near-global products, including Overture Maps, Global Building Atlas, 3D-GloBFP, Google Open Buildings 2.5D Temporal (OBT), Microsoft TEMPO, GHSL, and WSF Tracker, against harmonized reference footprints across 135 study areas. The evaluation combines complementary measures of detection, geometric agreement, and aggregate quantity accuracy, together with stratified analyses of settlement characteristics and diagnostic experiments on error size and temporal alignment. Overture achieved the highest median city-level vector F1 (0.786). Raster rankings were resolution-dependent: OBT achieved the highest median F1 at 10m (0.642), whereas WSF Tracker led at 100m (0.862). However, WSF Tracker substantially overestimated built-up area, emphasizing that when using raster products, it is important for the user to understand whether the raster identifies only buildings or includes additional impervious surfaces. Raster accuracy increased consistently with building density (Spearman \r{ho} = 0.58-0.75), while small candidate buildings were disproportionately associated with false positives in the vector products. Temporally aligning WSF Tracker with reference imagery increased mean F1 by 0.060 (median +0.037), indicating that the reported accuracies are conservative in rapidly growing areas. The study establishes a reproducible benchmark for comparing heterogeneous global urban and settlement layer datasets across geographic and settlement contexts.

    https://arxiv.org/abs/2609.28154


    Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination

    oai:arXiv.org:2609.28182v1

    arXiv:2609.28182v1 Announce Type: cross Abstract: Coordinating large population of flexible grid-edge devices can alleviate the need for time-consuming and capital-intensive network upgrades, and AI-based control methods such as multi-agent reinforcement learning or imitation learning are promising in their real-time decision scalability. However, system operators still need an independent and rigorous way to decide whether a given AI system is safe enough for deployment. This paper develops a finite-sample probabilistic safety certification framework for black-box AI decision models in closed-loop grid operation. The central idea is to reduce the complete input--AI--grid evaluator workflow to a binary unsafe outcome under an operator-defined safety specification, and then use exact binomial inference to certify the corresponding unsafe operation probability. Given a set of held-out calibration scenarios, the framework returns the tightest one-sided upper certificate and an accept/reject deployment criterion that controls the probability of false safety certification. Because the certification is for the calibration distribution that may deviate from the future operation, we further combine the nominal certificate with physically interpretable sample-space adversarial attacks, a concept widely used in AI to investigate the fragility of AI models. Case studies on grid-edge flexibility coordination with 1{,}000-agent AI models (independent parameters) verify the finite-sample safety guarantee and the value of integrating adversarial attacks into a rolling-window training-certification-deployment flow.

    https://arxiv.org/abs/2609.28182


    From ECG Signals to Representative-Morphology Heatmaps for Biometric Recognition

    oai:arXiv.org:2609.28183v1

    arXiv:2609.28183v1 Announce Type: cross Abstract: Electrocardiography (ECG) contains subject-specific morphology that supports biometric recognition, yet image-based performance depends on how the waveform is rendered. We introduce representative-morphology heatmaps, a deterministic ECG-to-image representation adapted from ECGXtractor. Within each block of ten aligned beats, the five beats closest to the block mean are averaged into a 400 by L matrix and rendered either as a conventional trace or as a dense cardiac-time-by-lead heatmap. Since both representations contain identical physiological samples, their comparison isolates the effect of rendering. We evaluate verification and closed-set identification on PTB, ECG-ID, and MIMIC-IV-ECG-DEMO. Five compact models, including ZACH-ViT, are trained from scratch, while six ImageNet-pretrained CNN and transformer backbones assess model scale and visual transfer. Heatmaps improve both FNMR operating points and both identification ranks in all 15 compact model-dataset comparisons, while EER improves in 14. Across the matched experiments, EER decreases by 9.59 percentage points and Rank-1 increases by 24.69 points on average. ConvNeXt-Tiny reaches 2.43% EER on PTB and 5.79% on ECG-ID, whereas DeiT-Base reaches 14.92% on MIMIC-DEMO. ImageNet initialization clearly benefits the two multilead datasets but has a mixed effect on ECG-ID, and performance does not increase monotonically with model size. The best heatmap systems approach the strongest signal-domain EER on PTB and ECG-ID, while DeiT-Base provides the strongest evaluated performance on MIMIC-DEMO. Lead-channel ablation further shows that useful channel combinations depend on the cohort and biometric task. Overall, representative-morphology heatmaps provide an effective image representation for ECG verification and identification.

    https://arxiv.org/abs/2609.28183


    Geospatial embeddings detect old-growth forests but buffered spatial validation narrows their advantage over Sentinel features

    oai:arXiv.org:2609.28194v1

    arXiv:2609.28194v1 Announce Type: cross Abstract: Old-growth forests develop over centuries under minimal anthropogenic disturbance, producing structurally complex and biodiverse stands. In Europe, protecting them requires mapping that is accurate for individual forest parcels yet deployable continent-wide. Geospatial foundation model (GFM) embeddings enable label-scarce land classification, but their value for old-growth detection remains unknown. Here, we map old-growth forests across 211,893 ha of Romania's Southern Carpathians, a beech-spruce landscape typical of the Alpine Biogeographic Region. We construct high-confidence, expert-informed reference labels for old-growth and non-old-growth parcels. We add AlphaEarth, TESSERA v2 and Sentinel-1/2 features to a common baseline of topographic and human-access predictors, then compare them under spatially blocked validation with and without 10 km train-test buffers to limit residual autocorrelation. With buffering, GFM and Sentinel-1/2 predictors increase precision-recall AUC by 0.21-0.25 [95% CIs: 0.15-0.34] relative to baseline, indicating spectral data contain a spatially robust old-growth signal. With a PR-AUC of 0.84 [0.79-0.88], TESSERA outperforms Sentinel-1/2 (+0.08 [+0.05 to +0.11]) and AlphaEarth (+0.08 [+0.04 to +0.12]) under unbuffered spatial validation. At a 10 km buffer, however, this advantage narrows to +0.04 [-0.01 to +0.11] and +0.03 [-0.04 to +0.10], intervals consistent with no difference. At 10 m resolution, convolutional neural networks add no benefit over pixel-based XGBoost. Comparisons with four national- and continental-scale products show the importance of non-old-growth labels, and reveal 81% agreement between our predictions and a field-calibrated map. We conclude that buffered spatial validation is vital when transferring old-growth detection models to unseen landscapes, and provide our labels and predictions for future work.

    https://arxiv.org/abs/2609.28194


    Do Electromagnetic Side-Channel Attacks Threaten Electronic Polling Stations? Scenarios and Recommendations

    oai:arXiv.org:2609.28209v1

    arXiv:2609.28209v1 Announce Type: cross Abstract: This paper investigates the threat to ballot secrecy in the Brazilian electronic voting machine (UEB) posed by electromagnetic side-channel attacks, also known as TEMPEST attacks. In these attacks, screen content can be reconstructed remotely by intercepting electromagnetic emanations associated with the target device's video signal. This work is motivated by a recent ruling by a Brazilian electoral court concerning an attempt to violate ballot secrecy using electronic equipment. Based on publicly available information about the electoral system, attack scenarios against polling stations are proposed. Experiments using software-defined radio show that the effectiveness of TEMPEST attacks strongly depends on the lack of oversight resulting from public unawareness of the threat. Finally, awareness guidelines are proposed for voters, poll workers, and party representatives to mitigate attack risks within a polling station.

    https://arxiv.org/abs/2609.28209


    Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward

    oai:arXiv.org:2609.28396v1

    arXiv:2609.28396v1 Announce Type: cross Abstract: This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platforms operating under unknown dynamics and strict actuator limits to satisfy complex high-level specifications. We denote these high-level specifications using Signal Temporal Logic (STL) and propose a novel time-aware Reinforcement Learning (RL) framework that leverages the geometric properties of Spatiotemporal Tubes (STTs). While traditional analytical STT controllers often struggle to enforce input constraints, and existing RL approaches rely on memory-intensive state history, our method natively overcomes both limitations. By mapping the logical and temporal complexities of the full class of STL into time-varying geometric boundaries, we directly constrain the multidimensional system state without relying on scalar robustness metrics. Augmenting the state space with time, we train a time-aware Soft Actor-Critic (SAC) agent using a continuous, geometry-aware reward function that eliminates the need to explicitly evaluate complex logical semantics during execution. The proposed framework offers a history-free, computationally efficient approach to learn continuous control policies that ensure robust satisfaction of specifications while strictly adhering to system input constraints.

    https://arxiv.org/abs/2609.28396


    Predicting the Progression of Adolescent Idiopathic Scoliosis

    oai:arXiv.org:2609.28434v1

    arXiv:2609.28434v1 Announce Type: cross Abstract: Adolescent Idiopathic Scoliosis is defined as a lateral curvature of the spine that develops during adolescence, without known cause. The condition can result in significant pain and disability, and often progresses rapidly during adolescence. The objective of this paper is to predict the progression of the condition in a temporal sequence from ages 9 to 24, as measured from a sequence of Dual X-ray Absorptiometry (DXA) scans. To this end, we train a transformer model that takes in the curve of the spine to predict curve progression. The model is trained using a large-scale synthetic dataset of spine curves and their time series, covering different curve types and different progression patterns. We show that the model is able to generalise from synthetic to real data by evaluating it on a dataset of real DXA scans covering multiple time points. We find that fine-tuning the model on real data gives a significant boost to performance. The model is able to accurately predict spine curve progression in both scoliosis and normal cases.

    https://arxiv.org/abs/2609.28434


    Towards optimal algorithms for the recovery of low-dimensional models with linear rates

    oai:arXiv.org:2410.06607v4

    arXiv:2410.06607v4 Announce Type: replace Abstract: We consider the problem of recovering elements of a low-dimensional model from linear measurements. From signal and image processing to inverse problems in data science, this question has been at the center of many applications. Lately, with the success of models and methods relying on deep neural networks, there has been a multiplication of different algorithms and recovery results. Comparing the performance of recovery algorithms becomes a complex task without a unifying framework. In this article, as a first step for the study of general algorithms for low-dimensional recovery, we study a class of generalized projected gradient descent algorithms that can recover a given low-dimensional model with linear rates. The obtained rates decouple the impact of the quality of the measurements with respect to the model from the geometry of the properties of the chosen generalized projection: we can directly measure performance through a restricted Lipschitz constant of the projection with respect to the low dimensional model. By optimizing this constant, we define an optimal generalized projected gradient descent. Our general approach provides an optimality result in the case of sparse recovery. Moreover, our framework allows for a common interpretation of linear rates of recovery in the context of both sparse models and models induced by some ``plug-and-play'' imaging methods that rely on deep neural networks. These rates of recovery are observed in experiments on synthetic and real data.

    https://arxiv.org/abs/2410.06607


    Tracking Driving Stressors through Multimodal Physiological Monitoring

    oai:arXiv.org:2507.14146v2

    arXiv:2507.14146v2 Announce Type: replace Abstract: Understanding and mitigating driving stress is important for improving road safety and driver well-being. Reliable estimation, however, requires distinguishing biobehavioral responses to individual stressors from gradual physiological and contextual changes. We collected physiological data and vehicle telemetry from 31 participants across 44 simulated-driving sessions containing controlled stressor events. Under cross-validation, a multimodal classifier achieved an AUROC of 0.768 when distinguishing the stressor phase from an earlier baseline, reflecting both stressor effects and temporal drift. Controlling for drift retained an AUROC of 0.661, but revealed stronger responses to sustained than brief stressors, and shifted feature attribution toward phasic cardiac and electrodermal markers. We further quantified the interaction between model-estimated physiological stress and observable changes in vehicle control through simulation telemetry. Our findings show that stressor-aware modeling can identify physiologically grounded responses that correspond to meaningful changes in driving behavior.

    https://arxiv.org/abs/2507.14146


    Discrete optimal transport is a strong audio adversarial attack

    oai:arXiv.org:2509.14959v4

    arXiv:2509.14959v4 Announce Type: replace Abstract: In this paper, we investigate discrete optimal transport (DOT) as a black-box attack against modern automatic speaker verification (ASV) and anti-spoofing countermeasure (CM) systems. Our attack operates as a post-processing distribution-alignment step. Frame-level WavLM embeddings of generated speech (or another person speech) are aligned to an unpaired bona fide speech pool using entropic optimal transport and a top-k barycentric projection, followed by neural vocoding. Unlike gradient-based attacks, the proposed method requires no access to model parameters, gradients, or training data. Experiments on ASVspoof2019 and ASVspoof5 demonstrate that DOT attack substantially increases CM EER and substantially degrades ASV performance across multiple spoofing attacks. The attack transfers across datasets and remains effective after CM fine-tuning. Analysis using speaker similarity, Fr\'echet Audio Distance, and visualization of embedding distributions suggests that DOT succeeds by shifting source speech toward bona fide regions of the representation space rather than by maximizing speaker similarity. These results indicate that optimal-transport-based distribution alignment represents a previously underexplored attack vector for contemporary ASV and anti-spoofing systems.

    https://arxiv.org/abs/2509.14959


    KELP: K-space-conditioned Estimation of Learned Sampling Patterns for Scan-Adaptive Multi-Coil MRI

    oai:arXiv.org:2509.16846v2

    arXiv:2509.16846v2 Announce Type: replace Abstract: Deep learning techniques have gained considerable attention for accelerating MRI acquisition while maintaining image quality. In this work, we present a convolutional neural network (CNN)-based framework for predicting scan-adaptive undersampling patterns directly from low-frequency multi-coil $k$-space data. Unlike approaches that optimize sampling patterns during training or rely on nearest-neighbor search at inference, our method is trained using precomputed scan-adaptive optimized masks as supervised labels and predicts a scan-specific sampling pattern in a single forward pass. The training procedure alternates between optimizing the sampling network and a reconstruction network, allowing the learned masks to adapt to reconstruction performance. Experiments on the fastMRI multi-coil knee dataset demonstrate competitive reconstruction quality compared with existing population- and scan-adaptive sampling approaches at $4\times$ and $8\times$ acceleration factors. In addition, the proposed method enables efficient scan-specific mask prediction with inference cost independent of the training dictionary size.

    https://arxiv.org/abs/2509.16846


    Adaptive double-phase Rudin--Osher--Fatemi denoising model

    oai:arXiv.org:2510.04382v3

    arXiv:2510.04382v3 Announce Type: replace Abstract: Even though more than 30 years have passed since the seminal Rudin--Osher--Fatemi (ROF) paper on total variation (TV) denoising, it remains relevant due to its simplicity, robustness and interpretability. However, it is known to suffer from artifacts such as the staircasing effect. Many variants of the model have been proposed with the aim of countering this. Recently, against the backdrop of immense research output on double-phase problems in the mathematical analysis community, a double-phase type integral functional, comprising of TV and a weighted term of quadratic growth, was suggested as a regularizer for image restoration. Here, we propose an adaptive variant of the ROF denoising model based on that regularizer. Variable growth of the double-phase functional allows for qualitatively different behavior at image contours, which are captured by an initial ROF reconstruction step. The model is designed to reduce staircasing with respect to the classical ROF model, while preserving the edges of the image in a similar fashion. We derive a closed-form resolvent formula and adapt the primal-dual Chambolle--Pock scheme for the numerical solution of the model. We also propose a practical noise-dependent parameter prescription and evaluate its performance on synthetic and natural images over a range of noise levels. Compared to established models with similar interpretability, we observe an improved or similar performance in terms of similarity metrics SSIM, PSNR, and LPIPS, while the staircasing effect is visibly reduced.

    https://arxiv.org/abs/2510.04382


    Observability and parameter estimation of a generic model for aggregated distributed energy resources

    oai:arXiv.org:2510.10892v2

    arXiv:2510.10892v2 Announce Type: replace Abstract: We propose a novel framework for estimating the parameters of an aggregated distributed energy resources (DER A) model. First, we introduce a rigorous method to determine whether all model parameters are estimable. When they are not, our approach identifies the subset of parameters that can be estimated. The proposed framework offers new insights into the number and specific parameters that can be reliably estimated based on commonly available measurements. It also highlights the limitations of calibrating such models. Second, we introduce a Kalman filtering method to calibrate the DER A model. Since we account for nonlinear effects such as saturation and deadbands, we develop a specific mechanism to handle smoothing functions within the Kalman filter. Specifically, we consider the extended and the unscented Kalman filter. We demonstrate the effectiveness of the proposed framework on a modified IEEE 34-node distribution feeder with inverter- based resources. Our findings align with the North American Electric Reliability Corporation's parameterization guideline and underscore the importance of model calibration in accurately capturing the collective dynamics of distributed energy resources installed on distribution systems.

    https://arxiv.org/abs/2510.10892


    Remote Magnetic Levitation Using Reduced Attitude Control and Parametric Field Models

    oai:arXiv.org:2512.15207v3

    arXiv:2512.15207v3 Announce Type: replace Abstract: Electromagnetic navigation systems (eMNS) are increasingly used in minimally invasive procedures such as endovascular interventions and targeted drug delivery due to their ability to generate fast and precise magnetic fields. In this paper, we utilize the OctoMag and a custom 13-coil eMNS to achieve remote levitation and control of multiple rigid bodies across large air gaps, showcasing the dynamic capabilities of such systems. A compact parametric analytical model maps coil currents to the forces and torques acting on the levitating object, eliminating the need for computationally expensive simulations or lookup tables and establishing a levitator- and platform-agnostic control framework. Translational motion is stabilized using linear quadratic regulators. A nonlinear time-invariant controller is used to regulate the reduced attitude accounting for the inherent uncontrollability of rotations about the dipole axis and stabilizing the full five degrees of freedom controllable pose subspace. We analyze key design limitations and evaluate the approach through trajectory tracking experiments across different objects and actuation platforms. Notably, our proposed controller demonstrates superiority over an equivalent baseline PID formulation, reliably tracking large spatial angles up to 65 degrees. This work demonstrates the dynamic capabilities and potential of feedback control in electromagnetic navigation, which is likely to open up new medical applications.

    https://arxiv.org/abs/2512.15207


    Advances in Diffusion-Based Generative Compression

    oai:arXiv.org:2601.18932v2

    arXiv:2601.18932v2 Announce Type: replace Abstract: Popularized by their strong image generation performance, diffusion and related methods for generative modeling have found widespread success in visual media applications. In particular, diffusion methods have enabled new approaches to data compression, where realistic reconstructions can be generated at extremely low bit-rates. This article provides a unifying review of recent diffusion-based methods for generative lossy compression, with a focus on image compression. These methods generally encode the source into an embedding and use a diffusion model to iteratively refine it during decoding, so that the reconstruction approximately follows the true data distribution. The embedding can take various forms and is typically transmitted via an auxiliary entropy model, and recent methods also explore the use of diffusion models themselves for information transmission via channel simulation. We review representative approaches through the lens of rate-distortion-perception theory, highlighting the role of common randomness and connections to inverse problems, and identify open challenges.

    https://arxiv.org/abs/2601.18932


    Resolution-Aliasing Trade-off in Near-Field Localisation

    oai:arXiv.org:2602.01947v2

    arXiv:2602.01947v2 Announce Type: replace Abstract: Extremely Large-scale MIMO (XL-MIMO) systems operating in Near-Field (NF) introduce new degrees of freedom for accurate source localisation, but make dense arrays impractical. Sparse or distributed arrays can reduce hardware complexity while maintaining high resolution, yet sub-Nyquist spatial sampling introduces aliasing artefacts in the localisation ambiguity function. This paper presents a unified framework to jointly characterise resolution and aliasing in NF localisation and study the trade-off between the two. Leveraging the concept of local chirp spatial frequency, we derive analytical expressions linking array geometry and sampling density to the spatial bandwidth of the received field. We introduce two geometric tools--Critical Antenna Elements (CAEs) and the Non-Contributive Zone (NCZ)--to intuitively identify how individual antennas contribute to resolution and/or aliasing. Our analysis reveals that resolution and aliasing are not always strictly coupled, e.g., increasing the array aperture can improve resolution without necessarily aggravating aliasing. These results provide practical guidelines for designing NF arrays that optimally balance resolution and aliasing, supporting efficient XL-MIMO deployment.

    https://arxiv.org/abs/2602.01947


    Structured Learning for Electromagnetic Field Modeling and Real-Time Inversion

    oai:arXiv.org:2602.06618v2

    arXiv:2602.06618v2 Announce Type: replace Abstract: Precise magnetic field modeling is fundamental to the closed-loop control of electromagnetic navigation systems (eMNS) and the analytical Multipole Expansion Model (MPEM) is the current standard. However, the MPEM relies on strict physical assumptions regarding source symmetry and isolation, and requires optimization-based calibration that is highly sensitive to initialization. These constraints limit its applicability to systems with complex or irregular coil geometries. This work introduces an alternative modeling paradigm based on multi-layer perceptrons that learns nonlinear magnetic mappings while strictly preserving the linear dependence on currents. As a result, the field models enable fast, closed-form minimum-norm inversion with evaluation times of approximately 1 ms, which is critical for high-bandwidth magnetic control. For model training and evaluation we use large-scale, high-density datasets collected from the research-grade OctoMag and clinical-grade Navion systems. Our results demonstrate that data-driven models achieve predictive fidelity equivalent to the MPEM while maintaining comparable data efficiency, and we further assess their suitability for real-time magnetic control in a closed-loop tracking experiment running at 100 Hz. Furthermore, we demonstrate that straightforward design choices effectively eliminate spurious workspace ill-conditioning frequently reported in MPEM-based calibration. To facilitate future research, we release the complete codebase and datasets open source.

    https://arxiv.org/abs/2602.06618


    Attack-Dependent Robustness of Neural Audio Codecs for Adversarial ASR

    oai:arXiv.org:2603.09034v2

    arXiv:2603.09034v2 Announce Type: replace Abstract: Neural audio codecs impose a discrete bottleneck through residual vector quantization (RVQ), making them a useful class of inference-time transformations for reducing adversarial perturbations before ASR inference. We study how codec quantization depth affects defended ASR under non-adaptive, standard adaptive, and quantization-aware adaptive untargeted $\ell_\infty$ attacks. Under non-adaptive attacks, intermediate RVQ depths yield the lowest word error rates and outperform traditional compression at comparable bitrates. However, this apparent optimum is not stable under adaptive evaluation. The standard identity-gradient adaptive baseline (BPDA+EOT) can overestimate robustness, while an implementation of an RVQ-relaxed adaptive attack (SoftVQ-PGD) substantially changes the observed depth trend and largely removes the intermediate-depth advantage. Overall, neural codecs can improve defended ASR under specific threat models. However, the relationship between robustness and RVQ depth depends on the attack used for evaluation, rather than on the codec architecture alone.

    https://arxiv.org/abs/2603.09034


    PHONOS: PHOnetic Neutralization for Online Streaming Applications

    oai:arXiv.org:2603.27001v2

    arXiv:2603.27001v2 Announce Type: replace Abstract: Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accent cues intact, which is problematic because such cues can reveal a speaker's first-language or geographic background and narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that performs accent neutralization in a privacy sense: reducing accent-origin cues by converting non-native segmental realizations toward a chosen target accent domain. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test accentedness ratings consistent with this shift. PHONOS also moves outputs away from the original speaker in embedding space, suggesting lower linkability under an embedding-based proxy, while running with $\leq241\,\mathrm{ms}$ end-to-end latency on a single GPU.

    https://arxiv.org/abs/2603.27001


    Salted Fisher Information for Hybrid Systems

    oai:arXiv.org:2603.29862v2

    arXiv:2603.29862v2 Announce Type: replace Abstract: Discrete events change how parameter-influence propagates in hybrid systems. Prevailing Fisher information for- mulations assume that sensitivities evolve smoothly according to continuous-time variational equations and therefore neglect the sensitivity updates induced by discrete events. This paper derives a Fisher information matrix formulation compatible with hybrid systems. To do so, we use the saltation matrix, which encodes the first-order transformation of sensitivities induced by discrete events. We call the resulting formulation the salted Fisher information matrix (SFIM). The proposed framework unifies continuous information accumulation during flows with discrete updates at event times. We also show that hybrid persistence of excitation is sufficient for the SFIM to be positive definite

    https://arxiv.org/abs/2603.29862


    Min-Max Grassmannian Optimization for Online Subspace Tracking

    oai:arXiv.org:2604.00825v2

    arXiv:2604.00825v2 Announce Type: replace Abstract: We propose GeRoST (Geometrically Robust Subspace Tracking), an online subspace tracking algorithm that models uncertainty in a subspace using a Grassmannian ball. We derive an exact scalar dual for the worst-case subspace problem, establish conditions for a unique worst-case subspace and a Riemannian gradient, and characterize the minimum radius needed to cover a dimensional extension of the target subspace. Each update uses either a spectral direction computed in a reduced subspace or the gradient of the window reconstruction loss. Our numerical experiments show that GeRoST achieves lower mean post-fault prediction error than GREAT in system identification. In video separation, it achieves higher precision and a better precision--recall balance, as measured by the F$_1$ score, than both GREAT and GRASTA at the reported thresholds, with lower recall and longer runtime.

    https://arxiv.org/abs/2604.00825


    Safe learning-based control via function-based uncertainty quantification

    oai:arXiv.org:2604.01173v2

    arXiv:2604.01173v2 Announce Type: replace Abstract: Uncertainty quantification is essential when deploying learning-based control methods in safety-critical systems. This is commonly realized by constructing uncertainty tubes that enclose the unknown function of interest, e.g., the reward and constraint functions or the underlying dynamics model, with high probability. However, existing approaches for uncertainty quantification typically rely on restrictive assumptions that encode smoothness properties of the unknown function, such as a known norm in a function space. Moreover, these methods usually struggle with discontinuities. In this paper, we model the unknown function as a random function from which independent and identically distributed realizations can be generated. We then construct uncertainty tubes via the scenario approach that hold with high probability. Our uncertainty tubes rely solely on sampled realizations and can therefore accommodate discontinuities represented by the sampling model. We integrate these uncertainty tubes into a safe Bayesian optimization algorithm with which we safely tune control parameters on a real Furuta pendulum.

    https://arxiv.org/abs/2604.01173


    How Sensor Attacks Transfer Across Lie Groups

    oai:arXiv.org:2604.03461v2

    arXiv:2604.03461v2 Announce Type: replace Abstract: Sensor spoofing analysis in cyber-physical systems is predominantly confined to linear state spaces, where an attack's persistence implies the existence of an integrator somewhere in the loop. We extend this insight to a commutativity condition on Lie groups, where noncommutative dynamics can distort sensor attacks, exposing nominally stealthy attacks by complex maneuvers. We present a geometric framework characterizing when a sensor attack can transfer across operating conditions, preserving both its physical impact and stealthiness. We prove that successful transfer requires the attack to commute with the flow (a Lie bracket condition), isolating transferable attacks to an invariant subspace. For small deviations from transferable attacks, our decomposition theorem reveals a fundamental asymmetry: the flow's Adjoint action distorts the physical impact of the bracket-violating component. Furthermore, even if the attack's impact isn't distorted, the subsequent residual could be. Finally, we demonstrate how turning maneuvers on a Dubins unicycle collapse the transferable subspace to a single direction, verifying that imperfect attacks remain within theoretical detection bounds.

    https://arxiv.org/abs/2604.03461


    Enhanced ShockBurst for Ultra Low-Power On-Demand Sensing

    oai:arXiv.org:2604.07188v2

    arXiv:2604.07188v2 Announce Type: replace Abstract: On-demand sensing requires battery-powered Internet-of-Things (IoT) and implantable medical devices to remain in deep sleep and activate wireless communication only when data transmission is required. In such systems, battery lifetime depends strongly on radio active time. This work investigates how communication architecture and physical layer (PHY) configuration influence radio active time by comparing connection-oriented Bluetooth Low Energy (BLE) with connectionless Enhanced ShockBurst (ESB) on identical BLE-compatible hardware. Under identical 2 Mbps PHY configurations, ESB reduces wake-up latency and energy consumption to approximately one-twentieth of BLE by eliminating connection establishment and maintenance overhead. Increasing the ESB PHY rate from 2 to 4 Mbps further shortens packet airtime by approximately 52% and reduces transmission energy by approximately 43%. Finally, a first-in, first-out (FIFO)-triggered implantable loop recorder prototype demonstrates that jointly optimizing communication architecture, PHY configuration, and buffered transmission enables sleep-wake operation and reduces total system power consumption by approximately 60% compared with conventional BLE operation. These results identify minimizing radio active time as a key design principle for ultra-low-power on-demand sensing and provide practical guidance for battery-powered sensing systems.

    https://arxiv.org/abs/2604.07188


    Receding Horizon Multi-Agent Deceptive Path Planner

    oai:arXiv.org:2605.14085v2

    arXiv:2605.14085v2 Announce Type: replace Abstract: Deceptive path planning enables autonomous agents to obscure their true goals from observers by deviating from an expected optimal path. Prior work largely solves full-horizon, end-to-end optimization for single agents, which is expensive to recompute online and difficult to scale or adapt en route. We propose a unified framework for deceptive path planning using a Boltzmann distribution, computing over short-horizon candidate trajectories within a receding-horizon loop. By param- By iterating a user-defined cost that captures deception, resources, and smoothness, and optionally includes coupling terms between agents, the framework yields stochastic policies that balance the tradeoff between optimal paths and deceptive deviation. Policies are updated locally and do not require training. The level of deception and adherence to constraints can be dynamically tuned, enabling online adaptation to changes in goals and constraints such as obstacles. This step-by-step tuning opens the door to new forms of dynamic deception. Simulation studies demonstrate the flexibility of our approach, maintaining deception while adapting to environmental and constraint updates, avoiding the recomputation required by full-horizon methods, and supporting intuitive tuning via a small set of parameters

    https://arxiv.org/abs/2605.14085


    STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

    oai:arXiv.org:2605.23137v3

    arXiv:2605.23137v3 Announce Type: replace Abstract: Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--language spaces, making direct cross-modal alignment unstable. To address this, we propose STAMBRIDGE, a versatile two-stage framework that sequentially tackles feature conditioning and cross-modal alignment. First, we introduce a Spectral-Temporal Amplitude-aware Modulation (STAM) to extract well-conditioned EEG representations. By replacing hard frequency masking with amplitude-derived soft channel weighting and multi-scale temporal convolutions, STAM explicitly preserves frequency-aware transients while reducing the risk of time-domain ringing artifacts. Building upon these robust neural features, we further introduce a model-agnostic Mid-Feature Semantic Bridge (MFSB) that constructs a regularized intermediate space through directed cross-modal interactions, enabling staged distillation and more stable semantic alignment. Experiments on the THINGS-EEG benchmark show competitive 200-way zero-shot retrieval performance, with 34.50\% Top-1 and 65.95\% Top-5 accuracy. In addition, embeddings learned by STAMBRIDGE produce semantically coherent image reconstructions with a diffusion model, demonstrating robust EEG-to-vision semantic alignment. The code is available at: https://github.com/thabeatmjh/STAMBRIDGE.

    https://arxiv.org/abs/2605.23137


    Ambiguity Analysis and Design of Sparse Arrays via Generalized Vandermonde Rank Conditions

    oai:arXiv.org:2606.00360v2

    arXiv:2606.00360v2 Announce Type: replace Abstract: Sparse linear arrays obtained by thinning a uniform linear array (ULA) achieve large effective apertures with a reduced number of physical sensors and have become a key enabling technology across radar, sonar, communications, and integrated sensing and communications. The price of thinning, however, is the emergence of ambiguities in the array manifold: distinct sets of directions of arrival that produce identical sensor measurements, precluding unique identification of multiple sources. Conventional sparse-array design criteria, based on beampattern shaping or estimation-performance optimization, do not fully capture how multiple steering vectors interact jointly to produce such ambiguities. This paper develops an algebraic framework for the complete multi-source identifiability analysis of thinned ULAs, characterizing all ambiguity sets for arbitrary source number rather than only at the maximum identifiable number. By relating the rank deficiency of the generalized Vandermonde matrix associated with the sparse steering matrix to that of a thinned Toeplitz matrix, and further to a rank condition on an augmented full-ULA steering matrix with prescribed generators, we obtain a systematic characterization of the ambiguity sets in large sparse arrays together with constructive design guidelines for ambiguity-free geometries. Algebraic and numerical examples demonstrate that the proposed framework resolves the ambiguity structure of thinned arrays of a few dozen grid positions, for which previous methods yield only individual ambiguity sets.

    https://arxiv.org/abs/2606.00360


    Diffusion Models for Radio Map Estimation: Theoretical Performance Analysis and Sampling Rate Guideline

    oai:arXiv.org:2606.25310v2

    arXiv:2606.25310v2 Announce Type: replace Abstract: Radio maps, which characterize the spatial distribution of radio frequency metrics, such as the received signal strength, are essential for a wide range of wireless applications. The problem of radio map estimation involves constructing a radio map from a limited set of radio samples measured by sparsely distributed sensors.Recently, diffusion models have been increasingly adopted for this problem, yet their theoretical performance remains largely unexamined. Consequently, it is difficult to evaluate their performance when ground-truth radio maps are unavailable, as is often the case in practice. To bridge this gap, we first formulate radio map estimation as a non-linear matrix completion problem. We then derive a theoretical expression for the estimation error of a diffusion model, capturing the effects of the mismatch between the training and deployment environments, the model design, and the sampling strategy on the quality of the estimated radio map. Moreover, considering that the derived error expression depends on certain information that is difficult to obtain in practice, we propose an empirical approximation that is readily computable from observable data. Finally, extensive simulations demonstrate that the empirical formula closely approximates the theoretical error expression, validating its effectiveness for practical deployment. Our results provide guidelines not only for evaluating the performance of a diffusion model when ground truth is unavailable, but also for determining how many sensors are required to spatially sample the region to achieve a given estimation accuracy. They further reveal a critical sensor count beyond which the estimation performance of diffusion models converges.

    https://arxiv.org/abs/2606.25310


    Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

    oai:arXiv.org:2606.25672v3

    arXiv:2606.25672v3 Announce Type: replace Abstract: Classifier-free guidance (CFG) is widely used in flow-matching-based zero-shot text-to-speech (TTS), where generation is conditioned on text content and a speech prompt. Standard CFG uses a single guidance weight for their joint conditional effect, while branch-selective guidance emphasizes text or speaker conditioning and can introduce a trade-off between text accuracy and speaker similarity. In this paper, we revisit CFG under independently masked conditions and decompose the guidance field into text, speaker, and joint residuals. We show that condition-specific branch differences couple the joint residual with the corresponding text or speaker residual under a shared weight. Trajectory analysis further shows that the joint residual varies over flow time and contains information that cannot be represented by reweighting the text and speaker residuals alone. Based on these observations, we propose joint residual reweighting, which assigns independent weights to the three residuals. Experiments on F5-TTS, CosyVoice2, and GLM-TTS across three evaluation sets show overall improvements in speaker similarity and text accuracy over the default CFG settings without retraining.

    https://arxiv.org/abs/2606.25672


    SCI-Mamba: Unsupervised Learning based Low-Light Image Enhancement for Non-Cooperative Spacecraft

    oai:arXiv.org:2607.08033v2

    arXiv:2607.08033v2 Announce Type: replace Abstract: Low-light visual perception acts as the core visual foundation for on-orbit servicing missions targeting non-cooperative spacecraft, supporting autonomous rendezvous, pose estimation, component detection and robotic capture operations. Spaceborne imagery suffers from severe low-light degradation, while the extreme scarcity of paired normal/low-light space samples severely limits the generalization capacity of supervised enhancement algorithms. To address this practical bottleneck, this paper proposes SCI-Mamba, an unsupervised enhancement network for low-light orbital spacecraft observations. The proposed framework unites self-calibrated unsupervised learning, linear-complexity VMamba architecture and Retinex physical priors, delivering a lightweight enhancement pipeline adaptable to resource-limited spaceborne hardware. We construct Space Dark-1.0, a dedicated low-light spacecraft dataset integrating real orbital footage, darkroom hardware-in-the-loop measurements and physically constrained synthetic data covering diverse illumination, motion and attitude conditions. Comprehensive comparisons with CNN-, Transformer- and prevailing Mamba-based approaches verify the advantages of SCI-Mamba in visual authenticity, color fidelity and inference speed. The proposed framework provides a practical low-light enhancement solution for close-proximity non-cooperative space operations. The code is available at https://github.com/bitswh/SCI-Mamba

    https://arxiv.org/abs/2607.08033


    A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification

    oai:arXiv.org:2608.14824v2

    arXiv:2608.14824v2 Announce Type: replace Abstract: We present a parameter-free episodic evaluation of nearest-centroid classification of elephant vocalisations on fixed pretrained embeddings, for the Elephant Voices (EV) and Linguistic Data Consortium (LDC) datasets. We ask not which embedding yields the best classifier trained on all labelled data, but how the simplest classifier performs as the number of exemplars per class varies. There are no learnable parameters, because each class is modelled as the mean of its support embeddings and each query is assigned to the nearest centroid under squared Euclidean distance. Evaluation covers the fixed Perch (ver. 1), Perch (ver. 2) and HuBERT (base, layer 2) embeddings, alongside mel frequency cepstral coefficient (MFCC) features, $N$-way $k$-shot, under the same stratified $K$-fold cross-validation protocol as the trained classifiers. None of these embedding models was trained to distinguish elephant call types. On the smaller EV dataset the centroid classifier is markedly data-efficient. Using Perch (ver. 1) or Perch (ver. 2) embeddings it overtakes in mean average precision (mAP) the fully-trained logistic regression (LR) baseline from one or two exemplars and the recurrent baseline from two. Over the reduced set of call types on which the strongly-supervised end-to-end baseline was trained, the centroid classifier using Perch (ver. 2) embeddings overtakes that baseline in mAP as well, from two exemplars. On the larger LDC dataset the recurrent baselines retain their advantage for all considered values of $k$. Only LR is overtaken, and only in mAP. Nearest-centroid classification is therefore preferable precisely when exemplars are few and the fixed embedding already separates the call types.

    https://arxiv.org/abs/2608.14824


    Coherent Direct D-MIMO Localization: Analysis of Coherence Levels

    oai:arXiv.org:2608.24880v2

    arXiv:2608.24880v2 Announce Type: replace Abstract: Distributed multiple-input multiple-output (D-MIMO) offers favorable geometry for localization and sensing, with its greatest potential arising from joint coherent processing across panels. However, stringent frequency-synchronization and phase-calibration requirements, together with multimodal likelihood functions, complicate the estimation problem. Consequently, most existing methods process the panels noncoherently, potentially sacrificing localization accuracy. We present a unified family of Bayesian state filters based on concentrated Type-I and marginal Type-II likelihoods for wideband near-field D-MIMO systems. The Type-I filters realize (i) noncoherent, (ii) coherent, and (iii) carrier-phase-based processing. We show that a zero-mean Type-II model is inherently noncoherent under distributed processing, whereas observation stacking restores coherence. A recently proposed nonzero-mean Type-II model adapts to the coherence available in the data, a property termed ``soft coherence.'' We derive posterior Cram\'er--Rao lower bounds (PCRLBs) for all three coherence levels and show that each level is fundamentally tied to the number of phase parameters used for positioning or treated as nuisance parameters. Numerical results show that the filters closely approach their respective PCRLBs and that coherent processing substantially outperforms noncoherent processing. Our particle-based implementations parallelize over particles and panels, scale linearly with the observed data, and achieve runtimes of tens of milliseconds per time step under GPU acceleration.

    https://arxiv.org/abs/2608.24880


    A Foundation Model for Large-Scale Wireless Network Planning , Operation and Optimization

    oai:arXiv.org:2609.08482v3

    arXiv:2609.08482v3 Announce Type: replace Abstract: Wireless cellular networks form the connective tissue of human society, sustained by a continuous physical dialogue between engineered infrastructure and its surroundings. Radio signals emitted from base stations traverse terrain, diffract around buildings and scatter through streets before reaching billions of users. Together, these interactions produce the city-wide radio environment on which every network decision rests. Shaping this environment through deployment and optimization determines the connectivity societies rely on, yet learning it effectively at city scale and generalizing across diverse cities and deployments remain open challenges. Here we answer positively by introducing ChaRT, a foundation model that learns transferable radio representations from measurement reports generated by operational networks. These reports provide abundant multi-cell, multi-beam observations without dedicated campaigns, forming a scalable data foundation for city-scale learning. ChaRT embeds beam-level angular structure, network hierarchy and propagation-regime diversity in its architecture, and is pretrained through context-aware masked beam modelling and self-distillation with channel-model-constrained augmentation. We pretrain ChaRT on over one billion reports comprising 18.2 billion beam-level observations from 3,503 cells in one city. With a single set of weights, ChaRT reconstructs radio environments in unseen cities and transfers to radio map construction, new-site prediction and network parameter tuning. With only 1% of labelled data, it supports user localization, beam prediction, propagation scenario classification and estimation of the signal-to-interference-plus-noise ratio. The learned representation further enables beamspace clustering for reusable radio-grid construction. These results establish ChaRT as a transferable foundation for network-wide intelligence.

    https://arxiv.org/abs/2609.08482


    Source-Adaptive Data Curation for Bilingual NVV-Aware ASR

    oai:arXiv.org:2609.09929v2

    arXiv:2609.09929v2 Announce Type: replace Abstract: Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) systems often discard. We present a bilingual Mandarin-English system for Track 1 of the NVVSpeech Challenge at ISCSLP 2026, which requires joint transcription of lexical content and 16 NVV categories at their transcript-relative positions. Our NVV-Aware Whisper adapts Whisper-medium through checkpoint-compatible vocabulary remapping, enabling lexical tokens and inline NVV tags to be decoded within a unified autoregressive sequence without expanding the vocabulary. To provide reliable and diverse supervision, we further introduce a source-adaptive data curation strategy that refines public NVV corpora through acoustic augmentation and multimodal LLM filtering, while mining spontaneous NVVs from in-the-wild media through automated preprocessing and annotation. Under the official bilingual evaluation protocol, the proposed system improves final score from 33.32 to 53.61, with ablations confirming the complementary benefits of the proposed data-curation components.

    https://arxiv.org/abs/2609.09929


    NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding

    oai:arXiv.org:2609.09940v2

    arXiv:2609.09940v2 Announce Type: replace Abstract: Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.

    https://arxiv.org/abs/2609.09940


    Diarization Error Decomposition Under Pause Annotation Ambiguity

    oai:arXiv.org:2609.11007v2

    arXiv:2609.11007v2 Announce Type: replace Abstract: Speaker diarization evaluation is sensitive to ambiguity in pause annotation, which can inflate diarization error rate (DER) or obscure genuine model errors. We show that morphological closing, which has been used for pause-tolerant diarization evaluation, discards segment-level distinctions. Instead, we propose an exact, overlap-aware decomposition of standard DER into a pause-attributable component, consisting of errors compatible with pause filling, and a residual core component that can serve as a proxy for intrinsic diarization errors. The decomposition leaves DER unchanged, while the pause-attributable and core components vary monotonically with the pause threshold and eventually saturate. Experiments spanning synthetic transformations, annotation mismatch, cross-domain evaluation, and tight-boundary diarization show that the decomposition reveals error sources not apparent from standard DER.

    https://arxiv.org/abs/2609.11007


    Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

    oai:arXiv.org:2609.11260v2

    arXiv:2609.11260v2 Announce Type: replace Abstract: We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressive speech models. To construct preference data without human preference annotation, we build a bilingual prompt corpus by combining NVV-injected real transcripts with LLM-generated semantically aligned prompts, perform stochastic model rollouts, and use a LALM to rank candidate utterances and form same-prompt chosen--rejected pairs. We then adopt a two-stage optimization strategy: Rejection Sampling Fine-Tuning (RSFT) first adapts the model to LALM-selected high-scoring samples, followed by Anchored Flow-DPO, which formulates pairwise preference optimization using utterance-level flow-matching loss and retains the chosen-sample flow-matching objective as an SFT anchor. This design enables DPO-style preference learning without explicit sequence likelihoods while preserving direct supervision on preferred realizations. On the official 1,600-utterance NVVSpeech Challenge Track~2 test set, our method achieves a Final Track2Score of \textbf{75.80} (79.39 ZH / 72.21 EN), outperforming the VoxCPM2 baseline by \textbf{+1.84}. The improvements are mainly driven by higher NVV Accuracy and NVV Perceptual Effect, while Overall Quality remains stable.

    https://arxiv.org/abs/2609.11260


    UQMIA: An Open, Hands-On Tutorial on Uncertainty Quantification in Medical Imaging Analysis with Large Language Model-Based Assessment of Educational Content

    oai:arXiv.org:2609.23241v2

    arXiv:2609.23241v2 Announce Type: replace Abstract: Uncertainty quantification (UQ) is increasingly recognized as an important component of reliable machine learning in medical imaging, yet practical resources connecting UQ theory, implementation, and evaluation remain limited. We developed Uncertainty Quantification in Medical Imaging Analysis (UQMIA), an open-access, hands-on tutorial comprising 19 sessions covering major UQ approaches, including variational inference, Monte Carlo dropout, deep ensembles, evidential deep learning, and conformal prediction, together with methods for evaluating uncertainty reliability. The tutorial is designed for researchers and practitioners with experience in medical imaging, progressing from foundational concepts to implementation and evaluation, with notebooks executable through Kaggle. Beyond presenting the tutorial, we introduce a framework using large language models (LLMs) to evaluate whether technical educational resources contain retrievable and usable knowledge. Using 100 four-option multiple-choice questions derived from primary methodological literature, we evaluated 20 instruction-tuned LLMs from six model families with and without retrieved tutorial context. Retrieval of UQMIA improved accuracy in 18 of 20 models, increasing mean accuracy from 0.680 to 0.742 (+0.062; Holm-adjusted P=0.00032), and improved area under the receiver operating characteristic curve (AUROC) in 18 of 20 models, increasing from 0.725 to 0.788 (+0.063; Holm-adjusted P=0.00032). UQMIA provides an accessible resource for learning UQ in medical imaging and demonstrates an LLM-based approach for evaluating technical educational material. The tutorial is available at https://benyamin-gheiji.github.io/Uncertainty-Quantification-Medical-Imaging-Analysis/ and its source code and notebooks at https://github.com/benyamin-gheiji/Uncertainty-Quantification-Medical-Imaging-Analysis/.

    https://arxiv.org/abs/2609.23241


    Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge

    oai:arXiv.org:2609.23462v2

    arXiv:2609.23462v2 Announce Type: replace Abstract: Non-verbal vocalizations (NVVs) carry important paralinguistic information but are often omitted by conventional automatic speech recognition (ASR) systems. The ISCSLP NVVSpeech Challenge requires joint transcription of lexical content and 16 NVV categories under limited and highly imbalanced supervision. We present a data-centric NVV-aware ASR pipeline based on cross-dataset label harmonization and a two-stage sampling schedule. We map heterogeneous source labels to the official taxonomy and exclude samples without a reliable mapping. Our schedule first uses square-root category sampling to moderate the long-tailed distribution and then applies uniform-category fine-tuning. On a fixed local validation split, square-root category sampling performs best among the tested single-stage settings. The final two-stage system obtains an official score of 63.86 and ranks fourth in Track 1.

    https://arxiv.org/abs/2609.23462


    Robust Strictly Positive Real Synthesis for Sixth-Order Interval Polynomial Families

    oai:arXiv.org:2609.26541v2

    arXiv:2609.26541v2 Announce Type: replace Abstract: Every Hurwitz-stable interval family of monic real polynomials of degree six admits a single real numerator of degree six that makes all the associated transfer functions strictly positive real. We give a constructive proof. The complete existence theorem has been formalized in Lean4.

    https://arxiv.org/abs/2609.26541


    Radiomics and artificial Intelligence for thyroid cancer diagnosis: Concepts, challenges, and solutions

    oai:arXiv.org:2404.07239v2

    arXiv:2404.07239v2 Announce Type: replace-cross Abstract: Thyroid cancer is an increasing global health concern that requires advanced diagnostic methods. The application of AI and radiomics to thyroid cancer diagnosis is examined in this review. A review of multiple databases was conducted in compliance with PRISMA guidelines until October 2024. A combination of keywords led to the discovery of an English academic publication on thyroid cancer and related subjects. 368 papers were returned from the original search after 112 duplicates were removed. Relevant studies were selected according to predetermined criteria after 176 articles were eliminated based on an examination of their abstract and title. After the comprehensive analysis, an additional six studies were excluded. Among the 42 included studies, radiomics analysis, which incorporates ultrasound (US) images, demonstrated its effectiveness in diagnosing thyroid cancer. Various results were noted, some of the studies presenting new strategies that outperformed the status quo. The literature has emphasized various challenges faced by AI models, including interpretability issues, dataset constraints, and operator dependence. The synthesized findings of the 42 included studies mentioned the need for standardization efforts and prospective multicenter studies to address these concerns. Furthermore, approaches to overcome these obstacles were identified, such as advances in explainable AI technology and personalized medicine techniques. The review focuses on how AI and radiomics could transform the diagnosis and treatment of thyroid cancer. Despite challenges, future research on multidisciplinary cooperation, clinical applicability validation, and algorithm improvement holds the potential to improve patient outcomes and diagnostic precision in the treatment of thyroid cancer.

    https://arxiv.org/abs/2404.07239


    LiSeCo: Linear Semantic Control for Language Generation

    oai:arXiv.org:2405.15454v5

    arXiv:2405.15454v5 Announce Type: replace-cross Abstract: The prevalence of Large Language Models (LLMs) in critical applications highlights the need for controlled language generation methods that are both computationally efficient and enjoy performance guarantees. To address this need, we use a common model of concept semantics as linearly represented in an LLM's latent space. In particular, we take the view that natural language generation traces a trajectory in this continuous semantic space, realized by the language model's hidden activations. This view permits a control-theoretic treatment of text generation in latent space, in which we propose Linear Semantic Control (LiSeCo), a lightweight, gradient-free intervention that dynamically steers trajectories away from regions corresponding to undesired meanings. In particular, we propose to directly intervene, in an online fashion, the activations of the token that is being generated in embedding space. Crucially, LiSeCo does not simply steer activations towards a desirable region. Instead, it relies on classical techniques from control theory to precisely control activations in a context-dependent way, and guarantees that they are brought into a specific pre-defined region of embedding space that corresponds to allowed semantics. The intervention is computed in closed form according to an optimal controller formulation, minimally impacting generation time. This control of the activations in embedding space allows for fine-grained steering of attributes of the generated sequence. We demonstrate that our approach is effective on different tasks -- toxicity, sentiment, and language (English/Spanish) steering -- while maintaining text quality.

    https://arxiv.org/abs/2405.15454


    Retrieval Augmented (Knowledge Graph), and Large Language Model-Driven Design Structure Matrix (DSM) Generation of Cyber-Physical Systems

    oai:arXiv.org:2602.16715v2

    arXiv:2602.16715v2 Announce Type: replace-cross Abstract: We explore the potential of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and Graph-based RAG (GraphRAG) for generating Design Structure Matrices (DSMs). We test these methods on two distinct use cases--a power screwdriver and a CubeSat with known architectural references--evaluating their performance on two key tasks: determining relationships between predefined components, and the more complex challenge of identifying components and their subsequent relationships. We measure the performance by assessing each element of the DSM and overall architecture. Despite design and computational challenges, we identify opportunities for automated DSM generation, with all code publicly available for reproducibility and further feedback from the domain experts.

    https://arxiv.org/abs/2602.16715


    Linearly Solvable Continuous-Time General-Sum Stochastic Differential Games

    oai:arXiv.org:2604.07479v2

    arXiv:2604.07479v2 Announce Type: replace-cross Abstract: This paper introduces a class of continuous-time, finite-player stochastic general-sum differential games that admit solutions through an exact linear PDE system. We formulate a distribution planning game utilizing the cross-log-likelihood ratio to naturally model multi-agent spatial conflicts, such as congestion avoidance. By applying a generalized multivariate Cole-Hopf transformation, we decouple the associated non-linear Hamilton-Jacobi-Bellman (HJB) equations into a system of linear partial differential equations. This reduction enables the efficient, grid-free computation of feedback Nash equilibrium strategies via the Feynman-Kac path integral method, effectively overcoming the curse of dimensionality.

    https://arxiv.org/abs/2604.07479


    Functional Connectivity-Guided Band Selection for Motor Imagery Brain-Computer Interfaces

    oai:arXiv.org:2605.00746v2

    arXiv:2605.00746v2 Announce Type: replace-cross Abstract: Reliable control in motor imagery brain-computer interfaces (MI-BCIs) requires the precise decoding of user-specific neural rhythms, which vary significantly across individuals. The Common Spatial Pattern (CSP) algorithm is a cornerstone of MI-BCI decoding, yet its performance depends strongly on the spectral range of the input EEG data. Although Filter Bank CSP (FBCSP) extends this as a data-driven decoding framework, its frequency sub-bands are predefined rather than selected using subject-specific physiological criteria. This paper presents a proof-of-concept study of static functional connectivity (FC)-guided band selection for MI-BCI, demonstrated using a conventional FBCSP-based pipeline. The proposed method identifies the most discriminative spectral bands by calculating phase-based connectivity across four sensorimotor channels using wPLI, PLV, and PLI. Nine bands in a 4-40 Hz filter bank are ranked by the effect size of their hemispheric coupling differences and pruned to the top K bands for feature extraction and classification via FBCSP and a Support Vector Regressor. This framework was tested for K values ranging from 1 to 8 across the BCI Competition IV-2a (n = 9) and OpenBMI (n = 54) datasets. Performance was benchmarked against standard nine-band FBCSP and random ablation to determine the minimum number of bands (K*) required to maintain accuracy within a 2% baseline equivalence zone. Results show FC-guided selection can outperform random ablation and achieve near-baseline performance while reducing required CSP fits by 22.2% to 77.8%. PLV enables the most aggressive dimensionality reduction by prioritizing the {\mu} and low-\b{eta} ranges, while wPLI demonstrates superior inter-session robustness by mitigating volume conduction. These findings establish FC-guided selection as a principled and interpretable alternative to heuristic filter bank designs.

    https://arxiv.org/abs/2605.00746


    TactileReflex: Noise-Statistics-Driven Vision-Tactile Reflex Control for Force-Sensitive Manipulation

    oai:arXiv.org:2605.23568v4

    arXiv:2605.23568v4 Announce Type: replace-cross Abstract: Manipulating fragile deformable containers, such as disposable plastic cups filled with liquid, demands real-time grip-force adaptation within an extremely narrow force margin: insufficient force causes slip, while excessive force irreversibly deforms the thin wall. Existing approaches struggle to achieve such force-sensitive manipulation tasks. We propose a noise-statistics-based calibration-driven reflex control paradigm with vision-based tactile sensing: by analyzing the sensor's intrinsic noise characteristics (via a brief static-hold-and-unload protocol), we directly derive all controller thresholds, eliminating external force calibration, trial-and-error manual tuning, or material-specific physical models. Instantiating this paradigm, we present TactileReflex, a three-channel closed-loop controller that extracts three image-level proxies, shear intensity ($S_y$), contact intensity ($F_n$), and center of pressure ($C$), from dual visuo-tactile sensors and drives prioritized reflex channels at ~12 Hz for slip suppression, weight-adaptive release, and force protection. Each channel closes the loop directly on its proxy via noise-derived thresholds. Ablation demonstrates that only the full three-channel system is able to prevent irreversible container deformation (5/5 success vs. at most 1/5 for partial configurations). In a dynamic pouring task, fixed-effort baselines fail in all 10 attempts due to pose drift, while TactileReflex achieves 9/10 success across two water volumes. As a self-contained and interpretable controller, TactileReflex can serve as a plug-and-play safety layer beneath high-level manipulation pipelines, including haptic-free VR teleoperation and vision-language-action (VLA) policies.

    https://arxiv.org/abs/2605.23568


    Microwave Linear Analog Computer (MiLAC) for Simultaneous Active and Passive Beamforming

    oai:arXiv.org:2605.31549v2

    arXiv:2605.31549v2 Announce Type: replace-cross Abstract: Microwave linear analog computers (MiLACs) have recently emerged to enable high-performance and efficient beamforming in the analog domain. In this paper, we introduce a dual-functionality framework for MiLAC-aided transceivers. Beyond analog-domain precoding/combining (active beamforming), a MiLAC and its antenna array can simultaneously act as a reconfigurable intelligent surface (RIS) (passive beamforming). This allows the MiLAC to execute beamforming for transmission/reception while reflecting external incident signals. We provide an optimal reconfiguration strategy for this dual-functional MiLAC, and characterize the fundamental limits on the trade-off between active and passive rate, namely the capacity region bounds and the sum-rate capacity.

    https://arxiv.org/abs/2605.31549


    Repeated Binary Direct Collinear Impacts Under Incremental Contact Laws With Permanent Indentation: A Hybrid Systems Formulation

    oai:arXiv.org:2609.06138v2

    arXiv:2609.06138v2 Announce Type: replace-cross Abstract: Incremental contact laws specify the normal contact force through a differential equation carrying an internal state, driven by the indentation and its rate. In some, the force is extinguished at a nonzero indentation, whether by plastic deformation or by an elastic aftereffect, so that a residual deformation remains at the separation. Such laws sit uneasily within rigid body dynamics, which admits no deformation. The tension is tolerable when the indentation is small relative to the bodies, so that it may be carried constitutively rather than geometrically. Even then, the contact law alone does not determine the interaction of the bodies. Because force and indentation no longer vanish together, conditions for the commencement and termination of contact must be supplied separately. So must the fate of the deformation and internal state at separation, neither of which the equations of motion contain. This article formulates the repeated direct collinear impact of two convex bodies under external forces as a hybrid dynamical system. The contact interface is modeled as a massless element carrying the contact law and its state, coupled to the bodies through relative velocity and an interaction force dictated by the contact state. Consequently, all switching and resets are confined to the interface, leaving the geometry and the equations of motion of the bodies unaltered. The principal analytical properties of the resulting formulations are established, among them passivity and completeness; the branching of solutions at the onset and termination of contact is also examined. The framework is demonstrated through numerical simulations.

    https://arxiv.org/abs/2609.06138


    Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

    oai:arXiv.org:2609.10181v2

    arXiv:2609.10181v2 Announce Type: replace-cross Abstract: AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. Nonetheless, operational networks usually span many devices and administrative domains. Realizing an operator's intent requires coordinating agents with distinct authority scopes that define the resources they can access, the operations they can invoke, and the network state they can observe. This division limits the blast radius of an erroneous action but fragments the evidence needed to assess the network-wide outcome. Successful execution of a configuration action proposed by one agent does not establish that remote devices responded as intended or that routing changes reached the required devices. A valid observation may also become stale after a subsequent change. Before the coordinated operation can be declared complete, a trusted assurance layer must collect current observations from the required scopes and determine whether they collectively support the operator's intended network-wide outcome. To address the completion admission problem, we present EvidenceNet, a runtime assurance layer for deciding whether coordinated agent operations have achieved an operator's network intent. Its broker collects the post-change observations required by a completion contract, and its admission gate checks that the evidence comes from the required scopes, remains current, and satisfies the task rules. A verifier agent provides an additional assessment of the observation content. Experiments on live routing networks show that post-change state checks recognize successful outcomes that configuration-action records alone cannot establish. Controlled interventions further show that EvidenceNet rejects completion when otherwise satisfactory observations have the wrong source, have been substituted, or are stale.

    https://arxiv.org/abs/2609.10181


    VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

    oai:arXiv.org:2609.12831v2

    arXiv:2609.12831v2 Announce Type: replace-cross Abstract: As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe sets. In this paper, we propose \emph{VertexCBF}, a framework for learning neural CBFs in a scalable, systematic, and explainable way. We approximate the stationary Hamilton--Jacobi value function using a neural network trained via a combination of physics-informed and sparsely supervised learning. By exploiting control-affine dynamics and a convex polytope control set, under which the Hamiltonian is maximized at the control vertices, we efficiently generate supervision points via GPU-parallel vertex-restricted tree search, while a residual architecture guarantees that the learned CBF is never larger than the specified constraint function. We evaluate the method on 15 systems and compare it against relevant baselines, showing that it reliably recovers large safe sets where the baselines are conservative or fail completely. In addition, we perform a hardware experiment in which a mobile robot safely avoids pedestrians using a neural CBF trained with our method.

    https://arxiv.org/abs/2609.12831


    Realtime-Venus: A full-duplex interaction system with asynchronous delegation

    oai:arXiv.org:2609.13814v3

    arXiv:2609.13814v3 Announce Type: replace-cross Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.

    https://arxiv.org/abs/2609.13814