-
SMART: Self-learning Meta-strategy Agent for Reasoning Tasks
Authors:
Rongxing Liu,
Kumar Shridhar,
Manish Prajapat,
Patrick Xia,
Mrinmaya Sachan
Abstract:
Tasks requiring deductive reasoning, especially those involving multiple steps, often demand adaptive strategies such as intermediate generation of rationales or programs, as no single approach is universally optimal. While Language Models (LMs) can enhance their outputs through iterative self-refinement and strategy adjustments, they frequently fail to apply the most effective strategy in their f…
▽ More
Tasks requiring deductive reasoning, especially those involving multiple steps, often demand adaptive strategies such as intermediate generation of rationales or programs, as no single approach is universally optimal. While Language Models (LMs) can enhance their outputs through iterative self-refinement and strategy adjustments, they frequently fail to apply the most effective strategy in their first attempt. This inefficiency raises the question: Can LMs learn to select the optimal strategy in the first attempt, without a need for refinement? To address this challenge, we introduce SMART (Self-learning Meta-strategy Agent for Reasoning Tasks), a novel framework that enables LMs to autonomously learn and select the most effective strategies for various reasoning tasks. We model the strategy selection process as a Markov Decision Process and leverage reinforcement learning-driven continuous self-improvement to allow the model to find the suitable strategy to solve a given task. Unlike traditional self-refinement methods that rely on multiple inference passes or external feedback, SMART allows an LM to internalize the outcomes of its own reasoning processes and adjust its strategy accordingly, aiming for correct solutions on the first attempt. Our experiments across various reasoning datasets and with different model architectures demonstrate that SMART significantly enhances the ability of models to choose optimal strategies without external guidance (+15 points on the GSM8K dataset). By achieving higher accuracy with a single inference pass, SMART not only improves performance but also reduces computational costs for refinement-based strategies, paving the way for more efficient and intelligent reasoning in LMs.
△ Less
Submitted 21 October, 2024;
originally announced October 2024.
-
Dirty-Waters: Detecting Software Supply Chain Smells
Authors:
Raphina Liu,
Sofia Bobadilla,
Benoit Baudry,
Martin Monperrus
Abstract:
Using open-source dependencies is essential in modern software development. However, this practice implies significant trust in third-party code, while there is little support for developers to assess this trust. As a consequence, attacks have been increasingly occurring through third-party dependencies. These are called software supply chain attacks. In this paper, we target the problem of projec…
▽ More
Using open-source dependencies is essential in modern software development. However, this practice implies significant trust in third-party code, while there is little support for developers to assess this trust. As a consequence, attacks have been increasingly occurring through third-party dependencies. These are called software supply chain attacks. In this paper, we target the problem of projects that use dependencies while unaware of the potential risks posed by their software supply chain. We define the novel concept of software supply chain smell and present Dirty-Waters, a novel tool for detecting software supply chain smells. We evaluate Dirty-Waters on three JavaScript projects across nine versions and demonstrate the prevalence of all proposed software supply chain smells. Not only are there smells in all projects, but there are many of them, which immediately reveal potential risks and provide clear indicators for developers to act on the security of their supply chain.
△ Less
Submitted 21 October, 2024;
originally announced October 2024.
-
A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning
Authors:
Shengjie Sun,
Runze Liu,
Jiafei Lyu,
Jing-Wen Yang,
Liangpeng Zhang,
Xiu Li
Abstract:
Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often involves human intervention, numerous LLM queries, or repetitive RL training. To address these issues, we propose CARD, a LLM-driven Reward Design framework that iteratively generates and improves reward function code.…
▽ More
Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often involves human intervention, numerous LLM queries, or repetitive RL training. To address these issues, we propose CARD, a LLM-driven Reward Design framework that iteratively generates and improves reward function code. Specifically, CARD includes a Coder that generates and verifies the code, while a Evaluator provides dynamic feedback to guide the Coder in improving the code, eliminating the need for human feedback. In addition to process feedback and trajectory feedback, we introduce Trajectory Preference Evaluation (TPE), which evaluates the current reward function based on trajectory preferences. If the code fails the TPE, the Evaluator provides preference feedback, avoiding RL training at every iteration and making the reward function better aligned with the task objective. Empirical results on Meta-World and ManiSkill2 demonstrate that our method achieves an effective balance between task performance and token efficiency, outperforming or matching the baselines across all tasks. On 10 out of 12 tasks, CARD shows better or comparable performance to policies trained with expert-designed rewards, and our method even surpasses the oracle on 3 tasks.
△ Less
Submitted 18 October, 2024;
originally announced October 2024.
-
Joint Space-Time Adaptive Processing and Beamforming Design for Cell-Free ISAC Systems
Authors:
Rang Liu,
Ming Li,
Qian Liu
Abstract:
In this paper, we explore cooperative sensing and communication within cell-free integrated sensing and communication (ISAC) systems. Specifically, multiple transmit access points (APs) collaboratively serve multiple communication users while simultaneously illuminating a potential target, with a separate sensing AP dedicated to collecting echo signals for target detection. To improve the performa…
▽ More
In this paper, we explore cooperative sensing and communication within cell-free integrated sensing and communication (ISAC) systems. Specifically, multiple transmit access points (APs) collaboratively serve multiple communication users while simultaneously illuminating a potential target, with a separate sensing AP dedicated to collecting echo signals for target detection. To improve the performance of identifying a moving target in the presence of strong interference originating from transmit APs, we employ the space-time adaptive processing (STAP) technique and jointly optimize the transmit/receive beamforming. Our goal is to maximize the radar output signal-to-interference-plus-noise ratio (SINR), subject to the communication SINR requirements and the power budget. An efficient alternating algorithm is developed to solve the resulting non-convex optimization problem. Simulations demonstrate significant performance improvements in target detection and validate the advantages of the proposed joint STAP and beamforming design for cell-free ISAC systems.
△ Less
Submitted 18 October, 2024;
originally announced October 2024.
-
Vision-Language Navigation with Energy-Based Policy
Authors:
Rui Liu,
Wenguan Wang,
Yi Yang
Abstract:
Vision-language navigation (VLN) requires an agent to execute actions following human instructions. Existing VLN models are optimized through expert demonstrations by supervised behavioural cloning or incorporating manual reward engineering. While straightforward, these efforts overlook the accumulation of errors in the Markov decision process, and struggle to match the distribution of the expert…
▽ More
Vision-language navigation (VLN) requires an agent to execute actions following human instructions. Existing VLN models are optimized through expert demonstrations by supervised behavioural cloning or incorporating manual reward engineering. While straightforward, these efforts overlook the accumulation of errors in the Markov decision process, and struggle to match the distribution of the expert policy. Going beyond this, we propose an Energy-based Navigation Policy (ENP) to model the joint state-action distribution using an energy-based model. At each step, low energy values correspond to the state-action pairs that the expert is most likely to perform, and vice versa. Theoretically, the optimization objective is equivalent to minimizing the forward divergence between the occupancy measure of the expert and ours. Consequently, ENP learns to globally align with the expert policy by maximizing the likelihood of the actions and modeling the dynamics of the navigation states in a collaborative manner. With a variety of VLN architectures, ENP achieves promising performances on R2R, REVERIE, RxR, and R2R-CE, unleashing the power of existing VLN models.
△ Less
Submitted 18 October, 2024;
originally announced October 2024.
-
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
Authors:
Shuwei He,
Rui Liu,
Haizhou Li
Abstract:
Visual Text-to-Speech (VTTS) aims to take the spatial environmental image as the prompt to synthesize the reverberation speech for the spoken content. Previous research focused on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address the issues, we propose a novel multi-s…
▽ More
Visual Text-to-Speech (VTTS) aims to take the spatial environmental image as the prompt to synthesize the reverberation speech for the spoken content. Previous research focused on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address the issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS$^2$KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and semantic captions from image understanding LLM as supplementary sources. Afterwards, we propose a serial interaction mechanism to deeply engage with both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on their contributions.This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive spatial speech experience.Experimental results demonstrate that the MS$^2$KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https://github.com/MS2KU-VTTS/MS2KU-VTTS.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
Authors:
Renpu Liu,
Ruida Zhou,
Cong Shen,
Jing Yang
Abstract:
An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating based on the contextual information provided by the corresponding input-output demonstration pairs. It has been theoretically proved that ICL is enabled by the capability of Transformers to perform gradient-descent algor…
▽ More
An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating based on the contextual information provided by the corresponding input-output demonstration pairs. It has been theoretically proved that ICL is enabled by the capability of Transformers to perform gradient-descent algorithms (Von Oswald et al., 2023a; Bai et al., 2024). This work takes a step further and shows that Transformers can perform learning-to-optimize (L2O) algorithms. Specifically, for the ICL sparse recovery (formulated as LASSO) tasks, we show that a K-layer Transformer can perform an L2O algorithm with a provable convergence rate linear in K. This provides a new perspective explaining the superior ICL capability of Transformers, even with only a few layers, which cannot be achieved by the standard gradient-descent algorithms. Moreover, unlike the conventional L2O algorithms that require the measurement matrix involved in training to match that in testing, the trained Transformer is able to solve sparse recovery problems generated with different measurement matrices. Besides, Transformers as an L2O algorithm can leverage structural information embedded in the training tasks to accelerate its convergence during ICL, and generalize across different lengths of demonstration pairs, where conventional L2O algorithms typically struggle or fail. Such theoretical findings are supported by our experimental results.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
Differentiable Robot Rendering
Authors:
Ruoshi Liu,
Alper Canberk,
Shuran Song,
Carl Vondrick
Abstract:
Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap between visual data and action data. We introduce differentiable robot rendering, a method allowing the visual appearance of a robot body to be directly differentiable with respect to…
▽ More
Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap between visual data and action data. We introduce differentiable robot rendering, a method allowing the visual appearance of a robot body to be directly differentiable with respect to its control parameters. Our model integrates a kinematics-aware deformable model and Gaussians Splatting and is compatible with any robot form factors and degrees of freedom. We demonstrate its capability and usage in applications including reconstruction of robot poses from images and controlling robots through vision language models. Quantitative and qualitative results show that our differentiable rendering model provides effective gradients for robotic control directly from pixels, setting the foundation for the future applications of vision foundation models in robotics.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
High-temperature ferromagnetism and ferroelasticity in ultraflexible atomically thin square-shaped lattices
Authors:
Xinyuan Huang,
Yueqiao Qu,
Yu Liao,
Qian Zheng,
Ran Liu,
Yu Chen,
Liang Liu,
Junzhong Wang,
Gang Yao
Abstract:
The coexistence of high-temperature intrinsic ferromagnetic ordering, large magnetic anisotropy, along with novel mechanical properties such as ferroelasticity and flexibility, in experimental feasible two-dimensional (2D) crystals is greatly appealing for nanoscale spintronics. However, the progress in identifying such materials is limited. Here, by first-principles calculations, we report the fi…
▽ More
The coexistence of high-temperature intrinsic ferromagnetic ordering, large magnetic anisotropy, along with novel mechanical properties such as ferroelasticity and flexibility, in experimental feasible two-dimensional (2D) crystals is greatly appealing for nanoscale spintronics. However, the progress in identifying such materials is limited. Here, by first-principles calculations, we report the findings of an extraordinary combination of the above qualities for the first time in a new 2D exfoliated FeSi nanosheet in the P4/nmm space group. Due to the strong anion-mediated superexchange interaction, the monolayer FeSi (ML-FeSi) exhibits a Curie temperature Tc as high as 830 K, surpassing the current experimental record (344 K for ML-Cr3Te4). Furthermore, including FeSi, such isostructural lattices all demonstrate exceptional softness, as evidenced by their ultra-low in-plane stiffness. Remarkably, the transition metal atom and square-shaped crystal form work together to give this family of ML materials unique properties that can transition from Ising-like 2D ferromagnets in FeSi, MnP, MnAs, CrP, FeI, and VAs to 2D-XY ones in CrAs, VP, and multiferroic MnGe and TiTe. Overall, our work highlights such 2D lattices as promising candidates in emerging multifunctional device applications and nontrivial topological spintronics.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
Cryogenic Digital Image Correlation as a Probe of Strain in Iron-Based Superconductors
Authors:
Ziye Mo,
Chunyi Li,
Wenting Zhang,
Chang Liu,
Yongxin Sun,
Ruixian Liu,
Xingye Lu
Abstract:
Uniaxial strain is a powerful tuning parameter that can control symmetry and anisotropic electronic properties in iron-based superconductors. However, accurately characterizing anisotropic strain can be challenging and complex. Here, we utilize a cryogenic optical system equipped with a high-spatial-resolution microscope to characterize surface strains in iron-based superconductors using the digit…
▽ More
Uniaxial strain is a powerful tuning parameter that can control symmetry and anisotropic electronic properties in iron-based superconductors. However, accurately characterizing anisotropic strain can be challenging and complex. Here, we utilize a cryogenic optical system equipped with a high-spatial-resolution microscope to characterize surface strains in iron-based superconductors using the digital image correlation method. Compared with other methods such as high-resolution X-ray diffraction, strain gauge, and capacitive sensor, digital image correlation offers a non-contact, full-field measurement approach, acting as an optical virtual strain gauge that provides high spatial resolution. The results measured on detwinned {\BFA} are quantitatively consistent with the distortion measured by X-ray diffraction and neutron Larmor diffraction. These findings highlight the potential of cryogenic digital image correlation as an effective and accessible tool for probing the isotropic and anisotropic strains, facilitating the application of uniaxial strain tuning in the study of quantum materials.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
Evolution of pairing symmetry in FeSe$_{1-x}$S$_x$ as probed by uniaxial-strain tuning of $T_c$
Authors:
Ruixian Liu,
Qi Tang,
Chang Liu,
Chunyi Li,
Kaijuan Zhou,
Qiaoyu Wang,
Xingye Lu
Abstract:
In iron-based superconductors (FeSCs), the interplay between electronic nematicity and superconductivity is essential for understanding the exotic superconducting ground state. In the nematic regime, uniaxial-strain ($\varepsilon$) tuning of the superconducting transition temperature $T_c$ [$ΔT_c(\varepsilon)=α\varepsilon+β\varepsilon^2$] offers a unique approach to investigating the evolution of…
▽ More
In iron-based superconductors (FeSCs), the interplay between electronic nematicity and superconductivity is essential for understanding the exotic superconducting ground state. In the nematic regime, uniaxial-strain ($\varepsilon$) tuning of the superconducting transition temperature $T_c$ [$ΔT_c(\varepsilon)=α\varepsilon+β\varepsilon^2$] offers a unique approach to investigating the evolution of pairing symmetry if both $s$ and $d$ wave pairing instabilities are relevant. Here, we employ uniaxial strain to tune the $T_c$ of FeSe$_{1-x}$S$_x$, in which both nematicity and superconductivity undergo significant changes with doping. While $T_c$ is usually suppressed quadratically with $\varepsilon$ in optimally doped BaFe$_2$As$_2$, $ΔT_c(\varepsilon)$ in FeSe$_{1-x}$S$_x$ dominated by $ΔT_c(\varepsilon)=β\varepsilon^2$ changes its sign from $β$ < $0$ in FeSe to $β$ > $0$ in FeSe$_{1-x}$S$_x$ ($x\gtrsim0.10$), indicating an evolution of the pairing symmetry from an $s_{\pm}$ state towards an $s+d$ wave state. These findings highlight the $ΔT_c(\varepsilon)$ as a powerful probe for elucidating the superconducting pairing symmetry in the nematic regime of FeSCs and provide new insights into the evolution of pairing symmetry in FeSCs.
△ Less
Submitted 18 October, 2024; v1 submitted 17 October, 2024;
originally announced October 2024.
-
DOA Estimation-Oriented Joint Array Partitioning and Beamforming Designs for ISAC Systems
Authors:
Rang Liu,
Ming Li,
Qian Liu,
A. Lee Swindlehurst
Abstract:
Integrated sensing and communication has been identified as an enabling technology for forthcoming wireless networks. In an effort to achieve an improved performance trade-off between multiuser communications and radar sensing, this paper considers a dynamically-partitioned antenna array architecture for monostatic ISAC systems, in which each element of the array at the base station can function a…
▽ More
Integrated sensing and communication has been identified as an enabling technology for forthcoming wireless networks. In an effort to achieve an improved performance trade-off between multiuser communications and radar sensing, this paper considers a dynamically-partitioned antenna array architecture for monostatic ISAC systems, in which each element of the array at the base station can function as either a transmit or receive antenna. To fully exploit the available spatial degrees of freedom for both communication and sensing functions, we jointly design the partitioning of the array between transmit and receive antennas together with the transmit beamforming in order to minimize the direction-of-arrival (DOA) estimation error, while satisfying constraints on the communication signal-to-interference-plus-noise ratio and the transmit power budget. An alternating algorithm based on Dinkelbach's transform, the alternative direction method of multipliers, and majorization-minimization is developed to solve the resulting complicated optimization problem. To reduce the computational complexity, we also present a heuristic three-step strategy that optimizes the transmit beamforming after determining the antenna partitioning. Simulation results confirm the effectiveness of the proposed algorithms in significantly reducing the DOA estimation error.
△ Less
Submitted 16 October, 2024;
originally announced October 2024.
-
Incorporating Metabolic Information into LLMs for Anomaly Detection in Clinical Time-Series
Authors:
Maxx Richard Rahman,
Ruoxuan Liu,
Wolfgang Maass
Abstract:
Anomaly detection in clinical time-series holds significant potential in identifying suspicious patterns in different biological parameters. In this paper, we propose a targeted method that incorporates the clinical domain knowledge into LLMs to improve their ability to detect anomalies. We introduce the Metabolism Pathway-driven Prompting (MPP) method, which integrates the information about metab…
▽ More
Anomaly detection in clinical time-series holds significant potential in identifying suspicious patterns in different biological parameters. In this paper, we propose a targeted method that incorporates the clinical domain knowledge into LLMs to improve their ability to detect anomalies. We introduce the Metabolism Pathway-driven Prompting (MPP) method, which integrates the information about metabolic pathways to better capture the structural and temporal changes in biological samples. We applied our method for doping detection in sports, focusing on steroid metabolism, and evaluated using real-world data from athletes. The results show that our method improves anomaly detection performance by leveraging metabolic context, providing a more nuanced and accurate prediction of suspicious samples in athletes' profiles.
△ Less
Submitted 19 October, 2024; v1 submitted 2 October, 2024;
originally announced October 2024.
-
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
Authors:
Renhang Liu,
Abhinaba Roy,
Dorien Herremans
Abstract:
In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First, we compute LLM embeddings for emotion labels and apply non-parametric clustering to group similar labels, across multiple datasets containing disjoint labels. We use these cluster…
▽ More
In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First, we compute LLM embeddings for emotion labels and apply non-parametric clustering to group similar labels, across multiple datasets containing disjoint labels. We use these cluster centers to map music features (MERT) to the LLM embedding space. To further enhance the model, we introduce an alignment regularization that enables dissociation of MERT embeddings from different clusters. This further enhances the model's ability to better adaptation to unseen datasets. We demonstrate the effectiveness of our approach by performing zero-shot inference on a new dataset, showcasing its ability to generalize to unseen labels without additional training.
△ Less
Submitted 17 October, 2024; v1 submitted 15 October, 2024;
originally announced October 2024.
-
Online Statistical Inference for Time-varying Sample-averaged Q-learning
Authors:
Saunak Kumar Panda,
Ruiqi Liu,
Yisha Xiang
Abstract:
Reinforcement learning (RL) has emerged as a key approach for training agents in complex and uncertain environments. Incorporating statistical inference in RL algorithms is essential for understanding and managing uncertainty in model performance. This paper introduces a time-varying batch-averaged Q-learning algorithm, termed sampleaveraged Q-learning, which improves upon traditional single-sampl…
▽ More
Reinforcement learning (RL) has emerged as a key approach for training agents in complex and uncertain environments. Incorporating statistical inference in RL algorithms is essential for understanding and managing uncertainty in model performance. This paper introduces a time-varying batch-averaged Q-learning algorithm, termed sampleaveraged Q-learning, which improves upon traditional single-sample Q-learning by aggregating samples of rewards and next states to better account for data variability and uncertainty. We leverage the functional central limit theorem (FCLT) to establish a novel framework that provides insights into the asymptotic normality of the sample-averaged algorithm under mild conditions. Additionally, we develop a random scaling method for interval estimation, enabling the construction of confidence intervals without requiring extra hyperparameters. Numerical experiments conducted on classic OpenAI Gym environments show that the time-varying sample-averaged Q-learning method consistently outperforms both single-sample and constant-batch Q-learning methods, achieving superior accuracy while maintaining comparable learning speeds.
△ Less
Submitted 14 October, 2024;
originally announced October 2024.
-
Dual-Path Mechanism of Amino Acid Racemization Mediated by Quantum Mechanical Tunneling
Authors:
Xinrui Yang,
Rui Liu,
Ruiqi Xu,
Zhaohua Cui,
Zhigang Wang
Abstract:
The racemization of amino acids constitutes one of the most elemental and critical reactions, holding primitive significance for understanding the life's origin and maintenance. Nevertheless, its mechanism at the atomic level has been persistently misunderstood for more than a century. In this work, we demonstrate that the racemization of amino acid molecules in aqueous environments can occur simu…
▽ More
The racemization of amino acids constitutes one of the most elemental and critical reactions, holding primitive significance for understanding the life's origin and maintenance. Nevertheless, its mechanism at the atomic level has been persistently misunderstood for more than a century. In this work, we demonstrate that the racemization of amino acid molecules in aqueous environments can occur simultaneously by two pathways via the carboxyl (COOH) and amino (NH2) groups. Behind this result, the quantum mechanical tunneling (QMT) effect plays a pivotal role, as evidenced by the tunneling hindrance of the NH2 reaction and the tunneling enhancement of the COOH reaction. Notably, the disparity in the QMT effect leads to a crossover between the COOH and NH2 reactions within 200-257 K, such that NH2 reactions dominate at high temperatures and COOH reactions dominate at low temperatures. Our work emphasizes the significance of QMT effect in the racemization of amino acids and therefore introduces a dual-path coexistence mechanism, offering valuable insights into the origin of homochirality in extreme environments of the early Earth.
△ Less
Submitted 14 October, 2024;
originally announced October 2024.
-
Individual solid-state nuclear spin qubits with coherence exceeding seconds
Authors:
James O'Sullivan,
Jaime Travesedo,
Louis Pallegoix,
Zhiyuan W. Huang,
Alexande May,
Boris Yavkin,
Patrick Hogan,
Sen Lin,
Renbao Liu,
Thierry Chaneliere,
Sylvain Bertaina,
Philippe Goldner,
Daniel Esteve,
Denis Vion,
Patrick Abgrall,
Patrice Bertet,
Emmanuel Flurin
Abstract:
The ability to coherently control and read out qubits with long coherence times in a scalable system is a crucial requirement for any quantum processor. Nuclear spins in the solid state have shown great promise as long-lived qubits. Control and readout of individual nuclear spin qubit registers has made major progress in the recent years using individual electron spin ancilla addressed either elec…
▽ More
The ability to coherently control and read out qubits with long coherence times in a scalable system is a crucial requirement for any quantum processor. Nuclear spins in the solid state have shown great promise as long-lived qubits. Control and readout of individual nuclear spin qubit registers has made major progress in the recent years using individual electron spin ancilla addressed either electrically or optically. Here, we present a new platform for quantum information processing, consisting of $^{183}$W nuclear spin qubits adjacent to an Er$^{3+}$ impurity in a CaWO$_4$ crystal, interfaced via a superconducting resonator and detected using a microwave photon counter at 10mK. We study two nuclear spin qubits with $T_2^*$ of $0.8(2)~$s and $1.2(3)~$s, $T_2$ of $3.4(4)~$s and $4.4(6)~$ s, respectively. We demonstrate single-shot quantum non-demolition readout of each nuclear spin qubit using the Er$^{3+}$ spin as an ancilla. We introduce a new scheme for all-microwave single- and two-qubit gates, based on stimulated Raman driving of the coupled electron-nuclear spin system. We realize single- and two-qubit gates on a timescale of a few milliseconds, and prepare a decoherence-protected Bell state with 88% fidelity and $T_2^*$ of $1.7(2)~$s. Our results are a proof-of-principle demonstrating the potential of solid-state nuclear spin qubits as a promising platform for quantum information processing. With the potential to scale to tens or hundreds of qubits, this platform has prospects for the development of scalable quantum processors with long-lived qubits.
△ Less
Submitted 14 October, 2024;
originally announced October 2024.
-
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
Authors:
Ying Liu,
Ge Bai,
Chenji Lu,
Shilong Li,
Zhang Zhang,
Ruifang Liu,
Wenbin Guo
Abstract:
Despite the remarkable advancements in Visual Question Answering (VQA), the challenge of mitigating the language bias introduced by textual information remains unresolved. Previous approaches capture language bias from a coarse-grained perspective. However, the finer-grained information within a sentence, such as context and keywords, can result in different biases. Due to the ignorance of fine-gr…
▽ More
Despite the remarkable advancements in Visual Question Answering (VQA), the challenge of mitigating the language bias introduced by textual information remains unresolved. Previous approaches capture language bias from a coarse-grained perspective. However, the finer-grained information within a sentence, such as context and keywords, can result in different biases. Due to the ignorance of fine-grained information, most existing methods fail to sufficiently capture language bias. In this paper, we propose a novel causal intervention training scheme named CIBi to eliminate language bias from a finer-grained perspective. Specifically, we divide the language bias into context bias and keyword bias. We employ causal intervention and contrastive learning to eliminate context bias and improve the multi-modal representation. Additionally, we design a new question-only branch based on counterfactual generation to distill and eliminate keyword bias. Experimental results illustrate that CIBi is applicable to various VQA models, yielding competitive performance.
△ Less
Submitted 14 October, 2024;
originally announced October 2024.
-
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
Authors:
Rui Liu,
Zhenqi Jia,
Jie Yang,
Yifan Hu,
Haizhou Li
Abstract:
Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interac…
▽ More
Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interaction scenarios, due to the scarcity of conversational emphasis datasets and the difficulty in context understanding. In this paper, we propose a novel Emphasis Rendering scheme for the CTTS model, termed ER-CTTS, that includes two main components: 1) we simultaneously take into account textual and acoustic contexts, with both global and local semantic modeling to understand the conversation context comprehensively; 2) we deeply integrate multi-modal and multi-scale context to learn the influence of context on the emphasis expression of the current utterance. Finally, the inferred emphasis feature is fed into the neural speech synthesizer to generate conversational speech. To address data scarcity, we create emphasis intensity annotations on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in emphasis rendering within a conversational setting. The code and audio samples are available at https://github.com/CodeStoreTTS/ER-CTTS.
△ Less
Submitted 12 October, 2024;
originally announced October 2024.
-
FlatQuant: Flatness Matters for LLM Quantization
Authors:
Yuxuan Sun,
Ruikang Liu,
Haoli Bai,
Han Bao,
Kang Zhao,
Yuening Li,
Jiaxin Hu,
Xianzhi Yu,
Lu Hou,
Chun Yuan,
Xin Jiang,
Wulong Liu,
Jun Yao
Abstract:
Recently, quantization has been widely used for the compression and acceleration of large language models~(LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with the equally spaced quantization points. Prior research explores various pre-quantization transformations to suppress outliers, such as per-channel scaling and Hadamard tran…
▽ More
Recently, quantization has been widely used for the compression and acceleration of large language models~(LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with the equally spaced quantization points. Prior research explores various pre-quantization transformations to suppress outliers, such as per-channel scaling and Hadamard transformation. However, we observe that these transformed weights and activations can still remain steep and outspread. In this paper, we propose FlatQuant (Fast and Learnable Affine Transformation), a new post-training quantization approach to enhance flatness of weights and activations. Our approach identifies optimal affine transformations tailored to each linear layer, calibrated in hours via a lightweight objective. To reduce runtime overhead, we apply Kronecker decomposition to the transformation matrices, and fuse all operations in FlatQuant into a single kernel. Extensive experiments show that FlatQuant sets up a new state-of-the-art quantization benchmark. For instance, it achieves less than $\textbf{1}\%$ accuracy drop for W4A4 quantization on the LLaMA-3-70B model, surpassing SpinQuant by $\textbf{7.5}\%$. For inference latency, FlatQuant reduces the slowdown induced by pre-quantization transformation from 0.26x of QuaRot to merely $\textbf{0.07x}$, bringing up to $\textbf{2.3x}$ speedup for prefill and $\textbf{1.7x}$ speedup for decoding, respectively. Code is available at: \url{https://github.com/ruikangliu/FlatQuant}.
△ Less
Submitted 12 October, 2024;
originally announced October 2024.
-
When Graph meets Multimodal: Benchmarking on Multimodal Attributed Graphs Learning
Authors:
Hao Yan,
Chaozhuo Li,
Zhigang Yu,
Jun Yin,
Ruochen Liu,
Peiyan Zhang,
Weihao Han,
Mingzheng Li,
Zhengxin Zeng,
Hao Sun,
Weiwei Deng,
Feng Sun,
Qi Zhang,
Senzhang Wang
Abstract:
Multimodal attributed graphs (MAGs) are prevalent in various real-world scenarios and generally contain two kinds of knowledge: (a) Attribute knowledge is mainly supported by the attributes of different modalities contained in nodes (entities) themselves, such as texts and images. (b) Topology knowledge, on the other hand, is provided by the complex interactions posed between nodes. The cornerston…
▽ More
Multimodal attributed graphs (MAGs) are prevalent in various real-world scenarios and generally contain two kinds of knowledge: (a) Attribute knowledge is mainly supported by the attributes of different modalities contained in nodes (entities) themselves, such as texts and images. (b) Topology knowledge, on the other hand, is provided by the complex interactions posed between nodes. The cornerstone of MAG representation learning lies in the seamless integration of multimodal attributes and topology. Recent advancements in Pre-trained Language/Vision models (PLMs/PVMs) and Graph neural networks (GNNs) have facilitated effective learning on MAGs, garnering increased research interest. However, the absence of meaningful benchmark datasets and standardized evaluation procedures for MAG representation learning has impeded progress in this field. In this paper, we propose Multimodal Attribute Graph Benchmark (MAGB)}, a comprehensive and diverse collection of challenging benchmark datasets for MAGs. The MAGB datasets are notably large in scale and encompass a wide range of domains, spanning from e-commerce networks to social networks. In addition to the brand-new datasets, we conduct extensive benchmark experiments over MAGB with various learning paradigms, ranging from GNN-based and PLM-based methods, to explore the necessity and feasibility of integrating multimodal attributes and graph topology. In a nutshell, we provide an overview of the MAG datasets, standardized evaluation procedures, and present baseline experiments. The entire MAGB project is publicly accessible at https://github.com/sktsherlock/ATG.
△ Less
Submitted 11 October, 2024;
originally announced October 2024.
-
Context-Aware Adapter Tuning for Few-Shot Relation Learning in Knowledge Graphs
Authors:
Ran Liu,
Zhongzhou Liu,
Xiaoli Li,
Yuan Fang
Abstract:
Knowledge graphs (KGs) are instrumental in various real-world applications, yet they often suffer from incompleteness due to missing relations. To predict instances for novel relations with limited training examples, few-shot relation learning approaches have emerged, utilizing techniques such as meta-learning. However, the assumption is that novel relations in meta-testing and base relations in m…
▽ More
Knowledge graphs (KGs) are instrumental in various real-world applications, yet they often suffer from incompleteness due to missing relations. To predict instances for novel relations with limited training examples, few-shot relation learning approaches have emerged, utilizing techniques such as meta-learning. However, the assumption is that novel relations in meta-testing and base relations in meta-training are independently and identically distributed, which may not hold in practice. To address the limitation, we propose RelAdapter, a context-aware adapter for few-shot relation learning in KGs designed to enhance the adaptation process in meta-learning. First, RelAdapter is equipped with a lightweight adapter module that facilitates relation-specific, tunable adaptation of meta-knowledge in a parameter-efficient manner. Second, RelAdapter is enriched with contextual information about the target relation, enabling enhanced adaptation to each distinct relation. Extensive experiments on three benchmark KGs validate the superiority of RelAdapter over state-of-the-art methods.
△ Less
Submitted 17 October, 2024; v1 submitted 11 October, 2024;
originally announced October 2024.
-
Cross-Currency Basis Swaps Referencing Backward-Looking Rates
Authors:
Yining Ding,
Ruyi Liu,
Marek Rutkowski
Abstract:
The financial industry has undergone a significant transition from the London Interbank Offered Rate (LIBOR) to Risk Free Rates (RFR) such as, e.g., the Secured Overnight Financing Rate (SOFR) in the U.S. and the AUD Overnight Index Average (AONIA) in Australia, as the primary benchmark rate for borrowing costs. The paper examines the pricing and hedging method for SOFR-related financial products…
▽ More
The financial industry has undergone a significant transition from the London Interbank Offered Rate (LIBOR) to Risk Free Rates (RFR) such as, e.g., the Secured Overnight Financing Rate (SOFR) in the U.S. and the AUD Overnight Index Average (AONIA) in Australia, as the primary benchmark rate for borrowing costs. The paper examines the pricing and hedging method for SOFR-related financial products in a cross-currency context with the special emphasis on the Compound SOFR vs Average AONIA cross-currency basis swaps. While the SOFR and AONIA serve as a particular case of a cross-currency basis swap (CCBS), the approach developed is able to handle backward-looking term rates for any two currencies. We give explicit pricing and hedging results for collateralized cross-currency basis swaps using interest rate and currency futures contracts as hedging tools within an arbitrage-free multi-curve setting.
△ Less
Submitted 10 October, 2024;
originally announced October 2024.
-
Opacity Enforcement by Edit Functions Under Incomparable Observations
Authors:
Wei Duan,
Ruotian Liu,
Maria Pia Fanti,
Christoforos N. Hadjicostis,
Zhiwu Li
Abstract:
As an information-flow privacy property, opacity characterizes whether a malicious external observer (referred to as an intruder) is able to infer the secret behavior of a system. This paper addresses the problem of opacity enforcement using edit functions in discrete event systems modeled by partially observed deterministic finite automata. A defender uses the edit function as an interface at the…
▽ More
As an information-flow privacy property, opacity characterizes whether a malicious external observer (referred to as an intruder) is able to infer the secret behavior of a system. This paper addresses the problem of opacity enforcement using edit functions in discrete event systems modeled by partially observed deterministic finite automata. A defender uses the edit function as an interface at the output of a system to manipulate actual observations through insertion, substitution, and deletion operations so that the intruder will be prevented from inferring the secret behavior of the system. Unlike existing work which usually assumes that the observation capabilities of the intruder and the defender are identical, we consider a more general setting where they may observe incomparable subsets of events generated by the system.To characterize whether the defender has the ability to enforce opacity of the system under this setting, the notion of \emph{$ic$-enforceability} is introduced. Then, the opacity enforcement problem is transformed to a two-player game, with imperfect information between the system and the defender, which can be used to determine a feasible decision-making strategy for the defender. Within the game scheme, an edit mechanism is constructed to enumerate all feasible edit actions following system behavior. We further show that an $ic$-enforcing edit function (if one exists) can be synthesized from the edit mechanism to enforce opacity.
△ Less
Submitted 10 October, 2024;
originally announced October 2024.
-
Generalizable autoregressive modeling of time series through functional narratives
Authors:
Ran Liu,
Wenrui Ma,
Ellen Zippi,
Hadi Pouransari,
Jingyun Xiao,
Chris Sandino,
Behrooz Mahasseni,
Juri Minxha,
Erdrin Azemi,
Eva L. Dyer,
Ali Moin
Abstract:
Time series data are inherently functions of time, yet current transformers often learn time series by modeling them as mere concatenations of time periods, overlooking their functional properties. In this work, we propose a novel objective for transformers that learn time series by re-interpreting them as temporal functions. We build an alternative sequence of time series by constructing degradat…
▽ More
Time series data are inherently functions of time, yet current transformers often learn time series by modeling them as mere concatenations of time periods, overlooking their functional properties. In this work, we propose a novel objective for transformers that learn time series by re-interpreting them as temporal functions. We build an alternative sequence of time series by constructing degradation operators of different intensity in the functional space, creating augmented variants of the original sample that are abstracted or simplified to different degrees. Based on the new set of generated sequence, we train an autoregressive transformer that progressively recovers the original sample from the most simplified variant. Analogous to the next word prediction task in languages that learns narratives by connecting different words, our autoregressive transformer aims to learn the Narratives of Time Series (NoTS) by connecting different functions in time. Theoretically, we justify the construction of the alternative sequence through its advantages in approximating functions. When learning time series data with transformers, constructing sequences of temporal functions allows for a broader class of approximable functions (e.g., differentiation) compared to sequences of time periods, leading to a 26\% performance improvement in synthetic feature regression experiments. Experimentally, we validate NoTS in 3 different tasks across 22 real-world datasets, where we show that NoTS significantly outperforms other pre-training methods by up to 6\%. Additionally, combining NoTS on top of existing transformer architectures can consistently boost the performance. Our results demonstrate the potential of NoTS as a general-purpose dynamic learner, offering a viable alternative for developing foundation models for time series analysis.
△ Less
Submitted 10 October, 2024;
originally announced October 2024.
-
Diversified and Adaptive Negative Sampling on Knowledge Graphs
Authors:
Ran Liu,
Zhongzhou Liu,
Xiaoli Li,
Hao Wu,
Yuan Fang
Abstract:
In knowledge graph embedding, aside from positive triplets (ie: facts in the knowledge graph), the negative triplets used for training also have a direct influence on the model performance. In reality, since knowledge graphs are sparse and incomplete, negative triplets often lack explicit labels, and thus they are often obtained from various sampling strategies (eg: randomly replacing an entity in…
▽ More
In knowledge graph embedding, aside from positive triplets (ie: facts in the knowledge graph), the negative triplets used for training also have a direct influence on the model performance. In reality, since knowledge graphs are sparse and incomplete, negative triplets often lack explicit labels, and thus they are often obtained from various sampling strategies (eg: randomly replacing an entity in a positive triplet). An ideal sampled negative triplet should be informative enough to help the model train better. However, existing methods often ignore diversity and adaptiveness in their sampling process, which harms the informativeness of negative triplets. As such, we propose a generative adversarial approach called Diversified and Adaptive Negative Sampling DANS on knowledge graphs. DANS is equipped with a two-way generator that generates more diverse negative triplets through two pathways, and an adaptive mechanism that produces more fine-grained examples by localizing the global generator for different entities and relations. On the one hand, the two-way generator increase the overall informativeness with more diverse negative examples; on the other hand, the adaptive mechanism increases the individual sample-wise informativeness with more fine-grained sampling. Finally, we evaluate the performance of DANS on three benchmark knowledge graphs to demonstrate its effectiveness through quantitative and qualitative experiments.
△ Less
Submitted 9 October, 2024;
originally announced October 2024.
-
Retrieved dropout imputation considering administrative study withdrawal
Authors:
Rong Liu,
Yongming Qu
Abstract:
The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) E9 (R1) Addendum provides a framework for defining estimands in clinical trials. Treatment policy strategy is the mostly used approach to handle intercurrent events in defining estimands. Imputing missing values for potential outcomes under the treatment policy strategy has been discussed…
▽ More
The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) E9 (R1) Addendum provides a framework for defining estimands in clinical trials. Treatment policy strategy is the mostly used approach to handle intercurrent events in defining estimands. Imputing missing values for potential outcomes under the treatment policy strategy has been discussed in the literature. Missing values as a result of administrative study withdrawals (such as site closures due to business reasons, COVID-19 control measures, and geopolitical conflicts, etc.) are often imputed in the same way as other missing values occurring after intercurrent events related to safety or efficacy. Some research suggests using a hypothetical strategy to handle the treatment discontinuations due to administrative study withdrawal in defining the estimands and imputing the missing values based on completer data assuming missing at random, but this approach ignores the fact that subjects might experience other intercurrent events had they not had the administrative study withdrawal. In this article, we consider the administrative study withdrawal censors the normal real-world like intercurrent events and propose two methods for handling the corresponding missing values under the retrieved dropout imputation framework. Simulation shows the two methods perform well. We also applied the methods to actual clinical trial data evaluating an anti-diabetes treatment.
△ Less
Submitted 9 October, 2024;
originally announced October 2024.
-
Continual Learning in the Frequency Domain
Authors:
Ruiqi Liu,
Boyu Diao,
Libo Huang,
Zijia An,
Zhulin An,
Yongjun Xu
Abstract:
Continual learning (CL) is designed to learn new tasks while preserving existing knowledge. Replaying samples from earlier tasks has proven to be an effective method to mitigate the forgetting of previously acquired knowledge. However, the current research on the training efficiency of rehearsal-based methods is insufficient, which limits the practical application of CL systems in resource-limited…
▽ More
Continual learning (CL) is designed to learn new tasks while preserving existing knowledge. Replaying samples from earlier tasks has proven to be an effective method to mitigate the forgetting of previously acquired knowledge. However, the current research on the training efficiency of rehearsal-based methods is insufficient, which limits the practical application of CL systems in resource-limited scenarios. The human visual system (HVS) exhibits varying sensitivities to different frequency components, enabling the efficient elimination of visually redundant information. Inspired by HVS, we propose a novel framework called Continual Learning in the Frequency Domain (CLFD). To our knowledge, this is the first study to utilize frequency domain features to enhance the performance and efficiency of CL training on edge devices. For the input features of the feature extractor, CLFD employs wavelet transform to map the original input image into the frequency domain, thereby effectively reducing the size of input feature maps. Regarding the output features of the feature extractor, CLFD selectively utilizes output features for distinct classes for classification, thereby balancing the reusability and interference of output features based on the frequency domain similarity of the classes across various tasks. Optimizing only the input and output features of the feature extractor allows for seamless integration of CLFD with various rehearsal-based methods. Extensive experiments conducted in both cloud and edge environments demonstrate that CLFD consistently improves the performance of state-of-the-art (SOTA) methods in both precision and training efficiency. Specifically, CLFD can increase the accuracy of the SOTA CL method by up to 6.83% and reduce the training time by 2.6$\times$.
△ Less
Submitted 10 October, 2024; v1 submitted 9 October, 2024;
originally announced October 2024.
-
A convex formulation of covariate-adjusted Gaussian graphical models via natural parametrization
Authors:
Ruobin Liu,
Guo Yu
Abstract:
Gaussian graphical models (GGMs) are widely used for recovering the conditional independence structure among random variables. Recently, several key advances have been made to exploit an additional set of variables for better estimating the GGMs of the variables of interest. For example, in co-expression quantitative trait locus (eQTL) studies, both the mean expression level of genes as well as th…
▽ More
Gaussian graphical models (GGMs) are widely used for recovering the conditional independence structure among random variables. Recently, several key advances have been made to exploit an additional set of variables for better estimating the GGMs of the variables of interest. For example, in co-expression quantitative trait locus (eQTL) studies, both the mean expression level of genes as well as their pairwise conditional independence structure may be adjusted by genetic variants local to those genes. Existing methods to estimate covariate-adjusted GGMs either allow only the mean to depend on covariates or suffer from poor scaling assumptions due to the inherent non-convexity of simultaneously estimating the mean and precision matrix. In this paper, we propose a convex formulation that jointly estimates the covariate-adjusted mean and precision matrix by utilizing the natural parametrization of the multivariate Gaussian likelihood. This convexity yields theoretically better performance as the sparsity and dimension of the covariates grow large relative to the number of samples. We verify our theoretical results with numerical simulations and perform a reanalysis of an eQTL study of glioblastoma multiforme (GBM), an aggressive form of brain cancer.
△ Less
Submitted 8 October, 2024;
originally announced October 2024.
-
TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data
Authors:
Jeremy Andrew Irvin,
Emily Ruoyu Liu,
Joyce Chuyi Chen,
Ines Dormoy,
Jinyoung Kim,
Samar Khanna,
Zhuo Zheng,
Stefano Ermon
Abstract:
Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal seque…
▽ More
Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than specialist models trained to perform these specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single EO image instruction-following model. We publicly release our data, models, and code at https://github.com/ermongroup/TEOChat .
△ Less
Submitted 8 October, 2024;
originally announced October 2024.
-
On the External Inverse Compton Scattering off the Prompt Emission in GRB 221009A
Authors:
Cui-Yuan Dai,
Jian-He Zheng,
Xiao-Hong Zhao,
Ruo-Yu Liu,
Xiang-Yu Wang
Abstract:
The light curve of the TeV emission in GRB 221009A displays a smooth transition from an initial rapid rise to a slower rise and eventually a decay phase. The smooth temporal profile of the TeV emission suggests that it mainly results from an external shock. The temporal overlap between the prompt KeV-MeV emission and the early TeV afterglow indicates that external inverse Compton scattering (EIC)…
▽ More
The light curve of the TeV emission in GRB 221009A displays a smooth transition from an initial rapid rise to a slower rise and eventually a decay phase. The smooth temporal profile of the TeV emission suggests that it mainly results from an external shock. The temporal overlap between the prompt KeV-MeV emission and the early TeV afterglow indicates that external inverse Compton scattering (EIC) between the prompt KeV-MeV photons and the afterglow electrons is inevitable. Since the energy density of the prompt emission is much higher than that of the afterglow during the early phase, the EIC process dominates the cooling of afterglow electrons. The EIC scattering rate is influenced by the anisotropy of the seed photon field, which depends on the radii of the internal dissipation ($R_{\rm dis}$), where the prompt emission is produced, and that of the external shock ($R_{\rm ext}$), where the afterglow emission is produced. We investigate the EIC process for different values of $R_{\rm dis}/R_{\rm ext}$. We find that, for varying \( R_{\rm dis}/R_{\rm ext} \), the EIC scattering rate can differ by a factor of $\sim 2$. For GRB 221009A, the EIC emission is dominated during the early rising phase of the TeV afterglow. It then transitions to a phase dominated by the synchrotron self-Compton (SSC) emission as the intensity of the prompt emission decreases. Additionally, we investigate the effect of $γγ$ absorption in the TeV afterglow caused by prompt MeV photons and find that it is insufficient to explain the early rapid rise in the TeV afterglow, even in the case of $R_{\rm dis}/R_{\rm ext} \sim 1$.
△ Less
Submitted 8 October, 2024;
originally announced October 2024.
-
Training Interactive Agent in Large FPS Game Map with Rule-enhanced Reinforcement Learning
Authors:
Chen Zhang,
Huan Hu,
Yuan Zhou,
Qiyang Cao,
Ruochen Liu,
Wenya Wei,
Elvis S. Liu
Abstract:
In the realm of competitive gaming, 3D first-person shooter (FPS) games have gained immense popularity, prompting the development of game AI systems to enhance gameplay. However, deploying game AI in practical scenarios still poses challenges, particularly in large-scale and complex FPS games. In this paper, we focus on the practical deployment of game AI in the online multiplayer competitive 3D F…
▽ More
In the realm of competitive gaming, 3D first-person shooter (FPS) games have gained immense popularity, prompting the development of game AI systems to enhance gameplay. However, deploying game AI in practical scenarios still poses challenges, particularly in large-scale and complex FPS games. In this paper, we focus on the practical deployment of game AI in the online multiplayer competitive 3D FPS game called Arena Breakout, developed by Tencent Games. We propose a novel gaming AI system named Private Military Company Agent (PMCA), which is interactable within a large game map and engages in combat with players while utilizing tactical advantages provided by the surrounding terrain.
To address the challenges of navigation and combat in modern 3D FPS games, we introduce a method that combines navigation mesh (Navmesh) and shooting-rule with deep reinforcement learning (NSRL). The integration of Navmesh enhances the agent's global navigation capabilities while shooting behavior is controlled using rule-based methods to ensure controllability. NSRL employs a DRL model to predict when to enable the navigation mesh, resulting in a diverse range of behaviors for the game AI. Customized rewards for human-like behaviors are also employed to align PMCA's behavior with that of human players.
△ Less
Submitted 7 October, 2024;
originally announced October 2024.
-
LHAASO detection of very-high-energy gamma-ray emission surrounding PSR J0248+6021
Authors:
Zhen Cao,
F. Aharonian,
Q. An,
Axikegu,
Y. X. Bai,
Y. W. Bao,
D. Bastieri,
X. J. Bi,
Y. J. Bi,
J. T. Cai,
Q. Cao,
W. Y. Cao,
Zhe Cao,
J. Chang,
J. F. Chang,
A. M. Chen,
E. S. Chen,
Liang Chen,
Lin Chen,
Long Chen,
M. J. Chen,
M. L. Chen,
Q. H. Chen,
S. H. Chen,
S. Z. Chen
, et al. (255 additional authors not shown)
Abstract:
We report the detection of an extended very-high-energy (VHE) gamma-ray source coincident with the locations of middle-aged (62.4~\rm kyr) pulsar PSR J0248+6021, by using the LHAASO-WCDA data of live 796 days and LHAASO-KM2A data of live 1216 days. A significant excess of \gray induced showers is observed both by WCDA in energy bands of 1-25~\rm TeV and KM2A in energy bands of $>$ 25~\rm TeV with…
▽ More
We report the detection of an extended very-high-energy (VHE) gamma-ray source coincident with the locations of middle-aged (62.4~\rm kyr) pulsar PSR J0248+6021, by using the LHAASO-WCDA data of live 796 days and LHAASO-KM2A data of live 1216 days. A significant excess of \gray induced showers is observed both by WCDA in energy bands of 1-25~\rm TeV and KM2A in energy bands of $>$ 25~\rm TeV with 7.3 $σ$ and 13.5 $σ$, respectively. The best-fit position derived through WCDA data is R.A. = 42.06$^\circ \pm$ 0.12$^\circ$ and Dec. = 60.24$^\circ \pm $ 0.13$^\circ$ with an extension of 0.69$^\circ\pm$0.15$^\circ$ and that of the KM2A data is R.A.= 42.29$^\circ \pm $ 0.13$^\circ$ and Dec. = 60.38$^\circ \pm$ 0.07$^\circ$ with an extension of 0.37$^\circ\pm$0.07$^\circ$. No clear extended multiwavelength counterpart of this LHAASO source has been found from the radio band to the GeV band. The most plausible explanation of the VHE \gray emission is the inverse Compton process of highly relativistic electrons and positrons injected by the pulsar. These electrons/positrons are hypothesized to be either confined within the pulsar wind nebula or to have already escaped into the interstellar medium, forming a pulsar halo.
△ Less
Submitted 6 October, 2024;
originally announced October 2024.
-
FluentEditor+: Text-based Speech Editing by Modeling Local Hierarchical Acoustic Smoothness and Global Prosody Consistency
Authors:
Rui Liu,
Jiatian Xi,
Ziyue Jiang,
Haizhou Li
Abstract:
Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to g…
▽ More
Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. While current TSE techniques focus on minimizing discrepancies between generated speech and reference targets within edited segments, they often neglect the importance of maintaining both local and global fluency in the context of the original discourse. Additionally, seamlessly integrating edited segments with unaltered portions of the audio remains challenging, typically requiring support from text-to-speech (TTS) systems. This paper introduces a novel approach, FluentEditor$\tiny +$, designed to overcome these limitations. FluentEditor$\tiny +$ employs advanced feature extraction techniques to capture both acoustic and prosodic characteristics, ensuring fluent transitions between edited and unedited regions. The model ensures segmental acoustic smoothness and global prosody consistency, allowing seamless splicing of speech while preserving the coherence and naturalness of the output. Extensive experiments on the VCTK and LibriTTS datasets show that FluentEditor$\tiny +$ surpasses existing TTS-based methods, including Editspeech, Campnet, $A^3T$ FluentSpeech, and Fluenteditor, in both fluency and prosody. Ablation studies further highlight the contributions of each module to the overall effectiveness of the system.
△ Less
Submitted 28 September, 2024;
originally announced October 2024.
-
Self-Assembly of a halogenated organic molecule on the Si(111) $\surd$3$\times$$\surd$3-Ag surface
Authors:
R Liu,
D. Marchese,
R. C. Mawhinney,
M. C. Gallagher
Abstract:
We study the self-assembly of halogen-based organic molecules on a passivated silicon surface. The room temperature adsorption of 2,4,6-tris(4-iodophenyl)-1,3,5-triazine (TIPT) on the Si(111)-$\surd$3$\times$$\surd$3-Ag surface is described. The adsorption is investigated primarily by room-temperature scanning tunneling microscopy (STM) and density-functional theoretical (DFT) calculations. The ex…
▽ More
We study the self-assembly of halogen-based organic molecules on a passivated silicon surface. The room temperature adsorption of 2,4,6-tris(4-iodophenyl)-1,3,5-triazine (TIPT) on the Si(111)-$\surd$3$\times$$\surd$3-Ag surface is described. The adsorption is investigated primarily by room-temperature scanning tunneling microscopy (STM) and density-functional theoretical (DFT) calculations. The experimental results is a dramatic example of how the substrate can influence the overall structure of the self-assembly. With increasing dose, the TIPT monomers form supramolecular structures defined by a two monomer, 2.07 $\pm$ 0.05 nm by 1.83 $\pm$ 0.05 nm rectangular cell. The unit cell is characterized by zig-zag rows of molecules aligned \pm13° from the high symmetry directions of the $\surd$3-Ag substrate. The 2.07 nm dimension along the zig-zag rows is very similar to self-assembled TIPT networks observed on HOPG, however the 1.83 nm dimension is extended considerably and commensurate with the $\surd$3-Ag substrate. The epitaxial relationship between the overlayer and the substrate, and the commensurate inter-row spacing indicate significant molecule-substrate interactions. In fact, DFT calculations of free standing TIPT hexamers reveal that increasing the inter row spacing comes at little energy cost. Experiments also indicate that the formation of supramolecular TIPT domains is extremely sensitive to the quality of the underlying $\surd3$-Ag reconstruction. Point defects in the $\surd3$-Ag reconstruction ultimately restricts the extent of the observed domains.
△ Less
Submitted 3 October, 2024;
originally announced October 2024.
-
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
Authors:
Qi Fan,
Hongyu Yuan,
Haolin Zuo,
Rui Liu,
Guanglai Gao
Abstract:
Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full modality integrity is often not applicable in reality and there always appears the situation that some modalities are missing. For example, video, audio, or text data is missing due to sensor failure or network bandwidth p…
▽ More
Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full modality integrity is often not applicable in reality and there always appears the situation that some modalities are missing. For example, video, audio, or text data is missing due to sensor failure or network bandwidth problems, which presents a great challenge to MER research. Traditional methods extract useful information from the complete modalities and reconstruct the missing modalities to learn robust multimodal joint representation. These methods have laid a solid foundation for research in this field, and to a certain extent, alleviated the difficulty of multimodal emotion recognition under missing modalities. However, relying solely on internal reconstruction and multimodal joint learning has its limitations, especially when the missing information is critical for emotion recognition. To address this challenge, we propose a novel framework of Retrieval Augment for Missing Modality Multimodal Emotion Recognition (RAMER), which introduces similar multimodal emotion data to enhance the performance of emotion recognition under missing modalities. By leveraging databases, that contain related multimodal emotion data, we can retrieve similar multimodal emotion information to fill in the gaps left by missing modalities. Various experimental results demonstrate that our framework is superior to existing state-of-the-art approaches in missing modality MER tasks. Our whole project is publicly available on https://github.com/WooyoohL/Retrieval_Augment_MER.
△ Less
Submitted 18 September, 2024;
originally announced October 2024.
-
C-MORL: Multi-Objective Reinforcement Learning through Efficient Discovery of Pareto Front
Authors:
Ruohong Liu,
Yuxin Pan,
Linjie Xu,
Lei Song,
Pengcheng You,
Yize Chen,
Jiang Bian
Abstract:
Multi-objective reinforcement learning (MORL) excels at handling rapidly changing preferences in tasks that involve multiple criteria, even for unseen preferences. However, previous dominating MORL methods typically generate a fixed policy set or preference-conditioned policy through multiple training iterations exclusively for sampled preference vectors, and cannot ensure the efficient discovery…
▽ More
Multi-objective reinforcement learning (MORL) excels at handling rapidly changing preferences in tasks that involve multiple criteria, even for unseen preferences. However, previous dominating MORL methods typically generate a fixed policy set or preference-conditioned policy through multiple training iterations exclusively for sampled preference vectors, and cannot ensure the efficient discovery of the Pareto front. Furthermore, integrating preferences into the input of policy or value functions presents scalability challenges, in particular as the dimension of the state and preference space grow, which can complicate the learning process and hinder the algorithm's performance on more complex tasks. To address these issues, we propose a two-stage Pareto front discovery algorithm called Constrained MORL (C-MORL), which serves as a seamless bridge between constrained policy optimization and MORL. Concretely, a set of policies is trained in parallel in the initialization stage, with each optimized towards its individual preference over the multiple objectives. Then, to fill the remaining vacancies in the Pareto front, the constrained optimization steps are employed to maximize one objective while constraining the other objectives to exceed a predefined threshold. Empirically, compared to recent advancements in MORL methods, our algorithm achieves more consistent and superior performances in terms of hypervolume, expected utility, and sparsity on both discrete and continuous control tasks, especially with numerous objectives (up to nine objectives in our experiments).
△ Less
Submitted 3 October, 2024;
originally announced October 2024.
-
Breaking the mold: The challenge of large scale MARL specialization
Authors:
Stefan Juang,
Hugh Cao,
Arielle Zhou,
Ruochen Liu,
Nevin L. Zhang,
Elvis Liu
Abstract:
In multi-agent learning, the predominant approach focuses on generalization, often neglecting the optimization of individual agents. This emphasis on generalization limits the ability of agents to utilize their unique strengths, resulting in inefficiencies. This paper introduces Comparative Advantage Maximization (CAM), a method designed to enhance individual agent specialization in multiagent sys…
▽ More
In multi-agent learning, the predominant approach focuses on generalization, often neglecting the optimization of individual agents. This emphasis on generalization limits the ability of agents to utilize their unique strengths, resulting in inefficiencies. This paper introduces Comparative Advantage Maximization (CAM), a method designed to enhance individual agent specialization in multiagent systems. CAM employs a two-phase process, combining centralized population training with individual specialization through comparative advantage maximization. CAM achieved a 13.2% improvement in individual agent performance and a 14.9% increase in behavioral diversity compared to state-of-the-art systems. The success of CAM highlights the importance of individual agent specialization, suggesting new directions for multi-agent system development.
△ Less
Submitted 2 October, 2024;
originally announced October 2024.
-
SegEarth-OV: Towards Traning-Free Open-Vocabulary Segmentation for Remote Sensing Images
Authors:
Kaiyu Li,
Ruixun Liu,
Xiangyong Cao,
Deyu Meng,
Zhi Wang
Abstract:
Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of remote sensing image applications; however, a prevalent limitation remains the need for extensive manual annotation. For this, we try to introduce open-vocabulary semantic segmentation (OVSS) into the remote sensing conte…
▽ More
Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of remote sensing image applications; however, a prevalent limitation remains the need for extensive manual annotation. For this, we try to introduce open-vocabulary semantic segmentation (OVSS) into the remote sensing context. However, due to the sensitivity of remote sensing images to low-resolution features, distorted target shapes and ill-fitting boundaries are exhibited in the prediction mask. To tackle this issue, we propose a simple and general upsampler, SimFeatUp, to restore lost spatial information in deep features in a training-free style. Further, based on the observation of the abnormal response of local patch tokens to [CLS] token in CLIP, we propose to execute a straightforward subtraction operation to alleviate the global bias in patch tokens. Extensive experiments are conducted on 17 remote sensing datasets spanning semantic segmentation, building extraction, road detection, and flood detection tasks. Our method achieves an average of 5.8%, 8.2%, 4%, and 15.3% improvement over state-of-the-art methods on 4 tasks. All codes are released. \url{https://earth-insights.github.io/SegEarth-OV}
△ Less
Submitted 2 October, 2024;
originally announced October 2024.
-
Open-vocabulary Multimodal Emotion Recognition: Dataset, Metric, and Benchmark
Authors:
Zheng Lian,
Haiyang Sun,
Licai Sun,
Lan Chen,
Haoyu Chen,
Hao Gu,
Zhuofan Wen,
Shun Chen,
Siyuan Zhang,
Hailiang Yao,
Mingyu Xu,
Kang Chen,
Bin Liu,
Rui Liu,
Shan Liang,
Ya Li,
Jiangyan Yi,
Jianhua Tao
Abstract:
Multimodal Emotion Recognition (MER) is an important research topic. This paper advocates for a transformative paradigm in MER. The rationale behind our work is that current approaches often rely on a limited set of basic emotion labels, which do not adequately represent the rich spectrum of human emotions. These traditional and overly simplistic emotion categories fail to capture the inherent com…
▽ More
Multimodal Emotion Recognition (MER) is an important research topic. This paper advocates for a transformative paradigm in MER. The rationale behind our work is that current approaches often rely on a limited set of basic emotion labels, which do not adequately represent the rich spectrum of human emotions. These traditional and overly simplistic emotion categories fail to capture the inherent complexity and subtlety of human emotional experiences, leading to limited generalizability and practicality. Therefore, we propose a new MER paradigm called Open-vocabulary MER (OV-MER), which encompasses a broader range of emotion labels to reflect the richness of human emotions. This paradigm relaxes the label space, allowing for the prediction of arbitrary numbers and categories of emotions. To support this transition, we provide a comprehensive solution that includes a newly constructed database based on LLM and human collaborative annotations, along with corresponding metrics and a series of benchmarks. We hope this work advances emotion recognition from basic emotions to more nuanced emotions, contributing to the development of emotional AI.
△ Less
Submitted 2 October, 2024;
originally announced October 2024.
-
PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
Authors:
Hongbo Wang,
Mingda Li,
Junyu Lu,
Hebin Xia,
Liang Yang,
Bo Xu,
Ruizhu Liu,
Hongfei Lin
Abstract:
Disclaimer: Samples in this paper may be harmful and cause discomfort!
Patronizing and condescending language (PCL) is a form of speech directed at vulnerable groups. As an essential branch of toxic language, this type of language exacerbates conflicts and confrontations among Internet communities and detrimentally impacts disadvantaged groups. Traditional pre-trained language models (PLMs) perf…
▽ More
Disclaimer: Samples in this paper may be harmful and cause discomfort!
Patronizing and condescending language (PCL) is a form of speech directed at vulnerable groups. As an essential branch of toxic language, this type of language exacerbates conflicts and confrontations among Internet communities and detrimentally impacts disadvantaged groups. Traditional pre-trained language models (PLMs) perform poorly in detecting PCL due to its implicit toxicity traits like hypocrisy and false sympathy. With the rise of large language models (LLMs), we can harness their rich emotional semantics to establish a paradigm for exploring implicit toxicity. In this paper, we introduce PclGPT, a comprehensive LLM benchmark designed specifically for PCL. We collect, annotate, and integrate the Pcl-PT/SFT dataset, and then develop a bilingual PclGPT-EN/CN model group through a comprehensive pre-training and supervised fine-tuning staircase process to facilitate implicit toxic detection. Group detection results and fine-grained detection from PclGPT and other models reveal significant variations in the degree of bias in PCL towards different vulnerable groups, necessitating increased societal attention to protect them.
△ Less
Submitted 30 September, 2024;
originally announced October 2024.
-
Pre-Chirp-Domain Index Modulation for Full-Diversity Affine Frequency Division Multiplexing towards 6G
Authors:
Guangyao Liu,
Tianqi Mao,
Zhenyu Xiao,
Ruiqi Liu,
Miaowen Wen
Abstract:
Affine frequency division multiplexing (AFDM), tailored as a superior multicarrier technique utilizing chirp signals for high-mobility communications, is envisioned as a promising candidate for the sixth-generation (6G) wireless network. AFDM is based on the discrete affine Fourier transform (DAFT) with two adjustable parameters of the chirp signals, termed as the pre-chirp and post-chirp paramete…
▽ More
Affine frequency division multiplexing (AFDM), tailored as a superior multicarrier technique utilizing chirp signals for high-mobility communications, is envisioned as a promising candidate for the sixth-generation (6G) wireless network. AFDM is based on the discrete affine Fourier transform (DAFT) with two adjustable parameters of the chirp signals, termed as the pre-chirp and post-chirp parameters, respectively. We show that the pre-chirp counterpart can be flexibly manipulated for additional degree-of-freedom (DoF). Therefore, this paper proposes a novel AFDM scheme with the pre-chirp index modulation (PIM) philosophy (AFDM-PIM), which can implicitly convey extra information bits through dynamic pre-chirp parameter assignment, thus enhancing both spectral and energy efficiency. Specifically, we first demonstrate that the subcarrier orthogonality is still maintained by applying distinct pre-chirp parameters to various subcarriers in the AFDM modulation process. Inspired by this property, each AFDM subcarrier is constituted with a unique pre-chirp signal according to the incoming bits. By such arrangement, extra binary bits can be embedded into the index patterns of pre-chirp parameter assignment without additional energy consumption. For performance analysis, we derive the asymptotically tight upper bounds on the average bit error rates (BERs) of the proposed schemes with maximum-likelihood (ML) detection, and validate that the proposed AFDM-PIM can achieve the optimal diversity order under doubly dispersive channels. Based on the derivations, we further propose an optimal pre-chirp alphabet design to enhance the BER performance via intelligent optimization algorithms. Simulations demonstrate that the proposed AFDM-PIM outperforms the classical benchmarks under doubly dispersive channel.
△ Less
Submitted 17 October, 2024; v1 submitted 30 September, 2024;
originally announced October 2024.
-
Building Real-time Awareness of Out-of-distribution in Trajectory Prediction for Autonomous Vehicles
Authors:
Tongfei,
Guo,
Taposh Banerjee,
Rui Liu,
Lili Su
Abstract:
Trajectory prediction describes the motions of surrounding moving obstacles for an autonomous vehicle; it plays a crucial role in enabling timely decision-making, such as collision avoidance and trajectory replanning. Accurate trajectory planning is the key to reliable vehicle deployments in open-world environment, where unstructured obstacles bring in uncertainties that are impossible to fully ca…
▽ More
Trajectory prediction describes the motions of surrounding moving obstacles for an autonomous vehicle; it plays a crucial role in enabling timely decision-making, such as collision avoidance and trajectory replanning. Accurate trajectory planning is the key to reliable vehicle deployments in open-world environment, where unstructured obstacles bring in uncertainties that are impossible to fully capture by training data. For traditional machine learning tasks, such uncertainties are often addressed reasonably well via methods such as continual learning. On the one hand, naively applying those methods to trajectory prediction can result in continuous data collection and frequent model updates, which can be resource-intensive. On the other hand, the predicted trajectories can be far away from the true trajectories, leading to unsafe decision-making. In this paper, we aim to establish real-time awareness of out-of-distribution in trajectory prediction for autonomous vehicles. We focus on the challenging and practically relevant setting where the out-of-distribution is deceptive, that is, the one not easily detectable by human intuition. Drawing on the well-established techniques of sequential analysis, we build real-time awareness of out-of-distribution by monitoring prediction errors using the quickest change point detection (QCD). Our solutions are lightweight and can handle the occurrence of out-of-distribution at any time during trajectory prediction inference. Experimental results on multiple real-world datasets using a benchmark trajectory prediction model demonstrate the effectiveness of our methods.
△ Less
Submitted 25 September, 2024;
originally announced September 2024.
-
Adaptive Self-Supervised Learning Strategies for Dynamic On-Device LLM Personalization
Authors:
Rafael Mendoza,
Isabella Cruz,
Richard Liu,
Aarav Deshmukh,
David Williams,
Jesscia Peng,
Rohan Iyer
Abstract:
Large language models (LLMs) have revolutionized how we interact with technology, but their personalization to individual user preferences remains a significant challenge, particularly in on-device applications. Traditional methods often depend heavily on labeled datasets and can be resource-intensive. To address these issues, we present Adaptive Self-Supervised Learning Strategies (ASLS), which u…
▽ More
Large language models (LLMs) have revolutionized how we interact with technology, but their personalization to individual user preferences remains a significant challenge, particularly in on-device applications. Traditional methods often depend heavily on labeled datasets and can be resource-intensive. To address these issues, we present Adaptive Self-Supervised Learning Strategies (ASLS), which utilizes self-supervised learning techniques to personalize LLMs dynamically. The framework comprises a user profiling layer for collecting interaction data and a neural adaptation layer for real-time model fine-tuning. This innovative approach enables continuous learning from user feedback, allowing the model to generate responses that align closely with user-specific contexts. The adaptive mechanisms of ASLS minimize computational demands and enhance personalization efficiency. Experimental results across various user scenarios illustrate the superior performance of ASLS in boosting user engagement and satisfaction, highlighting its potential to redefine LLMs as highly responsive and context-aware systems on-device.
△ Less
Submitted 25 September, 2024;
originally announced September 2024.
-
Reactive Multi-Robot Navigation in Outdoor Environments Through Uncertainty-Aware Active Learning of Human Preference Landscape
Authors:
Chao Huang,
Wenshuo Zang,
Carlo Pinciroli,
Zhi Jane Li,
Taposh Banerjee,
Lili Su,
Rui Liu
Abstract:
Compared with single robots, Multi-Robot Systems (MRS) can perform missions more efficiently due to the presence of multiple members with diverse capabilities. However, deploying an MRS in wide real-world environments is still challenging due to uncertain and various obstacles (e.g., building clusters and trees). With a limited understanding of environmental uncertainty on performance, an MRS cann…
▽ More
Compared with single robots, Multi-Robot Systems (MRS) can perform missions more efficiently due to the presence of multiple members with diverse capabilities. However, deploying an MRS in wide real-world environments is still challenging due to uncertain and various obstacles (e.g., building clusters and trees). With a limited understanding of environmental uncertainty on performance, an MRS cannot flexibly adjust its behaviors (e.g., teaming, load sharing, trajectory planning) to ensure both environment adaptation and task accomplishments. In this work, a novel joint preference landscape learning and behavior adjusting framework (PLBA) is designed. PLBA efficiently integrates real-time human guidance to MRS coordination and utilizes Sparse Variational Gaussian Processes with Varying Output Noise to quickly assess human preferences by leveraging spatial correlations between environment characteristics. An optimization-based behavior-adjusting method then safely adapts MRS behaviors to environments. To validate PLBA's effectiveness in MRS behavior adaption, a flood disaster search and rescue task was designed. 20 human users provided 1764 feedback based on human preferences obtained from MRS behaviors related to "task quality", "task progress", "robot safety". The prediction accuracy and adaptation speed results show the effectiveness of PLBA in preference learning and MRS behavior adaption.
△ Less
Submitted 24 September, 2024;
originally announced September 2024.
-
BARD: A seamless two-stage dose optimization design integrating backfill and adaptive randomization
Authors:
Yixuan Zhao,
Rachael Liu,
Jianchang Lin,
Ying Yuan
Abstract:
One common approach for dose optimization is a two-stage design, which initially conducts dose escalation to identify the maximum tolerated dose (MTD), followed by a randomization stage where patients are assigned to two or more doses to further assess and compare their risk-benefit profiles to identify the optimal dose. A limitation of this approach is its requirement for a relatively large sampl…
▽ More
One common approach for dose optimization is a two-stage design, which initially conducts dose escalation to identify the maximum tolerated dose (MTD), followed by a randomization stage where patients are assigned to two or more doses to further assess and compare their risk-benefit profiles to identify the optimal dose. A limitation of this approach is its requirement for a relatively large sample size. To address this challenge, we propose a seamless two-stage design, BARD (Backfill and Adaptive Randomization for Dose Optimization), which incorporates two key features to reduce sample size and shorten trial duration. The first feature is the integration of backfilling into the stage 1 dose escalation, enhancing patient enrollment and data generation without prolonging the trial. The second feature involves seamlessly combining patients treated in stage 1 with those in stage 2, enabled by covariate-adaptive randomization, to inform the optimal dose and thereby reduce the sample size. Our simulation study demonstrates that BARD reduces the sample size, improves the accuracy of identifying the optimal dose, and maintains covariate balance in randomization, allowing for unbiased comparisons between doses. BARD designs offer an efficient solution to meet the dose optimization requirements set by Project Optimus, with software freely available at www.trialdesign.org.
△ Less
Submitted 23 September, 2024;
originally announced September 2024.
-
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
Authors:
Xin Jiang,
Junwei Zheng,
Ruiping Liu,
Jiahang Li,
Jiaming Zhang,
Sven Matthiesen,
Rainer Stiefelhagen
Abstract:
As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our bench…
▽ More
As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework.
△ Less
Submitted 21 September, 2024;
originally announced September 2024.
-
GAInS: Gradient Anomaly-aware Biomedical Instance Segmentation
Authors:
Runsheng Liu,
Hao Jiang,
Yanning Zhou,
Huangjing Lin,
Liansheng Wang,
Hao Chen
Abstract:
Instance segmentation plays a vital role in the morphological quantification of biomedical entities such as tissues and cells, enabling precise identification and delineation of different structures. Current methods often address the challenges of touching, overlapping or crossing instances through individual modeling, while neglecting the intrinsic interrelation between these conditions. In this…
▽ More
Instance segmentation plays a vital role in the morphological quantification of biomedical entities such as tissues and cells, enabling precise identification and delineation of different structures. Current methods often address the challenges of touching, overlapping or crossing instances through individual modeling, while neglecting the intrinsic interrelation between these conditions. In this work, we propose a Gradient Anomaly-aware Biomedical Instance Segmentation approach (GAInS), which leverages instance gradient information to perceive local gradient anomaly regions, thus modeling the spatial relationship between instances and refining local region segmentation. Specifically, GAInS is firstly built on a Gradient Anomaly Mapping Module (GAMM), which encodes the radial fields of instances through window sliding to obtain instance gradient anomaly maps. To efficiently refine boundaries and regions with gradient anomaly attention, we propose an Adaptive Local Refinement Module (ALRM) with a gradient anomaly-aware loss function. Extensive comparisons and ablation experiments in three biomedical scenarios demonstrate that our proposed GAInS outperforms other state-of-the-art (SOTA) instance segmentation methods. The code is available at https://github.com/DeepGAInS/GAInS.
△ Less
Submitted 20 September, 2024;
originally announced September 2024.
-
Holistic and Historical Instance Comparison for Cervical Cell Detection
Authors:
Hao Jiang,
Runsheng Liu,
Yanning Zhou,
Huangjing Lin,
Hao Chen
Abstract:
Cytology screening from Papanicolaou (Pap) smears is a common and effective tool for the preventive clinical management of cervical cancer, where abnormal cell detection from whole slide images serves as the foundation for reporting cervical cytology. However, cervical cell detection remains challenging due to 1) hazily-defined cell types (e.g., ASC-US) with subtle morphological discrepancies caus…
▽ More
Cytology screening from Papanicolaou (Pap) smears is a common and effective tool for the preventive clinical management of cervical cancer, where abnormal cell detection from whole slide images serves as the foundation for reporting cervical cytology. However, cervical cell detection remains challenging due to 1) hazily-defined cell types (e.g., ASC-US) with subtle morphological discrepancies caused by the dynamic cancerization process, i.e., cell class ambiguity, and 2) imbalanced class distributions of clinical data may cause missed detection, especially for minor categories, i.e., cell class imbalance. To this end, we propose a holistic and historical instance comparison approach for cervical cell detection. Specifically, we first develop a holistic instance comparison scheme enforcing both RoI-level and class-level cell discrimination. This coarse-to-fine cell comparison encourages the model to learn foreground-distinguishable and class-wise representations. To emphatically improve the distinguishability of minor classes, we then introduce a historical instance comparison scheme with a confident sample selection-based memory bank, which involves comparing current embeddings with historical embeddings for better cell instance discrimination. Extensive experiments and analysis on two large-scale cytology datasets including 42,592 and 114,513 cervical cells demonstrate the effectiveness of our method. The code is available at https://github.com/hjiangaz/HERO.
△ Less
Submitted 20 September, 2024;
originally announced September 2024.
-
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
Authors:
Jiale Wei,
Junwei Zheng,
Ruiping Liu,
Jie Hu,
Jiaming Zhang,
Rainer Stiefelhagen
Abstract:
In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared with pinhole front-view images and panoramas. Traditional BEV methods, which rely on multiple narrow-field cameras and complex pose estimations, often face calibration and synchronization issues. To break the wall of the…
▽ More
In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared with pinhole front-view images and panoramas. Traditional BEV methods, which rely on multiple narrow-field cameras and complex pose estimations, often face calibration and synchronization issues. To break the wall of the aforementioned challenges, in this work, we introduce OneBEV, a novel BEV semantic mapping approach using merely a single panoramic image as input, simplifying the mapping process and reducing computational complexities. A distortion-aware module termed Mamba View Transformation (MVT) is specifically designed to handle the spatial distortions in panoramas, transforming front-view features into BEV features without leveraging traditional attention mechanisms. Apart from the efficient framework, we contribute two datasets, i.e., nuScenes-360 and DeepAccident-360, tailored for the OneBEV task. Experimental results showcase that OneBEV achieves state-of-the-art performance with 51.1% and 36.1% mIoU on nuScenes-360 and DeepAccident-360, respectively. This work advances BEV semantic mapping in autonomous driving, paving the way for more advanced and reliable autonomous systems.
△ Less
Submitted 20 September, 2024;
originally announced September 2024.