[{"type":"conference","title":"TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation","authors":[{"name":"Mingwei Li"},{"name":"Yi Yang"},{"name":"Hehe Fan"}],"institutions":["zgca"],"rawAffiliations":["1 College of Artificial Intelligence, Zhejiang University · 2 Zhongguancun Academy","1 Zhejiang University, 2 Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-09-06","year":2026,"venue":"ICML 2026","status":"published","topics":["cs.CV"],"abstract":"Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.","identifiers":{"arxiv":"2609.06665","doi":"10.48550/arXiv.2609.06665"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2609.06665"},{"label":"PDF","url":"https://arxiv.org/pdf/2609.06665"},{"label":"Code","url":"https://github.com/longxiang-ai/TransNormal-2"},{"label":"Model","url":"https://huggingface.co/black-forest-labs/FLUX.2-klein-base-9B"},{"label":"Dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2609.06665v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 College of Artificial Intelligence, Zhejiang University · 2 Zhongguancun Academy","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TransNormal-2"},{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, 2 Zhongguancun Academy","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TransNormal"}],"sources":["arXiv","Official GitHub project README"],"id":"doi-10-48550-arxiv-2609-06665","updatedAt":"2026-09-09T04:57:43Z"},{"type":"preprint","title":"How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition","authors":[{"name":"Hongyu Gu"},{"name":"Chang Liu"},{"name":"Jingwen Fu"}],"institutions":["zgca"],"rawAffiliations":["Hongyu Gu University of Science and Technology of China Hefei, China ustc_23ghy@mail.ustc.edu.cn Chang Liu Zhongguancun Academy and Technology of China Beijing, China liuchang@bza.edu.cn Jingwen Fu Zhongguancun Academy and Technology of China Beijing, China jwfu99@gmail.com † † thanks: Corresponding author. Abstract Large language models solve difficult problems by carrying intermediate computation across many reasoning steps. Conventional chain-of-thought writes this computation as tokens; recent continuous and recurrent approaches instead move part of it into fixed-dimensional latent states, where one thought can superpose several alternatives. This shift raises a basic design question: as reasoning proceeds, what should a continuous thought preserve? The natural strategy is to discard earlier computation and keep only the current frontier: storing more objects appears to dilute the st"],"relationType":"affiliation","publishedAt":"2026-09-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI"],"abstract":"Large language models solve hard problems through intermediate computations across multi-step reasoning. Traditional chain-of-thought encodes these computations as tokens. Recent continuous and recurrent methods instead move partial computations into fixed-dimensional latent states, where a single thought can superpose multiple alternatives. This raises a fundamental design question:what should continuous thoughts preserve as reasoning proceeds? An intuitive approach discards past computations and keeps only the current reasoning frontier. Storing more items seems to dilute states and waste limited representational capacity. We show this intuition can be incorrect. Under identical downstream computations, cumulative superposition retaining full reasoning history can require lower representational dimensions than frontier-only superposition holding only current alternatives. At fixed hidden width, this advantage allows latent reasoners to retain more valid evidence, distinguish more plausible downstream outcomes, and delay the point where compressed states turn unreliable. This counter-intuitive effect emerges because informative historical components coherently reinforce each other, while unrelated alternatives bring random interference. This perspective also answers a practical design question: how should models weight memories accumulated inside latent states when their future use is unknown? Across reusable weighted superpositions, prioritizing a small set of recent or salient items produces weakly-represented memories that bottleneck subsequent attention. Uniform cumulative weighting avoids this flaw, and we prove it is minimax-optimal for robust future reasoning. Our results turn superposition from an observed latent-space effect into a design principle: balanced cumulative memory lets a fixed representational budget support more reliable, reusable computations.","identifiers":{"arxiv":"2609.13747","doi":"10.48550/arXiv.2609.13747"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2609.13747"},{"label":"HTML","url":"https://arxiv.org/html/2609.13747v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2609.13747"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2609.13747v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hongyu Gu University of Science and Technology of China Hefei, China ustc_23ghy@mail.ustc.edu.cn Chang Liu Zhongguancun Academy and Technology of China Beijing, China liuchang@bza.edu.cn Jingwen Fu Zhongguancun Academy and Technology of China Beijing, China jwfu99@gmail.com † † thanks: Corresponding author. Abstract Large language models solve difficult problems by carrying intermediate computation across many reasoning steps. Conventional chain-of-thought writes this computation as tokens; recent continuous and recurrent approaches instead move part of it into fixed-dimensional latent states, where one thought can superpose several alternatives. This shift raises a basic design question: as reasoning proceeds, what should a continuous thought preserve? The natural strategy is to discard earlier computation and keep only the current frontier: storing more objects appears to dilute the st","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2609.13747v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2609-13747","updatedAt":"2026-09-28T06:11:42Z"},{"type":"preprint","title":"Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs","authors":[{"name":"Yang Yang"},{"name":"Jiawei Chen"},{"name":"Tairan Chen"},{"name":"Zhaoxia Yin"}],"institutions":["zgca"],"rawAffiliations":["Jiawei Chen Affiliation: Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-08-05","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through the reasoning chain and affect the final answer. Existing methods mainly improve spatial reasoning through training or additional spatial information, without considering whether the reasoning process itself is faithful to the model input. Our study shows that unfaithful reasoning chains significantly reduce final-answer accuracy. To address this issue, we propose a modular and training-free framework for spatial reasoning verification and correction. The framework constructs a Spatial Evidence Graph (SEG), which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence. Spatial Evidence Reliability Assessment (SERA) evaluates the reliability of visual evidence based on object existence, localization, and geometric measurements. The framework then identifies the earliest spatial evidence unit contradicted by reliable visual evidence and guides the original MLLM to revise the subsequent reasoning and final answer. Across 15 model-dataset settings, our method achieves an average accuracy of 68.94%, outperforming the compared baselines by 8.55 percentage points on average. Our code will be open-sourced.","identifiers":{"arxiv":"2608.04759","doi":"10.48550/arXiv.2608.04759"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2608.04759"},{"label":"HTML","url":"https://arxiv.org/html/2608.04759v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2608.04759"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2608.04759v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiawei Chen Affiliation: Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2608.04759v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2608-04759","updatedAt":"2026-08-13T14:57:08Z"},{"type":"preprint","title":"EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation","authors":[{"name":"Jiayi Luo"},{"name":"Hanxin Zhu"},{"name":"Chen Gao"},{"name":"Jiankun Wang"},{"name":"Cong Wang"},{"name":"Tianyu He"},{"name":"Jianxin Li"},{"name":"Zhibo Chen"}],"institutions":["zgca"],"rawAffiliations":["Jiayi Luo Affiliation: Beihang University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-08-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.","identifiers":{"arxiv":"2608.02990","doi":"10.48550/arXiv.2608.02990"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2608.02990"},{"label":"HTML","url":"https://arxiv.org/html/2608.02990v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2608.02990"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2608.02990v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiayi Luo Affiliation: Beihang University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2608.02990v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2608-02990","updatedAt":"2026-08-13T13:51:38Z"},{"type":"preprint","title":"RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?","authors":[{"name":"Hongjie Zhou"},{"name":"Shiqin Wang"},{"name":"Haoyang Chen"},{"name":"Haonan Guo"},{"name":"Di Wang"},{"name":"Juhua Liu"},{"name":"Fu Lin"},{"name":"Yong Luo"}],"institutions":["zgca"],"rawAffiliations":["Hongjie Zhou Affiliation: Fu Lin Affiliation: [3pt] Wuhan University Affiliation: Zhongguancun AcademyProject: https://HongjieZhou0329.github.io/RSVideo"],"relationType":"affiliation","publishedAt":"2026-08-03","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision-language models on continuous remote-sensing video understanding remains lacking. We introduce RSVideo-10K, a remote-sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo-Bench, contains 2,731 test instances and evaluates two complementary aspects of remote-sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision-language models still struggle to recover small local evidence, track short-lived states, and use scene-constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small-target spatiotemporal focusing that selects question-relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5-14B and attains the highest accuracy of 40.63% with Qwen3.6-27B across 26 open-source vision-language backbones. Codes will be available at https://github.com/HongjieZhou0329/RSVideo.","identifiers":{"arxiv":"2608.02039","doi":"10.48550/arXiv.2608.02039"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2608.02039"},{"label":"HTML","url":"https://arxiv.org/html/2608.02039v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2608.02039"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2608.02039v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hongjie Zhou Affiliation: Fu Lin Affiliation: [3pt] Wuhan University Affiliation: Zhongguancun AcademyProject: https://HongjieZhou0329.github.io/RSVideo","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2608.02039v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2608-02039","updatedAt":"2026-08-13T13:51:38Z"},{"type":"preprint","title":"CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition","authors":[{"name":"Lai Wei"},{"name":"Chengqi Li"},{"name":"Jiapeng Li"},{"name":"Ruina Hu"},{"name":"Yue Wang"},{"name":"Weiran Huang"}],"institutions":["zgca"],"rawAffiliations":["Lai Wei 1,2,∗ Chengqi Li 1,3,∗ Jiapeng Li 1,3,∗ Ruina Hu 2,∗ Yue Wang 2 Weiran Huang 1,3,† 1 School of Computer Science, Shanghai Jiao Tong University 2 Zhongguancun Academy 3 Shanghai Innovation Institute"],"relationType":"affiliation","publishedAt":"2026-07-28","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.","identifiers":{"arxiv":"2607.25294","doi":"10.48550/arXiv.2607.25294"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.25294"},{"label":"HTML","url":"https://arxiv.org/html/2607.25294v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.25294"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.25294v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lai Wei 1,2,∗ Chengqi Li 1,3,∗ Jiapeng Li 1,3,∗ Ruina Hu 2,∗ Yue Wang 2 Weiran Huang 1,3,† 1 School of Computer Science, Shanghai Jiao Tong University 2 Zhongguancun Academy 3 Shanghai Innovation Institute","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.25294v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-25294","updatedAt":"2026-09-25T05:30:40Z"},{"type":"preprint","title":"Distilled Reinforcement Learning for LLM Post-training","authors":[{"name":"Chen Wang"},{"name":"Zhaochun Li"},{"name":"Jionghao Bai"},{"name":"Yining Zhang"},{"name":"Hexuan Deng"},{"name":"Ge Lan"},{"name":"Yue Wang"}],"institutions":["zgca"],"rawAffiliations":["Chen Wang † † thanks: s-wc25@bza.edu.cn Affiliation: College of Elite Engineers, Nankai University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-07-19","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.","identifiers":{"arxiv":"2607.17247","doi":"10.48550/arXiv.2607.17247"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.17247"},{"label":"HTML","url":"https://arxiv.org/html/2607.17247v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.17247"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.17247v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chen Wang † † thanks: s-wc25@bza.edu.cn Affiliation: College of Elite Engineers, Nankai University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.17247v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-17247","updatedAt":"2026-09-23T05:18:13Z"},{"type":"preprint","title":"Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning","authors":[{"name":"Tieyuan Chen"},{"name":"Huabin Liu"},{"name":"Yi Wang"},{"name":"Chaofan Gan"},{"name":"Mingxi Lv"},{"name":"Ziran Qin"},{"name":"Shijie Li"},{"name":"Liquan Shen"},{"name":"Junhui Hou"},{"name":"Zheng Wang"},{"name":"Weiyao Lin"}],"institutions":["zgca"],"rawAffiliations":["1 Shanghai Jiao Tong University, 2 Zhongguancun Academy, 3 Shanghai AI Laboratory"],"relationType":"affiliation","publishedAt":"2026-07-19","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"","identifiers":{"arxiv":"2506.07811","doi":"10.48550/arXiv.2506.07811","publishedDoi":"10.1007/s11263-026-02925-w"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2506.07811"},{"label":"HTML","url":"https://arxiv.org/html/2506.07811v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2506.07811"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2506.07811v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Shanghai Jiao Tong University, 2 Zhongguancun Academy, 3 Shanghai AI Laboratory","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2506.07811v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2506-07811","updatedAt":"2026-08-08T18:30:46Z"},{"type":"preprint","title":"VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios","authors":[{"name":"Kailin Lyu"},{"name":"Long Xiao"},{"name":"Jianing Zeng"},{"name":"Di Wu"},{"name":"Lin Shu"},{"name":"Jie Hao"}],"institutions":["zgca"],"rawAffiliations":["Kailin Lyu Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-07-16","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic perception and human-machine interaction systems. However, existing methods depend on large-scale, diverse datasets from specific sensors and lack efficient data utilization and robust generalization capabilities, struggling in vision-limited environments. To address this, we introduce VQ-Touch, a tactile generation framework that supports both cross-sensor and multi-scenario applications. Specifically, to efficiently extract complex deformation and texture features from the data, we propose DM-VQGAN, an effective tactile representation learner. Furthermore, we introduce a discrete diffusion decoder with a unified conditioning interface, supporting multimodal generation tasks such as images and labels, and enhances the model's generalization capability through few-shot mixed training, thus achieving compatibility with current mainstream sensors and their variants. Experiments show that VQ-Touch surpasses state-of-the-art methods in multiple tasks.","identifiers":{"arxiv":"2607.14728","doi":"10.48550/arXiv.2607.14728"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.14728"},{"label":"HTML","url":"https://arxiv.org/html/2607.14728v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.14728"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.14728v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Kailin Lyu Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.14728v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-14728","updatedAt":"2026-09-23T05:18:13Z"},{"type":"preprint","title":"Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents","authors":[{"name":"Yixian Zhang"},{"name":"Huanming Zhang"},{"name":"Feng Gao"},{"name":"Xiao Li"},{"name":"Zhihao Liu"},{"name":"Chunyang Zhu"},{"name":"Jiaxing Qiu"},{"name":"Yuchen Yan"},{"name":"Jiyuan Liu"},{"name":"Wenhao Tang"},{"name":"Zhengru Fang"},{"name":"Yi Nie"},{"name":"Changxu Wei"},{"name":"Yu Wang"},{"name":"Wenbo Ding"},{"name":"C Shijia Yu"}],"institutions":["zgca"],"rawAffiliations":["Yixian Zhang 1,∗ Huanming Zhang 1,∗ Feng Gao 2 Xiao Li 3 Zhihao Liu 4 Chunyang Zhu 5 Jiaxing Qiu 5 Yuchen Yan 5 Jiyuan Liu 6 Wenhao Tang 1 Zhengru Fang 7 Yi Nie 1,2 Changxu Wei 1 Yu Wang 1 Wenbo Ding 1 Chao Yu 1,† 1 Tsinghua University 2 Striding AI 3 Purdue University 4 Institute of Automation, Chinese Academy of Sciences 5 Infinigence AI 6 Zhongguancun Academy 7 Hong Kong University of Science and Technology * Equal contribution. † Corresponding author: zoeyuchao@gmail.com Website: https://harnessvla.github.io/"],"relationType":"affiliation","publishedAt":"2026-07-09","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Language-conditioned manipulation requires both precise contact-rich control and robust reasoning over language, scenes, and long horizons. End-to-end Vision-Language-Action (VLA) models provide strong local visuomotor skills, but they are trained on in-distribution task trajectories and often fail under deployment perturbations such as semantic retargeting, goal re-binding, spatial-layout shifts, and unstable local contacts. LLM coding agents provide complementary semantic and compositional reasoning, but purely analytic primitives struggle with irregular grasping, constrained placement, and articulated-object interaction. We present Harness VLA, a memory-augmented agentic framework that exposes a frozen VLA as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without finetuning. Across perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation, Harness VLA improves over the strongest relevant baselines by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and reaches 58.4% on RoboTwin C2R. Code is available at https://github.com/RLinf/RPent.","identifiers":{"arxiv":"2607.08448","doi":"10.48550/arXiv.2607.08448"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.08448"},{"label":"HTML","url":"https://arxiv.org/html/2607.08448v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.08448"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.08448v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yixian Zhang 1,∗ Huanming Zhang 1,∗ Feng Gao 2 Xiao Li 3 Zhihao Liu 4 Chunyang Zhu 5 Jiaxing Qiu 5 Yuchen Yan 5 Jiyuan Liu 6 Wenhao Tang 1 Zhengru Fang 7 Yi Nie 1,2 Changxu Wei 1 Yu Wang 1 Wenbo Ding 1 Chao Yu 1,† 1 Tsinghua University 2 Striding AI 3 Purdue University 4 Institute of Automation, Chinese Academy of Sciences 5 Infinigence AI 6 Zhongguancun Academy 7 Hong Kong University of Science and Technology * Equal contribution. † Corresponding author: zoeyuchao@gmail.com Website: https://harnessvla.github.io/","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.08448v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-08448","updatedAt":"2026-09-22T05:34:24Z"},{"type":"conference","title":"TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios","authors":[{"name":"Kailin Lyu"},{"name":"Di Wu"},{"name":"Long Xiao"},{"name":"Jianning Zeng"},{"name":"Jianwei He"},{"name":"Chang Lin"},{"name":"Lianyu Hu"},{"name":"Lin Shu"},{"name":"Jie Hao"},{"name":"Ce Hao"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-07-06","year":2026,"venue":"IROS 2026","status":"published","topics":["cs.AI"],"abstract":"Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact and interaction in real-world environments. In this paper, we explore two key challenges of integrating tactile sensing into intelligent systems for multimodal reasoning: (i) insufficient modeling of dynamic tactile signals, which restricts reasoning over temporally evolving properties, and (ii) hallucination in tactile foundation models caused by the absence of explicit reasoning mechanisms, leading to unstable real-world inference. To address these challenges, we propose TacReasoner, a dynamic tactile-language framework for interactive reasoning in real-world scenarios. First, TacReasoner incorporates a Dynamic-aware Tactile Encoder to enhance the perception and representation of dynamic tactile signals. More importantly, we introduce TouchCoT-10k, the first tactile chain-of-thought dataset for structured reasoning over tactile inputs. Upon it, we establish DynTac-Bench to systematically evaluate dynamic tactile perception and real-world commonsense reasoning. Experimental results demonstrate that TacReasoner achieves competitive performance against state-of-the-art models across multiple datasets. Notably, despite using only 7B parameters, TacReasoner outperforms the 14B VTV-LLM model on most subtasks, highlighting its effectiveness and efficiency in tactile commonsense reasoning.","identifiers":{"arxiv":"2607.05131","doi":"10.48550/arXiv.2607.05131"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.05131"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.05131"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_9ag0vqmqfhud5uic6gwle37bsg2yz9hh"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.05131v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_9ag0vqmqfhud5uic6gwle37bsg2yz9hh"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2607-05131","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Don't Commit Alone: Joint Token Commitment in Diffusion Language Models","authors":[{"name":"L. Yao"}],"institutions":["zgca"],"rawAffiliations":["Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-07-05","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Diffusion language models (dLLMs) commit multiple tokens per denoising step by decoding each selected position independently from a shared context. When these positions are dependent, this factorization introduces an error captured by conditional total correlation, which confidence-based selection cannot infer from marginal probabilities alone. We propose CoCommit, a marker-gated coordination pass that delays commitment. After the usual bundle selection, a learned marker identifies the commit set, and the backbone's last n layers are re-applied to coordinate the marked positions before greedy argmax writes the tokens. This approximates joint-mode decoding while reusing existing weights, requiring only one partial forward pass and no auxiliary model. On LLaDA 2.1 with LoRA adapters and greedy inference, joint commitment improves five of the seven evaluated benchmarks over the released factorized decoder. The largest gains occur on code and reasoning tasks, while the remaining tasks are near parity.","identifiers":{"arxiv":"2607.04469","doi":"10.48550/arXiv.2607.04469"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.04469"},{"label":"HTML","url":"https://arxiv.org/html/2607.04469v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.04469"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.04469v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.04469v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-04469","updatedAt":"2026-09-21T05:33:39Z"},{"type":"preprint","title":"RL Forgets! Towards Continual Policy Optimization","authors":[{"name":"Mao-Lin Luo"},{"name":"Zhe-Xu Wang"},{"name":"Zi-Hao Zhou"},{"name":"Bo Ye"},{"name":"Jian Zhao"},{"name":"Min-Ling Zhang"},{"name":"Tong Wei"}],"institutions":["zgca","zgci"],"rawAffiliations":["Mao-Lin Luo 1 1 1 Equal contribution. , Zhe-Xu Wang 1 1 1 Equal contribution. , Zi-Hao Zhou 1 1 1 Equal contribution. , Bo Ye, Jian Zhao Affiliation: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China Affiliation: Key Laboratory of Computer Network and Information Integration (Southeast University) Ministry of Education, China Affiliation: Zhongguancun Academy Affiliation: Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-07-05","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent work has increasingly favored reinforcement learning over supervised fine-tuning, driven by the belief that reinforcement learning is inherently less prone to forgetting. However, the belief remains insufficiently validated, as existing evidence is largely drawn from outdated or homogeneous benchmarks. We revisit this assumption under recent and diverse multimodal reasoning tasks. To this end, we introduce MRCL, a Multimodal Reasoning Continual Learning benchmark. Experiments on MRCL show that standard reinforcement learning still suffers from severe catastrophic forgetting during continual post-training. We trace this failure to an objective mismatch: the KL regularization used in common policy optimization methods is evaluated on current-task data, whereas forgetting is caused by behavioral drift on prior-task distributions. To address this problem, we propose Continual Policy Optimization (CPO), a replay-free framework grounded in a prior-task behavioral KL objective. CPO relaxes the intractable historical KL constraint into sparse parameter-movement regularization, limiting policy drift without storing old data. Extensive experiments across multiple model scales show that CPO consistently reduces forgetting while preserving, and in some cases improving, pretrained model capabilities. On Qwen3-VL-8B, CPO reduces forgetting by 13.7% and improves pretrained capability by 7.0%. The implementation code is available at https://github.com/MaolinLuo/CPO.","identifiers":{"arxiv":"2607.04364","doi":"10.48550/arXiv.2607.04364"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.04364"},{"label":"HTML","url":"https://arxiv.org/html/2607.04364v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.04364"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.04364v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Mao-Lin Luo 1 1 1 Equal contribution. , Zhe-Xu Wang 1 1 1 Equal contribution. , Zi-Hao Zhou 1 1 1 Equal contribution. , Bo Ye, Jian Zhao Affiliation: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China Affiliation: Key Laboratory of Computer Network and Information Integration (Southeast University) Ministry of Education, China Affiliation: Zhongguancun Academy Affiliation: Zhongguancun Institute of Artificial Intelligence","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.04364v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Mao-Lin Luo 1 1 1 Equal contribution. , Zhe-Xu Wang 1 1 1 Equal contribution. , Zi-Hao Zhou 1 1 1 Equal contribution. , Bo Ye, Jian Zhao Affiliation: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China Affiliation: Key Laboratory of Computer Network and Information Integration (Southeast University) Ministry of Education, China Affiliation: Zhongguancun Academy Affiliation: Zhongguancun Institute of Artificial Intelligence","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.04364v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-04364","updatedAt":"2026-09-21T05:33:39Z"},{"type":"preprint","title":"WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection","authors":[{"name":"Xian Li"},{"name":"Yuanhe Tian"},{"name":"Yang Yang"},{"name":"Guoqing Wang"},{"name":"Y Song"}],"institutions":["zgca"],"rawAffiliations":["Xian Li Affiliation: University of Electronic Science and Technology of China Affiliation: Zhongguancun Academy Email: xianli@stu.uestc.edu.cn"],"relationType":"affiliation","publishedAt":"2026-07-05","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Online social media posts provide scalable signals for early depression screening, and recent studies mainly improve pre-classification evidence through risk-post selection, symptom grounding, and clinically informed feature construction. However, these screening-stage designs often leave final decisions to a single detector, overlooking how users heterogeneously express depressive risk after screening. A monolithic classifier must average across heterogeneous users, which may dilute localized evidence and cause misclassification, especially for non-self-disclosing users. To address this issue, we propose WPG-MoE, a weak-prior-guided dense mixture-of-experts framework built on a shared large language model (LLM) backbone. WPG-MoE derives user-level weak semantic priors to softly route users to experts matched to different evidence layouts. We formulate this process as learning using privileged information (LUPI): rich LLM-extracted structured evidence guides training-time routing, while inference retains only Patient Health Questionnaire-9 (PHQ-9) template screening and the deployable backbone. Experiments on Chinese and English datasets show that WPG-MoE outperforms strong baselines with interpretable routing behavior.","identifiers":{"arxiv":"2607.04350","doi":"10.48550/arXiv.2607.04350"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.04350"},{"label":"HTML","url":"https://arxiv.org/html/2607.04350v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.04350"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.04350v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xian Li Affiliation: University of Electronic Science and Technology of China Affiliation: Zhongguancun Academy Email: xianli@stu.uestc.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.04350v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-04350","updatedAt":"2026-09-21T05:33:39Z"},{"type":"preprint","title":"InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection","authors":[{"name":"Zichao Feng"},{"name":"Hao Zhu"},{"name":"Jingying Yang"},{"name":"S Xu"},{"name":"Yi Ren"},{"name":"Yuguang Yang"},{"name":"Xuhui Liu"},{"name":"Juan Zhang"},{"name":"Tian Wang"},{"name":"L Yang"},{"name":"Baochang Zhang"}],"institutions":["zgca"],"rawAffiliations":["Yangyang Ren 1,2 Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-07-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both RGB and infrared (IR) inputs, and treat them equally during both training and inference, which compromises their robustness when the RGB modality becomes unreliable or unavailable. In this case, we propose \\textbf{InfraNet}, an IR-centric quality-aware framework that regulates RGB guidance during training and supports flexible RGB--IR or IR-only deployment. InfraNet employs an asymmetric architecture where the primary IR pathway extracts multi-scale infrared features for predictions, while the auxiliary RGB pathway provides reliability-controlled supervisory signals. The core of InfraNet is \\textbf{QualGate}, a quality-aware fusion module that learns a task-oriented control signal to suppress unreliable RGB guidance and compensate IR features during cross-modal training. Built upon InfraNet, we design two architectural variants: a lightweight IR-only architecture InfraNet-IR and an RGB--IR architecture InfraNet-RGB-IR. Our method is evaluated through extensive experiments on four benchmark datasets (LLVIP, FLIR-Aligned, M$^3$FD, and DroneVehicle), showing strong or competitive accuracy in challenging low-light and adverse weather conditions. Notably, InfraNet maintains high efficiency in IR-only inference, making it both accurate and computationally efficient.","identifiers":{"arxiv":"2607.03795","doi":"10.48550/arXiv.2607.03795"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.03795"},{"label":"HTML","url":"https://arxiv.org/html/2607.03795v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.03795"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.03795v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yangyang Ren 1,2 Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.03795v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-03795","updatedAt":"2026-09-21T05:33:39Z"},{"type":"preprint","title":"SkillFab: An Agent-Native Skill Production Platform","authors":[{"name":"Anjie Xu"},{"name":"Yifeng Cai"},{"name":"Yi Li"},{"name":"Zixing Wang"},{"name":"Zhiyu Zhang"},{"name":"Jingfan Chen"},{"name":"Ruohan Xu"},{"name":"Leye Wang"}],"institutions":["zgca"],"rawAffiliations":["Anjie Xu Affiliation: Peking University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-07-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills. At runtime, agents first search for reusable skills; when no adequate skill exists, the unmet capability becomes a demand-first issue before any repository or implementation branch needs to exist. Development then proceeds through a SkillFab-managed repository, Git-ingested commit evidence, maintainer review, and registry publication. The same lifecycle is exposed through web, REST, and MCP surfaces, so humans, scripts, and external agents operate on shared state rather than separate task logs. The current system uses scoped Git push URLs, native range commit ingestion, workflow-state reads, and workflow-event histories to make long-running agent work reviewable and recoverable. We document the platform model, architecture, implemented capabilities, and three case studies: an end-to-end OS-detect skill run, a Docker research package that converts operational practice into reusable skill knowledge, and an external optimization case showing how improved skill artifacts can enter SkillFab as reviewable, versioned submissions. Deployment: https://skillfab.ai.","identifiers":{"arxiv":"2607.03780","doi":"10.48550/arXiv.2607.03780"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.03780"},{"label":"HTML","url":"https://arxiv.org/html/2607.03780v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.03780"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.03780v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Anjie Xu Affiliation: Peking University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.03780v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-03780","updatedAt":"2026-09-21T05:33:39Z"},{"type":"conference","title":"Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation","authors":[{"name":"Pengwei Zhang"},{"name":"Bin Xie"},{"name":"Xinpan Meng"},{"name":"Xinyu Guo"},{"name":"Ce Hao"},{"name":"Fang Deng"},{"name":"Long Cheng"},{"name":"Tiancai Wang"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面","Pengwei Zhang Affiliation: Pengwei Zhang, Xinpan Meng, and Long Cheng are with the School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China, and are also with the Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China. Affiliation: Pengwei Zhang, Ce Hao, Xinyu Guo, and Fang Deng are with Zhongguancun Academy, Beijing 100094, China."],"relationType":"official-output","publishedAt":"2026-07-03","year":2026,"venue":"IROS 2026","status":"published","topics":["cs.RO"],"abstract":"Tactile perception is indispensable for contact-rich manipulation, yet integrating it into Vision-Language-Action (VLA) models often induces modality collapse, where high-bandwidth visual features overshadow sparse tactile cues. Inspired by Predictive Coding, a neural mechanism where the brain attenuates predictable inputs to prioritize surprising stimuli, we propose ResTacVLA. Rather than treating tactile data as raw input, we reformulate it as a Residual Tactile Representation capturing the discrepancy between visual priors and physical sensations. By filtering out visually predictable dynamics, this formulation transforms sparse tactile signals into dense, high-value information gain, thereby inherently resolving the bandwidth mismatch. These residuals are discretized through a Vector Quantized (VQ) bottleneck into Latent Contact Primitives that capture critical events missed by vision. Analogous to the neural surprise signal, we leverage the uncertainty of the visual prior to adaptively gate tactile integration, prioritizing residuals specifically during visually unreliable phases to explicitly prevent visual dominance. Experimental results show that ResTacVLA consistently outperforms all baselines on a diverse set of contact-rich manipulation tasks, while remaining robust to unexpected dynamic disturbances. Project page: https://awilekong.github.io/ResTacVLA/","identifiers":{"arxiv":"2607.03387","doi":"10.48550/arXiv.2607.03387"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.03387"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.03387"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_ho6e7i6am89s329ptxsg21orx8s9qgv2"},{"label":"HTML","url":"https://arxiv.org/html/2607.03387v1"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.03387v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_ho6e7i6am89s329ptxsg21orx8s9qgv2"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Pengwei Zhang Affiliation: Pengwei Zhang, Xinpan Meng, and Long Cheng are with the School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China, and are also with the Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China. Affiliation: Pengwei Zhang, Ce Hao, Xinyu Guo, and Fang Deng are with Zhongguancun Academy, Beijing 100094, China.","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.03387v1"}],"sources":["arXiv","北京中关村学院官网","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-03387","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Flux-OPD: On-Policy Distillation with Evolving Contexts","authors":[{"name":"Yuran Wang"},{"name":"Zekun Wang"},{"name":"Bohan Zeng"},{"name":"Ruixu Zhang"},{"name":"Wenxuan Liu"},{"name":"Liu Yang"},{"name":"Yifan Dai"},{"name":"Yang Shi"},{"name":"Bozhou Li"},{"name":"Chengzhuo Tong"},{"name":"Daili Hua"},{"name":"Yuanxing Zhang"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Yuran Wang 1,2 , Zekun Wang 2 , Bohan Zeng 1 , Ruixu Zhang 3 , Wenxuan Liu 1 , Liu Yang 4 , Yifan Dai 4 , Yang Shi 1 , Bozhou Li 1 , Chengzhuo Tong 1 , Daili Hua 1 , Yuanxing Zhang 2 , Wentao Zhang 1,5,6 1 Peking University, 2 Kling Team, 3 Tsinghua University, 4 Shanghai Jiao Tong University, 5 Zhongguancun Academy, 6 Beijing Key Laboratory of Data Intelligence and Security (Peking University) † † thanks: Corresponding Author: wentao.zhang@pku.edu.cn"],"relationType":"affiliation","publishedAt":"2026-07-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI"],"abstract":"Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.","identifiers":{"arxiv":"2607.28022","doi":"10.48550/arXiv.2607.28022"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.28022"},{"label":"HTML","url":"https://arxiv.org/html/2607.28022v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.28022"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.28022v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuran Wang 1,2 , Zekun Wang 2 , Bohan Zeng 1 , Ruixu Zhang 3 , Wenxuan Liu 1 , Liu Yang 4 , Yifan Dai 4 , Yang Shi 1 , Bozhou Li 1 , Chengzhuo Tong 1 , Daili Hua 1 , Yuanxing Zhang 2 , Wentao Zhang 1,5,6 1 Peking University, 2 Kling Team, 3 Tsinghua University, 4 Shanghai Jiao Tong University, 5 Zhongguancun Academy, 6 Beijing Key Laboratory of Data Intelligence and Security (Peking University) † † thanks: Corresponding Author: wentao.zhang@pku.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.28022v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-28022","updatedAt":"2026-09-26T05:53:46Z"},{"type":"preprint","title":"Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning","authors":[{"name":"Haodong Zhu"},{"name":"Yangyang Ren"},{"name":"Yanjing Li"},{"name":"Sheng Xu"},{"name":"Haiguang Liu"},{"name":"Linlin Yang"},{"name":"Baochang Zhang"}],"institutions":["zgca"],"rawAffiliations":["Haodong Zhu † † thanks: Equal contribution. Affiliation: Beihang University Affiliation: Zhongguancun Academy Email: HaodongZhu@buaa.edu.cn"],"relationType":"affiliation","publishedAt":"2026-07-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.LG"],"abstract":"Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL's non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt's latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.","identifiers":{"arxiv":"2607.27610","doi":"10.48550/arXiv.2607.27610"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.27610"},{"label":"HTML","url":"https://arxiv.org/html/2607.27610v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.27610"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.27610v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haodong Zhu † † thanks: Equal contribution. Affiliation: Beihang University Affiliation: Zhongguancun Academy Email: HaodongZhu@buaa.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.27610v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-27610","updatedAt":"2026-09-26T05:53:46Z"},{"type":"preprint","title":"PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification","authors":[{"name":"Boya Zhang"},{"name":"Shuaiwen Zhou"},{"name":"Di Kong"},{"name":"Mingxu Wang"},{"name":"Wenbiao Du"},{"name":"Yiman Zhong"},{"name":"Yuexin Duan"},{"name":"Xiawei Yue"},{"name":"Liuquan Cheng"},{"name":"Xiru Li"}],"institutions":["zgca"],"rawAffiliations":["Xiru Li 2468li@sina.com organization=Nankai University, addressline=No. 94 Weijin Road, Nankai District, city=Tianjin, postcode=300071, country=China organization=Zhongguancun Academy, city=Beijing, country=China organization=The First Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Beijing University of Posts and Telecommunications, city=Beijing, country=China organization=The Six Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Tsinghua University, city=Beijing, country=China organization=Beijing Institute of Technology, city=Beijing, country=China organization=Beihang University, city=Beijing, country=China"],"relationType":"affiliation","publishedAt":"2026-07-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"Breast DCE-MRI AI is increasingly being explored for breast-level classification of no-lesion, benign, and malignant findings, beyond conventional lesion-centered diagnosis. Within this broader diagnostic scope, however, patient-specific background variability remains a major source of imaging confounding across classification tasks. Existing approaches predominantly focus on unilateral or lesion-centric analysis, whereas bilateral methods offer limited explicit modeling of spatially adaptive cross-breast correspondence. We propose PRISM-Net, a registration-free bilateral framework that leverages contralateral breast features as patient-specific references for background-aware representation learning. PRISM-Net integrates bilateral feature matching and asymmetry-aware attention to establish adaptive inter-breast correspondence and enhance representations of discriminative asymmetric patterns. On ODELIA, Macro AUC, Micro AUC, and quadratic weighted kappa were $84.11 \\pm 2.33$, $90.64 \\pm 1.61$, and $60.94 \\pm 5.64$ on the in-distribution test set, and $68.51 \\pm 4.54$, $80.74 \\pm 2.68$, and $43.45 \\pm 7.10$ on the held-out institution, respectively, outperforming the evaluated baseline methods across the primary evaluation metrics. PRISM-Net further demonstrated performance on independent institutional and background-complexity evaluations. Ablation experiments revealed that both bilateral relation modeling and asymmetry-aware reweighting contributed to improved classification performance. These findings highlight patient-specific bilateral reference modeling as a clinically grounded strategy for DCE-MRI interpretation, improving asymmetric pattern discrimination through explicit modeling of background complexity.","identifiers":{"arxiv":"2607.26799","doi":"10.48550/arXiv.2607.26799"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.26799"},{"label":"HTML","url":"https://arxiv.org/html/2607.26799v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.26799"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.26799v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xiru Li 2468li@sina.com organization=Nankai University, addressline=No. 94 Weijin Road, Nankai District, city=Tianjin, postcode=300071, country=China organization=Zhongguancun Academy, city=Beijing, country=China organization=The First Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Beijing University of Posts and Telecommunications, city=Beijing, country=China organization=The Six Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Tsinghua University, city=Beijing, country=China organization=Beijing Institute of Technology, city=Beijing, country=China organization=Beihang University, city=Beijing, country=China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.26799v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-26799","updatedAt":"2026-09-26T05:53:46Z"},{"type":"preprint","title":"FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification","authors":[{"name":"Zheming Fu"},{"name":"Ruizhe He"},{"name":"Wei Shang"},{"name":"Xiaoxiao Ma"},{"name":"Lei Wang"},{"name":"Chang Liu"},{"name":"Siming Fu"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-29","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which introduce training-inference inconsistencies and necessitates Classifier-Free Guidance (CFG). While implicit frameworks such as DiffusionNFT directly optimize forward-process velocity fields, its heuristic fixed-magnitude corrections prevent optimization strength from relative intra-group quality. We propose \\textit{Flow Advantage-Weighted Rectification} (\\textbf{FlowAWR}), a paradigm that recasts continuous generative policy optimization as supervised regression toward a theoretically optimal velocity field. Starting from the optimal policy of a KL-constrained reward maximization, FlowAWR derives the optimal velocity field that admits a magnitude-aware, advantage-weighted rectification form, yielding SDE-free optimization and CFG-free generation. In comparative evaluations on SD3.5-Medium, FlowAWR achieves improved alignment performance alongside a 2$\\times$ to 5$\\times$ convergence acceleration over DiffusionNFT (e.g., reaching a 24.12 PickScore in 1.2k steps, versus 23.82 in 2.0k steps for DiffusionNFT and 23.50 in $>$4k steps for FlowGRPO). Under multi-reward constraints, FlowAWR sustains generation quality, satisfying structural rules while maintaining stable out-of-domain performance.","identifiers":{"arxiv":"2606.30376","doi":"10.48550/arXiv.2606.30376"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.30376"},{"label":"HTML","url":"https://arxiv.org/html/2606.30376v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.30376"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.30376v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.30376v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-30376","updatedAt":"2026-09-20T05:28:00Z"},{"type":"preprint","title":"EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation","authors":[{"name":"Jie Zhao"},{"name":"Jie Feng"},{"name":"Can Rong"},{"name":"Zhihan Hou"},{"name":"Peng Lü"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Yong Li † † thanks: Jie Zhao and Yong Li are with the Department of Electronic Engineering, Tsinghua University, Beijing, China. (E-mail: liyong07@tsinghua.edu.cn). Jie Feng is with Zhongguancun Academy, Beijing, China. (E-mail: fengjie@bza.edu.cn). Can Rong is with Singapore-MIT Alliance for Research and Technology. Zhihan Hou is with the Xiuzhong College, Tsinghua University, Beijing, China. Peng Lu is with Department of Automation, Central South University, Changsha, China."],"relationType":"affiliation","publishedAt":"2026-06-26","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD models trained on routine mobility often degrade when extreme events abruptly alter regional functions and population activities, while retraining a new generator for each event is impractical under limited event-time supervision. We propose EventOD, an event-adaptive OD generation framework that steers a pretrained OD generator using structured event semantics. EventOD first uses a large language model to infer region-level functional and demographic control vectors from coarse event observations. It then learns two lightweight adaptation modules, AlphaNet and BetaNet, to calibrate the magnitude of these semantic shifts, and further introduces a retrieval-augmented fallback pathway for scenarios with sparse supervision. The resulting event-conditioned features are injected into a pretrained graph diffusion OD model through input-level modulation, enabling event-aware adaptation without updating generator parameters. Experiments on hurricane- and pandemic-induced mobility across U.S. counties show that EventOD consistently improves both reconstruction accuracy and distributional fidelity over strong baselines. Source code is available at https://anonymous.4open.science/r/EventOD-5C11/.","identifiers":{"arxiv":"2607.22655","doi":"10.48550/arXiv.2607.22655"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.22655"},{"label":"HTML","url":"https://arxiv.org/html/2607.22655v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.22655"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.22655v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yong Li † † thanks: Jie Zhao and Yong Li are with the Department of Electronic Engineering, Tsinghua University, Beijing, China. (E-mail: liyong07@tsinghua.edu.cn). Jie Feng is with Zhongguancun Academy, Beijing, China. (E-mail: fengjie@bza.edu.cn). Can Rong is with Singapore-MIT Alliance for Research and Technology. Zhihan Hou is with the Xiuzhong College, Tsinghua University, Beijing, China. Peng Lu is with Department of Automation, Central South University, Changsha, China.","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.22655v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-22655","updatedAt":"2026-09-25T05:30:40Z"},{"type":"preprint","title":"Self-Generated Error Training for Token Editing in Diffusion Language Models","authors":[{"name":"Lin Yao"}],"institutions":["zgca"],"rawAffiliations":["Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-06-15","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Token-to-token (T2T) editing lets LLaDA2.1 revise committed tokens during block-diffusion decoding. The released recipe trains this editor on random vocabulary corruptions, but at inference the editor sees the model's own fluent, high-confidence draft errors instead. We study this training-inference mismatch and propose self-generated T2T, which performs a no-gradient draft pass, fills masked positions with predicted tokens, and supervises recovery in a second pass under these self-generated corruptions. We implement the update as a short LoRA continued-pretraining pass on LLaDA2.1-mini and evaluate on several benchmarks under the official Q-Mode T2T procedure with unchanged inference parameters. The method generally improves accuracy while reducing T2T edit intensity, mitigating failure modes such as final-digit transcription errors after otherwise correct reasoning and excessive self-correction before short factual answers.","identifiers":{"arxiv":"2606.17175","doi":"10.48550/arXiv.2606.17175"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.17175"},{"label":"HTML","url":"https://arxiv.org/html/2606.17175v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.17175"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.17175v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.17175v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-17175","updatedAt":"2026-09-17T05:26:52Z"},{"type":"preprint","title":"MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs","authors":[{"name":"Yuanteng Chen"},{"name":"Nanxin Zeng"},{"name":"Peisong Wang"},{"name":"Zhilei Liu"},{"name":"Yuantian Shao"},{"name":"Shiqiang Lang"},{"name":"T Liu"},{"name":"Chuangyi Li"},{"name":"Qinghao Hu"},{"name":"Gang Li"},{"name":"Jing Liu"},{"name":"Jian Cheng"}],"institutions":["zgca"],"rawAffiliations":["Yuanteng Chen Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Zhongguancun Academy Email: peisong.wang@nlpr.ia.ac.cn"],"relationType":"affiliation","publishedAt":"2026-06-15","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. Among PTQ methods, expert-level mixed-precision quantization has proven effective for MoE-LLMs, yet suffers notable degradation on MoE-MLLMs due to two overlooked biases in expert importance estimation. (1) At the cross-modal level, the numerical dominance of vision tokens causes expert selection frequency to be dominated by vision tokens, masking experts that are critical to the text modality; (2) at the intra-vision level, the large proportion of redundant vision tokens further skew frequency statistics, obscuring experts critical for informative visual content. To bridge gaps, we propose MODE, a modality-decomposed expert-level mixed-precision quantization framework for MoE-MLLMs that decomposes expert selection frequency by modality, filters redundant vision tokens to obtain denoised visual frequency, and further evaluates quantization sensitivity per modality as a complementary signal to frequency-based estimation. These signals are integrated into an Integer Linear Programming formulation to assign per-expert bit-widths under a given budget. Extensive experiments show that MODE is particularly well-suited for MoE-MLLMs, limiting average performance loss to within 2.9% at W3A16, with larger gains at the extreme 2-bit setting.","identifiers":{"arxiv":"2606.17118","doi":"10.48550/arXiv.2606.17118"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.17118"},{"label":"HTML","url":"https://arxiv.org/html/2606.17118v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.17118"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.17118v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuanteng Chen Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Zhongguancun Academy Email: peisong.wang@nlpr.ia.ac.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.17118v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-17118","updatedAt":"2026-09-17T05:26:52Z"},{"type":"preprint","title":"DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents","authors":[{"name":"Minghang Zhu"},{"name":"Chuyang Wei"},{"name":"Junhao Xu"},{"name":"Yilin Cheng"},{"name":"Zhumin Chen"},{"name":"Jiyan He"}],"institutions":["zgca"],"rawAffiliations":["Minghang Zhu Chuyang Wei Junhao Xu Yilin Cheng Affiliation: Shandong University, Qingdao, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Fudan University, Shanghai, China Email: mhzhu@mail.sdu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-06-15","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Deep research agents synthesize long-form reports by searching and reasoning over retrieved evidence. Reinforcement learning with rubric-based rewards improves these agents by optimizing them against checkable criteria that translate report quality into reward signals, but its efficiency depends on whether those criteria reliably capture the task scope and evidence needs. Most existing studies ask an LLM to generate rubrics for a given query, but when the model fails to infer the underlying information needs, the generated rubrics may be incomplete and reduce RL efficiency. To obtain more reliable query--rubric supervision, we introduce DeepRubric, a data construction framework that reverses this process: instead of inferring evaluation criteria for a given query, it first determines what an evidence-backed report should be evaluated on and then synthesizes aligned query--rubric pairs from those evaluation targets. Starting from a sampled seed topic, DeepRubric builds an evidence tree by recursively expanding evidence-backed sub-questions, whose leaves serve as atomic and verifiable evaluation targets. It then uses the evidence tree to synthesize the training query and rubrics, ensuring that the reward evaluates exactly the information requested by the query. Using DeepRubric, we construct 9K query--rubric supervision examples and train DeepRubric-8B with rubric-based GRPO, achieving comparable performance to prior open state-of-the-art deep research models across three benchmarks with roughly 13x fewer RL GPU-hours.","identifiers":{"arxiv":"2606.17029","doi":"10.48550/arXiv.2606.17029"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.17029"},{"label":"HTML","url":"https://arxiv.org/html/2606.17029v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.17029"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.17029v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Minghang Zhu Chuyang Wei Junhao Xu Yilin Cheng Affiliation: Shandong University, Qingdao, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Fudan University, Shanghai, China Email: mhzhu@mail.sdu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.17029v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-17029","updatedAt":"2026-09-17T05:26:52Z"},{"type":"conference","title":"UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics","authors":[{"name":"Yanxin Xi"},{"name":"Xiang Su"},{"name":"Jie Feng"},{"name":"Yu Liu"},{"name":"Sasu Tarkoma"},{"name":"Pan Hui"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-06-14","year":2026,"venue":"KDD Datasets and Benchmarks Track 2026","status":"published","topics":["cs.AI"],"abstract":"Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWell, a large-scale benchmark designed to systematically evaluate the spatio-temporal reasoning capabilities of MLLMs for urban wellbeing analytics through joint modeling of satellite and street view imagery. UrbanWell spans 38 cities across multiple years and includes diverse indicators covering (1) environmental conditions (CO$_2$, NO$_2$, PM${2.5}$, and Normalized Difference Vegetation Index), (2) spatial accessibility (minimum distance to supermarkets and restaurants), (3) urban form (road length, road density, and land use), (4) urban vitality (population, economic activity diversity, and land use diversity), and (5) subjective perception attributes (e.g., safety, beauty, liveliness, wealth, and quietness). All indicators are aligned at grid level to enable standardized evaluation. Beyond static prediction, UrbanWell defines temporal reasoning tasks, including future value forecasting from historical observations and temporal trend classification. We benchmark 15 state-of-the-art representative MLLMs in a zero-shot setting, providing a comprehensive comparative evaluation across spatial and temporal dimensions. Experimental results indicate that while MLLMs capture salient spatial and perceptual cues, their performance varies substantially across heterogeneous urban indicators spanning environment and subjective perception. UrbanWell serves as a unified benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics, offering a standardized testbed for systematic assessment and future research on multimodal urban intelligence. Our codes and datasets are accessible via https://github.com/axin1301/UrbanWell-Benchmark.","identifiers":{"arxiv":"2606.15890","doi":"10.48550/arXiv.2606.15890"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.15890"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.15890"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_nq1x2ai9mrehwwm692ivxl957r4vvluw"},{"label":"Code","url":"https://github.com/axin1301/UrbanWell-Benchmark/"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.15890v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_nq1x2ai9mrehwwm692ivxl957r4vvluw"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2606-15890","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting","authors":[{"name":"Chen Su"},{"name":"Yuanhe Tian"},{"name":"Yan Song"}],"institutions":["zgca"],"rawAffiliations":["Yuanhe Tian Affiliation: University of Science and Technology of China Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-12","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose DiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. DiffDiff makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, DiffDiff outperforms six diffusion baselines, and our analysis confirms that DiffDiff concentrates the diffusion's generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.","identifiers":{"arxiv":"2607.22599","doi":"10.48550/arXiv.2607.22599"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.22599"},{"label":"HTML","url":"https://arxiv.org/html/2607.22599v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.22599"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.22599v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuanhe Tian Affiliation: University of Science and Technology of China Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.22599v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-22599","updatedAt":"2026-08-13T13:11:50Z"},{"type":"preprint","title":"VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification","authors":[{"name":"Xiaoxian Duan"},{"name":"Zequn Liu"},{"name":"Yingce Xia"}],"institutions":["zgca"],"rawAffiliations":["Xiaoxian Duan Affiliation: Institute of Automation, Chinese Academy of Sciences, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China Email: duanxiaoxian2026@ia.ac.cn"],"relationType":"affiliation","publishedAt":"2026-06-12","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Geometry problem generation is useful for AI-assisted education and multimodal mathematical reasoning, but reliable synthesis remains difficult because the problem statement, diagram, constraints, and solution should be mutually consistent. Existing methods often trade off controllability and reliability: seed-based rewriting is flexible but weakly verifiable, whereas diagram-first construction improves validity but is less suited to arbitrary user-specified constraints. We introduce VeriGeo, a controllable geometry generation framework grounded in executable reasoning traces. Given user constraints such as target concepts and difficulty, an Author agent generates a problem and diagram, and a Solver agent produces a proof-aligned solution. Both agents use a shared action sequence that connects natural language, diagrams, geometric constraints, and proof steps into a verifiable representation. A three-stage pipeline checks numerical consistency, analytical realizability, and global consistency, using verification-guided reflection to repair recoverable failures and reject unrecoverable ones. Across five LLM backbones, raw generations frequently fail these checks, while VeriGeo repairs a substantial fraction of the invalid attempts. Supervised fine-tuning on 8.7k examples generated by VeriGeo achieves the best reported GeoQA performance among end-to-end multimodal LLM-based solvers, and obtains strong results on PGPS9K and MathVista-GPS, demonstrating the effectiveness of verified synthetic data for improving multimodal geometry reasoning.","identifiers":{"arxiv":"2606.14176","doi":"10.48550/arXiv.2606.14176","publishedDoi":"10.48550/arxiv.2606.14176"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.14176"},{"label":"HTML","url":"https://arxiv.org/html/2606.14176v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.14176"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.14176v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xiaoxian Duan Affiliation: Institute of Automation, Chinese Academy of Sciences, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China Email: duanxiaoxian2026@ia.ac.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.14176v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-14176","updatedAt":"2026-10-01T06:38:45Z"},{"type":"preprint","title":"Position: The Systemic Lack of Agency in Visual Reasoning","authors":[{"name":"Yizhao Huang"},{"name":"Haoyang Chen"},{"name":"Shiqin Wang"},{"name":"Pohsun Huang"},{"name":"Jiayuan Li"},{"name":"Haoyuan Du"},{"name":"Yandong Shi"},{"name":"Zheng Wang"},{"name":"Zhixiang Wang"}],"institutions":["zgca"],"rawAffiliations":["Haoyang Chen Affiliation: National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University Affiliation: Hubei Key Laboratory of Multimedia and Network Communication Engineering Affiliation: Zhongguancun Academy, Beijing, China. 100094"],"relationType":"affiliation","publishedAt":"2026-06-11","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual evidence to bridge information gaps, rather than merely relying on explicitly specified targets. This capacity underlies human visual understanding and everyday reasoning. We argue that this limitation arises from a tendency to approach visual reasoning primarily as passive semantic retrieval, rather than as active, situated reasoning that depends on autonomous visual exploration. As a result, most existing benchmarks primarily assess Passive Capacity, leaving this aspect of reasoning largely unmeasured. To address this gap, we introduce the Visual Implicit Reasoning Diagnosing Benchmark (V-IRD), which targets this missing quadrant by requiring models to derive answers strictly through autonomous visual analysis. Our results show that, despite strong retrieval abilities, prominent VLMs struggle to utilize reference objects and to attend to visual evidence that requires self-directed inquiry. Simply put, strong semantic recognition does not equate to active visual exploration, revealing a critical gap in current VLMs. More information can be found at https://haoychen.github.io/Implicit-Reasoning/","identifiers":{"arxiv":"2606.14795","doi":"10.48550/arXiv.2606.14795"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.14795"},{"label":"HTML","url":"https://arxiv.org/html/2606.14795v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.14795"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.14795v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haoyang Chen Affiliation: National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University Affiliation: Hubei Key Laboratory of Multimedia and Network Communication Engineering Affiliation: Zhongguancun Academy, Beijing, China. 100094","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.14795v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-14795","updatedAt":"2026-09-17T05:26:52Z"},{"type":"preprint","title":"GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs","authors":[{"name":"Hengyi Feng"},{"name":"Zeang Sheng"},{"name":"Meiyi Qiang"},{"name":"Li Yang"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Hengyi Feng 1,2* , Zeang Sheng 2* , Meiyi Qiang 2 , Yang Li 3 , Wentao Zhang 2,4 † \\dagger † † thanks: * Equal contribution. † † thanks: $†$ Corresponding author. Affiliation: Affiliation: 1 University of Electronic Science and Technology of China 2 Peking University 3 Tencent Inc 4 Zhongguancun Academy hengyi.feng@std.uestc.edu.cn, {shengzeang18, wentao.zhang}@pku.edu.cn Abstract Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation networks, e-commerce platforms, social media, and web pages. Inspired by the remarkable semantic understanding ability of Large Language Models (LLMs), there have been numerous attempts to integrate LLMs into TAGs. However, existing methods still struggle to generalize across diverse graphs and tasks, and their ability to capture transferabl"],"relationType":"affiliation","publishedAt":"2026-06-10","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation networks, e-commerce platforms, social media, and web pages. Inspired by the remarkable semantic understanding ability of Large Language Models (LLMs), there have been numerous attempts to integrate LLMs into TAGs. However, existing methods still struggle to generalize across diverse graphs and tasks, and their ability to capture transferable graph structural patterns remains limited. To address this, we introduce the GraspLLM, a framework that combines Graph structural comprehension with semantic understanding prowess of LLMs to enhance the cross-dataset and cross-task generalizability. Specifically, we represent node texts from different graphs in a unified semantic space with a frozen general embedding model, on top of which we perform motif-aware contrastive learning across multiple motif-induced adjacency matrices to extract dataset-agnostic structural information. Then, with our proposed optimal contextual subgraph, we extract the most contextually relevant subgraph for each target node and align these subgraphs to the token space of LLM via an alignment projector. Extensive experiments on TAG benchmark datasets spanning diverse domains reveal that GraspLLM consistently outperforms previous LLM-based methods for TAGs, especially in zero-shot scenarios, highlighting its strong generalizability across different datasets and tasks. Our code is available at https://github.com/Heinz217/GraspLLM.","identifiers":{"arxiv":"2606.11898","doi":"10.48550/arXiv.2606.11898"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.11898"},{"label":"HTML","url":"https://arxiv.org/html/2606.11898v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.11898"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.11898v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hengyi Feng 1,2* , Zeang Sheng 2* , Meiyi Qiang 2 , Yang Li 3 , Wentao Zhang 2,4 † \\dagger † † thanks: * Equal contribution. † † thanks: $†$ Corresponding author. Affiliation: Affiliation: 1 University of Electronic Science and Technology of China 2 Peking University 3 Tencent Inc 4 Zhongguancun Academy hengyi.feng@std.uestc.edu.cn, {shengzeang18, wentao.zhang}@pku.edu.cn Abstract Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation networks, e-commerce platforms, social media, and web pages. Inspired by the remarkable semantic understanding ability of Large Language Models (LLMs), there have been numerous attempts to integrate LLMs into TAGs. However, existing methods still struggle to generalize across diverse graphs and tasks, and their ability to capture transferabl","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.11898v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-11898","updatedAt":"2026-09-16T05:23:59Z"},{"type":"preprint","title":"TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation","authors":[{"name":"Kailin Lyu"},{"name":"Di Wu"},{"name":"Pengwei Zhang"},{"name":"Yuhang Zheng"},{"name":"Yingxin Lai"},{"name":"Long Xiao"},{"name":"Kangyi Wu"},{"name":"Pengna Li"},{"name":"Chen Gao"},{"name":"Lianyu Hu"},{"name":"Xiaobin Hu"},{"name":"Jie Hao"},{"name":"Ce Hao"},{"name":"Weihao Yuan"},{"name":"Su Yan"}],"institutions":["zgca"],"rawAffiliations":["Kailin Lyu Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-10","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering 415 objects, 8 scenarios, and 7 sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.","identifiers":{"arxiv":"2606.11637","doi":"10.48550/arXiv.2606.11637"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.11637"},{"label":"HTML","url":"https://arxiv.org/html/2606.11637v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.11637"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.11637v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Kailin Lyu Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.11637v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-11637","updatedAt":"2026-09-16T05:23:59Z"},{"type":"preprint","title":"DynaOD: Dynamic Origin-Destination Flow Generation with Discrete-to-Continuous Temporal Semantic Modeling","authors":[{"name":"Jie Zhao"},{"name":"Xianqi Dai"},{"name":"Jie Feng"},{"name":"Huandong Wang"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Jie Feng Note: Corresponding author. Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-08","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Dynamic origin-destination (OD) flow generation seeks to synthesize realistic mobility dynamics from temporal context alone, without relying on historical OD observations. A key challenge is to translate semantic temporal signals into temporally coherent OD patterns while preserving the inherent spatial heterogeneity of urban regions. We propose DynaOD, a semantic-driven framework that models temporal dynamics through two complementary perspectives: discrete directional trends that characterize qualitative shifts in urban activity patterns, and continuous temporal evolution that captures how such shifts unfold over time. By jointly encoding these temporal semantics, the framework constructs time-varying region representations that condition pretrained static OD generators in a lightweight and plug-and-play fashion. This modular design further supports scalable deployment and cross-city transferability. Extensive experiments on large-scale real-world datasets show that our method consistently outperforms representative baselines in both predictive accuracy and distributional fidelity. Code is publicly available at https://github.com/csjiezhao/DynaOD.","identifiers":{"arxiv":"2606.09086","doi":"10.48550/arXiv.2606.09086"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.09086"},{"label":"HTML","url":"https://arxiv.org/html/2606.09086v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.09086"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.09086v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jie Feng Note: Corresponding author. Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.09086v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-09086","updatedAt":"2026-09-15T05:29:28Z"},{"type":"preprint","title":"EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control","authors":[{"name":"Haoyang Ge"},{"name":"Peng Ren"},{"name":"Yukun Shi"},{"name":"Cong Huang"},{"name":"Kun Li"},{"name":"Kai Chen"}],"institutions":["zgca","zgci"],"rawAffiliations":["Haoyang Ge Peng Ren Yukun Shi Cong Huang Kun Li Kai Chen Email: ∗ ghy0623@tju.edu.cn † Corresponding authors: Kun Li and Kai Chen Affiliation: Tianjin University, Tianjin, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Beihang University, Beijing, China Affiliation: Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: DeepCybo, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-06-07","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the same checkpoint. Experiments on Nymeria and EgoExo4D show that one checkpoint improves egocentric motion generation over UniEgoMotion while supporting reconstruction and forecasting; the generated SMPL motions can also be executed by a Unitree humanoid controller. These results indicate a practical path from scalable egocentric observations to generalizable and interactive humanoid motion priors.","identifiers":{"arxiv":"2606.08495","doi":"10.48550/arXiv.2606.08495"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.08495"},{"label":"HTML","url":"https://arxiv.org/html/2606.08495v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.08495"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.08495v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haoyang Ge Peng Ren Yukun Shi Cong Huang Kun Li Kai Chen Email: ∗ ghy0623@tju.edu.cn † Corresponding authors: Kun Li and Kai Chen Affiliation: Tianjin University, Tianjin, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Beihang University, Beijing, China Affiliation: Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: DeepCybo, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.08495v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Haoyang Ge Peng Ren Yukun Shi Cong Huang Kun Li Kai Chen Email: ∗ ghy0623@tju.edu.cn † Corresponding authors: Kun Li and Kai Chen Affiliation: Tianjin University, Tianjin, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Beihang University, Beijing, China Affiliation: Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: DeepCybo, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.08495v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-08495","updatedAt":"2026-09-15T05:29:28Z"},{"type":"preprint","title":"Towards World Models in Biomedical Research","authors":[{"name":"Guangyu Wang"},{"name":"Jingkun Yue"},{"name":"Siqi Zhang"},{"name":"Yu Liu (6938)"},{"name":"Xiaoyu Wang (182979)"},{"name":"Mingyuan Meng"},{"name":"Changwei Ji"},{"name":"Zongbo Han"},{"name":"Yulin Wang"},{"name":"Yang Yue"},{"name":"Frank Fu"},{"name":"Ting Chen"},{"name":"Song Wu"},{"name":"Ziwei Liu"},{"name":"Jiangning Song"},{"name":"Ming Li"},{"name":"Gao Huang"},{"name":"Xiaohong Liu (50816)"},{"name":"Athanasios V. Vasilakos"},{"name":"Xingcai Zhang"},{"name":"Ping Zhang (86495)"},{"name":"Yong Li"}],"institutions":["zgca","zgci"],"rawAffiliations":["Affiliation: Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China Affiliation: Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA Affiliation: Nanyang Technological University, Singapore Affiliation: Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia Affiliation: David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada Affiliation: Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway Affiliation: Corresponding authors: guangyu.wang24@gmail.com, yuejk@bupt.edu.cn"],"relationType":"affiliation","publishedAt":"2026-06-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbations, disease progression and therapeutic intervention. Although foundation models and large language models have accelerated biomedical data interpretation, most current systems remain focused on static pattern recognition rather than prospective simulation of biological futures. Here we propose biomedical world models as a paradigm for AI-driven discovery. These models learn latent representations of molecular, cellular, tissue and clinical states, together with intervention-conditioned dynamics that allow future trajectories to be simulated before actions are taken. We discuss how biomedical world models could function as data engines, environment simulators and scientific planning substrates across applications including virtual cells, organoids, virtual patients and surgical simulation. We outline the data infrastructure, evaluation benchmarks, safety constraints and governance frameworks required. Biomedical world models may provide a foundation for simulation-guided, closed-loop and experimentally actionable biomedical discovery.","identifiers":{"arxiv":"2606.05925","doi":"10.48550/arXiv.2606.05925"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.05925"},{"label":"HTML","url":"https://arxiv.org/html/2606.05925v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.05925"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.05925v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Affiliation: Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China Affiliation: Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA Affiliation: Nanyang Technological University, Singapore Affiliation: Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia Affiliation: David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada Affiliation: Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway Affiliation: Corresponding authors: guangyu.wang24@gmail.com, yuejk@bupt.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.05925v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Affiliation: Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China Affiliation: Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China Affiliation: Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA Affiliation: Nanyang Technological University, Singapore Affiliation: Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia Affiliation: David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada Affiliation: Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway Affiliation: Corresponding authors: guangyu.wang24@gmail.com, yuejk@bupt.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.05925v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-05925","updatedAt":"2026-09-15T05:29:28Z"},{"type":"preprint","title":"Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding","authors":[{"name":"Qiuyu Tian"},{"name":"Fengyi Chen"},{"name":"Yiding Li"},{"name":"Youyong Kong"},{"name":"Fan Guo"},{"name":"Yuyao Li"},{"name":"Jinjing Shen"},{"name":"Zhijing Xie"},{"name":"Yiyun Luo"},{"name":"Xin Zhang (35492)"},{"name":"Yingce Xia"},{"name":"Zequn Liu"}],"institutions":["zgca"],"rawAffiliations":["Qiuyu Tian 1,2 Fengyi Chen 3 Yiding Li 5 Youyong Kong 1 Fan Guo 5 Yuyao Li 5 Jinjing Shen 5 Zhijing Xie 5 Yiyun Luo 5 Xin Zhang 5 Yingce Xia 2 Zequn Liu 2 1 Southeast University, Nanjing, China 2 Beijing Zhongguancun Academy, Beijing, China 3 Nanjing Normal University, Nanjing, China 5 ZhuiWen Technology Co., Ltd., Beijing, China"],"relationType":"affiliation","publishedAt":"2026-06-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing character states, social relations, causal triggers, temporal position, and later consequences. Existing retrieval and graph-augmented generation methods improve evidence access, but their units--chunks, entities, relations, summaries, or tool actions--do not directly encode how evidence functions in a story. We introduce Narrative Knowledge Weaver(NKW), a source-grounded framework that aligns textual evidence, atomic facts, canonical graph structure, entity profiles, interactions, episodes, and storylines. At query time, NKW uses text, graph, and narrative tools with post-retrieval reading skills to assemble evidence and audit actor, scope, polarity, state, and temporal constraints. Across STAGE, FairytaleQA, and QuALITY, NKW is strongest on screenplay-level story-world QA while remaining competitive on more passage-centered benchmarks. Ablations, question-type analyses, graph-asset statistics, and case studies show complementary benefits for character, scene, temporal, causal, and narrative-progression reasoning.","identifiers":{"arxiv":"2606.05724","doi":"10.48550/arXiv.2606.05724"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.05724"},{"label":"HTML","url":"https://arxiv.org/html/2606.05724v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.05724"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.05724v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Qiuyu Tian 1,2 Fengyi Chen 3 Yiding Li 5 Youyong Kong 1 Fan Guo 5 Yuyao Li 5 Jinjing Shen 5 Zhijing Xie 5 Yiyun Luo 5 Xin Zhang 5 Yingce Xia 2 Zequn Liu 2 1 Southeast University, Nanjing, China 2 Beijing Zhongguancun Academy, Beijing, China 3 Nanjing Normal University, Nanjing, China 5 ZhuiWen Technology Co., Ltd., Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.05724v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-05724","updatedAt":"2026-09-15T05:29:28Z"},{"type":"preprint","title":"LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video","authors":[{"name":"Shiqiang Lang"},{"name":"Jing Liu"},{"name":"Haoyang He"},{"name":"Peiwen Sun"},{"name":"Yuanteng Chen"},{"name":"Tao Liu"},{"name":"Lan Yang"},{"name":"Longteng Guo"},{"name":"Honggang Zhang"}],"institutions":["zgca"],"rawAffiliations":["Shiqiang Lang Affiliation: Beijing University of Posts and Telecommunications Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-04","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more than recognizing the current view, as models must remember and retrieve previously observed spatial layouts, routes, viewpoint changes, and object states. To evaluate this capability, we introduce LongSpace-Bench, a room-tour video benchmark for long-horizon spatial memory, covering scene perception, spatial relations, and spatial memory. In this work, we further propose LongSpace, a memory framework for long-video spatial reasoning. LongSpace models long videos as sequential chunks, incorporates 3D structural cues into early decoder layers, and constructs layer-aware memory for question-guided retrieval. Experiments on multiple spatial reasoning benchmarks show that LongSpace improves long-video spatial understanding, further demonstrating explicit spatial memory as a key capability for long-horizon video MLLMs.","identifiers":{"arxiv":"2606.05677","doi":"10.48550/arXiv.2606.05677","publishedDoi":"10.48550/arxiv.2606.05677"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.05677"},{"label":"HTML","url":"https://arxiv.org/html/2606.05677v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.05677"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.05677v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Shiqiang Lang Affiliation: Beijing University of Posts and Telecommunications Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.05677v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-05677","updatedAt":"2026-10-01T06:38:45Z"},{"type":"preprint","title":"OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining","authors":[{"name":"Zhongzheng Li"},{"name":"Tiancan Feng"},{"name":"Wenhao Li"},{"name":"Qingsong Ran"},{"name":"Feng S"},{"name":"Xiaoyuan Zhang"},{"name":"Yue Wang"},{"name":"X Zhao"}],"institutions":["zgca"],"rawAffiliations":["Zhongzheng Li Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Affiliation: Zhongguancun Academy Email: zhangxiaoyuan@bza.edu.cn"],"relationType":"affiliation","publishedAt":"2026-06-02","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, state dynamics, numerical stability, implementation constraints, and empirical generalization. Existing automated optimizer discovery methods typically search either over unconstrained code spaces or within narrowly parameterized optimizer families. The former is flexible but often produces invalid or uninterpretable programs, while the latter is stable but limits novelty. We introduce OPTScientist, a theory-guided multi-agent framework for optimizer discovery in a typed domain-specific language (DSL). OPTScientist formulates optimizer design as a constrained scientific search process, where candidate updates are expressed through direction, scaling, preconditioning, regularization, state, and grouping modules. Four role agents, Theorist, Designer, Engineer, and Reviewer, collaborate within a single orchestration loop to propose hypotheses, synthesize DSL candidates, compile and evaluate optimizers, and critique results. To overcome the limitations of a fixed search space, OPTScientist combines evolutionary search over optimizer programs with a second-stage mechanism that proposes small DSL extensions when repeated failures reveal representational bottlenecks. Using this framework, we discover RS-MR, a reduced-state matrix optimizer that improves transformer pretraining over strong baselines under our native evaluation protocol. Our results suggest a path toward automated optimizer science grounded in theory, typed programs, compiler validation, and closed-loop experimentation.","identifiers":{"arxiv":"2607.20486","doi":"10.48550/arXiv.2607.20486"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.20486"},{"label":"HTML","url":"https://arxiv.org/html/2607.20486v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.20486"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.20486v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongzheng Li Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Affiliation: Zhongguancun Academy Email: zhangxiaoyuan@bza.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.20486v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-20486","updatedAt":"2026-09-24T05:29:54Z"},{"type":"preprint","title":"PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search","authors":[{"name":"Kailin Lyu"},{"name":"Zhiqiang Yuan"},{"name":"Jianwei He"},{"name":"Qiwei Yan"},{"name":"Xuanbo Su"},{"name":"Nanxing Hu"},{"name":"Yang Liu (4829)"},{"name":"Ce Hao"},{"name":"Su‐Juan Qin"},{"name":"Lianyu Hu"},{"name":"Jinchao Zhang"},{"name":"Jie Zhou"}],"institutions":["zgca"],"rawAffiliations":["Kailin Lyu Affiliation: Pattern Recognition Center, WeChat AI, Tencent Inc. Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-06-02","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and reactive, lacking persistent memory to maintain long-horizon context or transfer experience across tasks, which often leads to execution drift and experience isolation. To address these limitations, we propose PhotoCraft, a training-free, hierarchical memory system for photo-search agents. Inspired by human cognition, PhotoCraft equips MLLMs with working, episodic, and semantic memory, which are dynamically invoked during reasoning to preserve logical consistency and knowledge transferability throughout multi-step reasoning and answer generation. Extensive experiments on DISBench demonstrate that PhotoCraft consistently improves context-aware retrieval across diverse MLLM backbones, achieving gains of up to 18.5\\% and effectively mitigating key bottlenecks in memoryless deep image search, offering a practical path toward reliable and generalizable multimodal search agents.","identifiers":{"arxiv":"2606.03099","doi":"10.48550/arXiv.2606.03099"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.03099"},{"label":"HTML","url":"https://arxiv.org/html/2606.03099v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.03099"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.03099v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Kailin Lyu Affiliation: Pattern Recognition Center, WeChat AI, Tencent Inc. Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.03099v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-03099","updatedAt":"2026-09-14T05:35:39Z"},{"type":"preprint","title":"Decoding in Order-Agnostic Language Models: Chain-Rule Deviation and Uniform Spreading","authors":[{"name":"Lin Yao"}],"institutions":["zgca"],"rawAffiliations":["Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-31","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Order-agnostic language models (OALMs), including discrete diffusion language models (dLLMs), are trained to predict masked tokens under arbitrary conditioning sets, allowing sequences to be generated or scored under arbitrary reveal orders at inference time. In LLaDA-2.1, we report three findings. First, the learned conditionals are not exact factorizations of a coherent joint distribution: changing only the reveal order shifts target log-likelihood by up to 0.49 nats/token, so likelihood alone mixes content difficulty with path-dependent artifacts. Second, although confidence-first (CF) decoding is order-agnostic, its reveal orders are close to left-to-right (L2R) on content tokens. Third, we propose a complementary diagnostic based on the shape of the confidence trace. A uniform-spreading theorem shows that, at fixed total likelihood, target recoverability is maximized when per-step confidence is spread uniformly; the resulting deviation motivates $\\mathrm{Var}(\\log q_t)$ as a diagnostic for comparing decoding paths. Across C4 and four downstream benchmarks, low variance separates structured paths from random ordering, and variance is consistently associated with downstream correctness. These results support reporting mean confidence and confidence variance jointly when comparing OALM decoding paths.","identifiers":{"arxiv":"2606.00997","doi":"10.48550/arXiv.2606.00997"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.00997"},{"label":"HTML","url":"https://arxiv.org/html/2606.00997v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.00997"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.00997v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.00997v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-00997","updatedAt":"2026-09-13T06:31:22Z"},{"type":"preprint","title":"ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment","authors":[{"name":"Qiuyu Tian"},{"name":"Haojie Yin"},{"name":"Yingce Xia"},{"name":"Youyong Kong"},{"name":"Zequn Liu"}],"institutions":["zgca"],"rawAffiliations":["Qiuyu Tian 1,2 Zequn Liu 2 Yingce Xia 2 Youyong Kong 1 Haojie Yin 3 1 Southeast University, Nanjing, China 2 Beijing Zhongguancun Academy, Beijing, China 3 Duke Kunshan University, Kunshan, China"],"relationType":"affiliation","publishedAt":"2026-05-30","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce ForeSci, a temporally controlled benchmark for evaluating whether LLM agents can make such forward-looking research judgements from historical evidence. ForeSci contains 500 tasks across four fast-moving AI domains and four decision families. Each task is paired with a cutoff-aligned offline knowledge base; post-cutoff papers are hidden during generation and used only for validation. To avoid random future-event prediction, tasks are derived from pre-cutoff taxonomy branches and evidence signals, and answer-generation backbones are selected to precede the task cutoffs. We evaluate native LLMs, Hybrid RAG, and three research-agent adaptations across four backbones. Results show that explicit evidence organization improves traceability and factual support, but gains depend strongly on the decision family. Diagnostics reveal a recurring evidence-decision decoupling: agents may cite relevant evidence while forecasting the wrong research object. ForeSci turns forward-looking AI research judgement into a controlled benchmark for evaluating research agents as decision-making systems.","identifiers":{"arxiv":"2606.00644","doi":"10.48550/arXiv.2606.00644"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2606.00644"},{"label":"HTML","url":"https://arxiv.org/html/2606.00644v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.00644"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.00644v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Qiuyu Tian 1,2 Zequn Liu 2 Yingce Xia 2 Youyong Kong 1 Haojie Yin 3 1 Southeast University, Nanjing, China 2 Beijing Zhongguancun Academy, Beijing, China 3 Duke Kunshan University, Kunshan, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.00644v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2606-00644","updatedAt":"2026-09-13T06:31:22Z"},{"type":"preprint","title":"SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes","authors":[{"name":"Tianhui Liu"},{"name":"Jie Feng"},{"name":"Zhiheng Zheng"},{"name":"Shengyuan Wang"},{"name":"Yiming Guo"},{"name":"Yanxin Xi"},{"name":"Hangyu Fan"},{"name":"Yong Li"},{"name":"Pan Hui"}],"institutions":["zgca"],"rawAffiliations":["Yiming Guo Yanxin Xi Hangyu Fan Yong Li Pan Hui Affiliation: The Hong Kong University of Science and Technology (Guangzhou) Affiliation: Zhongguancun Academy Affiliation: Tsinghua University Affiliation: Helsinki University"],"relationType":"affiliation","publishedAt":"2026-05-29","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \\textbf{SpatialAct}, a simulator-grounded benchmark for probing \\textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.","identifiers":{"arxiv":"2605.31148","doi":"10.48550/arXiv.2605.31148"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.31148"},{"label":"HTML","url":"https://arxiv.org/html/2605.31148v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.31148"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.31148v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yiming Guo Yanxin Xi Hangyu Fan Yong Li Pan Hui Affiliation: The Hong Kong University of Science and Technology (Guangzhou) Affiliation: Zhongguancun Academy Affiliation: Tsinghua University Affiliation: Helsinki University","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.31148v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-31148","updatedAt":"2026-09-13T06:31:22Z"},{"type":"preprint","title":"VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models","authors":[{"name":"Dehao Huang"},{"name":"Aoxiang Gu"},{"name":"Chengjie Zhang"},{"name":"Bolin Zou"},{"name":"Wenlong Dong"},{"name":"Zilang Cen"},{"name":"Yue Wang"},{"name":"Hong Zhang"}],"institutions":["zgca"],"rawAffiliations":["Dehao Huang Affiliation: Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China. Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-05-28","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level signal for monitoring manipulation in open-world environments and supporting downstream decision-making. Existing methods typically construct task-success confidence from action-token probabilities. However, such probabilities are not naturally available in flow-matching policies, limiting their applicability to mainstream flow-matching VLAs. To address this issue, we propose VLAConf, a two-stage representation-level confidence framework that operates on frozen pretrained VLA representations. A step-conditioned Coin-Flip Network learns an uncalibrated inverse success-support score from successful demonstrations, while a low-capacity calibrator fitted on outcome-labeled successful and failed rollouts maps the aggregated score to task-success probability. Experimental results on the LIBERO benchmark demonstrate that VLAConf improves online task-success confidence estimation over alternative approaches. We further demonstrate its utility in selective expert assistance, where confidence-triggered handoffs improve task success over no intervention. Its applicability is also evaluated in real-robot experiments. To access the source code and supplementary videos, visit https://sites.google.com/view/vlaconf.","identifiers":{"arxiv":"2605.29605","doi":"10.48550/arXiv.2605.29605"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.29605"},{"label":"HTML","url":"https://arxiv.org/html/2605.29605v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.29605"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.29605v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Dehao Huang Affiliation: Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China. Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.29605v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-29605","updatedAt":"2026-09-24T05:29:54Z"},{"type":"preprint","title":"Mind-Omni: A Unified Multi-Task Framework for Brain-Vision-Language Modeling via Discrete Diffusion","authors":[{"name":"Yizhuo Lu"},{"name":"Changde Du"},{"name":"Qingyu Shi"},{"name":"Hang Chen"},{"name":"Jie Peng"},{"name":"Liuyun Jiang"},{"name":"Shuangchen Zhao"},{"name":"Huiguang He"}],"institutions":["zgca"],"rawAffiliations":["Changde Du Affiliation: NeuBCI Lab, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing, China Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-05-28","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Modeling the interplay between external stimuli and internal neural representations is a pivotal research area for Brain-Computer Interfaces (BCIs). A major limitation of prior work is the prevailing paradigm of specialized, single-task models, which curtails versatility and neglects inter-task synergies. To address this, we propose Mind-Omni, the first versatile framework that unifies seven distinct encoding and decoding tasks through a discrete diffusion paradigm. At its core is a novel Brain Tokenizer that transforms heterogeneous, continuous brain signals into standardized, discrete tokens. This enables direct, token-level interactions for mutual understanding and generation between any two or more modalities within a shared semantic space. To unlock advanced reasoning capabilities, we further curate a specialized Brain Question Answering (BQA) instruction-tuning dataset. Our model not only establishes a new state-of-the-art among multi-task unified frameworks but also provides strong evidence for multi-task synergy. By demonstrating performance competitive with, and at times superior to, larger specialized models, our work offers a powerful new paradigm for neural modeling and paves the way for foundation models of neural activity. The code is publicly available at https://github.com/ReedOnePeck/Mind-Omni.","identifiers":{"arxiv":"2605.29591","doi":"10.48550/arXiv.2605.29591"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.29591"},{"label":"HTML","url":"https://arxiv.org/html/2605.29591v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.29591"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.29591v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Changde Du Affiliation: NeuBCI Lab, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing, China Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.29591v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-29591","updatedAt":"2026-09-13T05:26:44Z"},{"type":"preprint","title":"TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning","authors":[{"name":"Chusen Li"},{"name":"Zhou Liu"},{"name":"Shuigeng Zhou"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Chusen Li Affiliation: Fudan University Affiliation: Zhongguancun Academy Email: lichusen@whu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-27","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directly applying single-agent reinforcement learning to multi-turn multi-agent systems faces following dilemmas: i) Sparse rewards, role-level free-riding and excessive training overhead. ii) Agents only imitate to collaborate. iii) Fixed collaboration protocol falls into oscillating local optimum. We introduce TRACER, a turn-level reinforcement framework for cooperative multi-LLM reasoning. TRACER separates collaborative decision making into a controller-regret layer, where controllers learn whether the agents should speak or skip the current round through regret matching, and a generation-credit layer, which optimizes proposer and reviewer utterances with role-specific GSPO rewards. This design i) assigns credit at the level of both action modes and generated utterances, thus avoiding free-riding and sparse rewards. We only expand the choices made by the controllers, thus greatly reducing computational cost of training. Moreover, ii) agents acquire collaborative capability as they learn when to utter and what to speak. Finally, iii) by designing binary actions ingeniously, we extend classical game theory established for finite action spaces to deep learning, thus achieving mathematically rigorous convergence. We train all local RL-style methods on the GSM8K training split and evaluate on held-out GSM8K, MATH500, and GPQA-Diamond to measure in-domain accuracy, cross-benchmark generalization, inference cost, and correction-preservation behavior. The resulting framework provides a compact and reproducible testbed for studying learned collaboration policies beyond fixed debate, voting, or aggregation protocols. Code is available at https://github.com/Shark-Forest/TRACER.","identifiers":{"arxiv":"2605.28699","doi":"10.48550/arXiv.2605.28699"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.28699"},{"label":"HTML","url":"https://arxiv.org/html/2605.28699v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.28699"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.28699v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chusen Li Affiliation: Fudan University Affiliation: Zhongguancun Academy Email: lichusen@whu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.28699v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-28699","updatedAt":"2026-09-13T05:26:44Z"},{"type":"preprint","title":"DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding","authors":[{"name":"Jianfei Zhao"},{"name":"Feng Zhang"},{"name":"Xin Sun"},{"name":"Chong Feng"},{"name":"Bing Wang"},{"name":"Zhixing Tan"}],"institutions":["zgca"],"rawAffiliations":["Jianfei Zhao Affiliation: School of Computer Science and Technology, Beijing Institute of Technology Affiliation: Zhongguancun Academy Email: zhqingan@bit.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-26","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to text tokens. Visual tokens, the core carriers of visual information, are optimized only implicitly as part of the context, leading to coarse-grained visual understanding. Prior works attempt to supervise visual inputs but inevitably rely on auxiliary components such as additional decoders or forward passes, because visual tokens lack readily interpretable labels. This limits their practical applicability. In this work, we propose \\textbf{D}irect \\textbf{V}ision \\textbf{S}upervised \\textbf{F}ine-\\textbf{T}uning (DV-SFT), which constructs explicit, token-level supervision for visual tokens and trains them through the same next-token prediction objective used for text. Specifically, we exploit the direct vision--text correspondence in OCR-related scenarios and automatically label each visual token with the word in its corresponding image patch. DV-SFT treats the MLLM as a black box, requiring no architectural modifications or additional forward passes. Extensive experiments demonstrate the superiority of direct vision supervision. DV-SFT consistently outperforms standard SFT across three in-domain and four out-of-domain benchmarks. Further analyses show that vision supervision effectively enhances fine-grained visual understanding and achieves higher multimodal alignment efficiency.","identifiers":{"arxiv":"2605.26656","doi":"10.48550/arXiv.2605.26656"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.26656"},{"label":"HTML","url":"https://arxiv.org/html/2605.26656v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.26656"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.26656v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jianfei Zhao Affiliation: School of Computer Science and Technology, Beijing Institute of Technology Affiliation: Zhongguancun Academy Email: zhqingan@bit.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.26656v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-26656","updatedAt":"2026-09-12T05:10:53Z"},{"type":"preprint","title":"SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation","authors":[{"name":"Yuanhe Tian"},{"name":"Yan Song"}],"institutions":["zgca"],"rawAffiliations":["Yuanhe Tian Affiliation: Zhongguancun Academy Email: yhtian94@gmail.com"],"relationType":"affiliation","publishedAt":"2026-05-23","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices. Existing methods typically learn a direct image-to-text mapping, providing limited mechanisms for modeling how CT evidence evolves across slices or how reports respond to controlled changes in latent lesion-related factors. We propose SliceWorld, a CT-specific world-state framework that treats an axial CT scan as an ordered sequence along the z-axis. SliceWorld encodes prefix CT evidence into factor-aware latent states containing anatomy, lesion, and uncertainty components, and projects these states into world tokens used for multi-step future-slice feature prediction, lesion-factor intervention, and LLM-based report generation. The model is first pretrained on CT slice sequences with predictive, factor-aware, and counterfactual objectives, and is then fine-tuned on paired CT-report data. Experiments on M3D-Cap and CT-RATE show that SliceWorld improves natural language generation metrics and clinically oriented automatic evaluation. Further analyses demonstrate multi-horizon future-slice prediction, measurable factor alignment, reduced-slice robustness, and selective lesion-sensitive report modulation.","identifiers":{"arxiv":"2605.24371","doi":"10.48550/arXiv.2605.24371"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.24371"},{"label":"HTML","url":"https://arxiv.org/html/2605.24371v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.24371"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.24371v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuanhe Tian Affiliation: Zhongguancun Academy Email: yhtian94@gmail.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.24371v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-24371","updatedAt":"2026-08-13T12:34:14Z"},{"type":"preprint","title":"Observing Joinings: A Distance-Array Characterization of Furstenberg Disjointness","authors":[{"name":"Ao Xu"}],"institutions":["zgca"],"rawAffiliations":["Ao Xu Email: xuao24@mails.jlu.edu.cn Affiliation: School of Artificial Intelligence, Jilin University, No. 2699 Qianjin Street, Changchun, 130012, China Affiliation: Zhongguancun Academy, Daniufang 2nd Ring Road, Haidian District, Beijing, 100094, China"],"relationType":"affiliation","publishedAt":"2026-05-22","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Joinings are fundamental global objects in ergodic theory, yet in compact metric models one naturally observes only finite orbit-distance patterns. We bridge this gap by introducing multi-particle distance arrays, which sample finite orbit segments and record their joint metric evolution. In the anchored fixed-model setting, this framework yields a purely finite-observable characterization of Furstenberg disjointness: two systems are disjoint if and only if all their anchored multi-orbit distance-array projections are independent. The structural engine behind this criterion is a marked and colored version of the Gromov--Vershik reconstruction principle for exchangeable arrays; unanchored arrays reconstruct the intrinsic twin-free quotient, while anchors recover the actual joining in a fixed model. To quantify this independence, we introduce Wasserstein dependence coefficients, establishing an all-order zero criterion for disjointness, and show that weak neighborhoods of the product joining always admit finite distance-array certificates. Examples from compact rotations, Bernoulli and reversible Markov shifts, common factors, Kronecker factors, and weak mixing demonstrate the strict necessity of the multi-particle level and the broad scope of this approach.","identifiers":{"arxiv":"2605.23349","doi":"10.48550/arXiv.2605.23349"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.23349"},{"label":"HTML","url":"https://arxiv.org/html/2605.23349v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.23349"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.23349v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Ao Xu Email: xuao24@mails.jlu.edu.cn Affiliation: School of Artificial Intelligence, Jilin University, No. 2699 Qianjin Street, Changchun, 130012, China Affiliation: Zhongguancun Academy, Daniufang 2nd Ring Road, Haidian District, Beijing, 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.23349v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-23349","updatedAt":"2026-09-12T05:10:53Z"},{"type":"preprint","title":"SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models","authors":[{"name":"Zizhao Tong"},{"name":"Yeying Jin"},{"name":"Hongfeng Lai"},{"name":"Zeqing Wang"},{"name":"Zhaohu Xing"},{"name":"Kexu Cheng"},{"name":"Haoran Xu"},{"name":"Zhao Pu"},{"name":"Shangwen Zhu"},{"name":"Ruili Feng"},{"name":"Jian Zhao"},{"name":"Yan Zhang"},{"name":"Hao Tang"},{"name":"Ling Shao"}],"institutions":["zgci"],"rawAffiliations":["Jian Zhao Affiliation: Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-05-22","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing methods inject actions globally and train on single titles, failing under dense FPS inputs. We observe that FPS actions are spatially selective: discrete events such as firing or reloading affect only a localized region around the weapon (the scope), while continuous camera and movement signals govern stable surroundings. We propose SCOPE, which inserts a conditioning module into each transformer block of a pretrained video diffusion model. It reshapes features into per-pixel temporal sequences so that each position computes its action response from local visual content. This separates in-scope effects from out-of-scope generation without segmentation labels. We also introduce CrossFPS, the first multi-game FPS dataset with frame-aligned action telemetry. It comprises 69K clips from 7 titles with 10-DoF controller signals, curated to remove gameplay bias. The model learns general visual-to-action mappings rather than game-specific patterns, enabling zero-shot transfer to unseen scenes. Experiments confirm strong action responsiveness, precise scope separation, and effective cross-game generalization.","identifiers":{"arxiv":"2605.23345","doi":"10.48550/arXiv.2605.23345"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.23345"},{"label":"HTML","url":"https://arxiv.org/html/2605.23345v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.23345"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.23345v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Jian Zhao Affiliation: Zhongguancun Institute of Artificial Intelligence","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.23345v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-23345","updatedAt":"2026-09-12T05:10:53Z"},{"type":"preprint","title":"AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists","authors":[{"name":"Junshu Pan"},{"name":"Panzhong Lu"},{"name":"Yixuan Weng"},{"name":"Qiyao Sun"},{"name":"Fang Guo"},{"name":"Zijie Yang"},{"name":"Qiji Zhou"},{"name":"Yue Zhang"}],"institutions":["zgca"],"rawAffiliations":["Qiyao Sun 1 1 footnotemark: 1 Affiliation: Westlake University Affiliation: Zhongguancun Academy https://airaxiv.com"],"relationType":"affiliation","publishedAt":"2026-05-20","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Recent advances in artificial intelligence (AI) have accelerated the growth of both human-authored and AI-generated research outputs, placing increasing strain on traditional academic publishing systems and challenging the scalability of conference- and journal-centered paradigms amid rising submission volumes, reviewer workload, and venue size. To address these challenges, we explore an AI-era publishing paradigm in which both human and AI scientists participate as authors and readers, and papers evolve through continuous, feedback-driven iteration. We propose AiraXiv, an AI-driven open-access platform built on open preprints, AI-augmented analysis and review, and reader feedback. AiraXiv supports human scientists through an interactive UI and AI scientists through Model Context Protocol (MCP)-based interactions. We validate AiraXiv through real-world deployments, including serving as the submission platform for ICAIS 2025, demonstrating its potential as a fast, inclusive, and scalable research infrastructure for the AI era. AiraXiv is publicly available at https://airaxiv.com.","identifiers":{"arxiv":"2605.21481","doi":"10.48550/arXiv.2605.21481"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.21481"},{"label":"HTML","url":"https://arxiv.org/html/2605.21481v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.21481"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.21481v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Qiyao Sun 1 1 footnotemark: 1 Affiliation: Westlake University Affiliation: Zhongguancun Academy https://airaxiv.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.21481v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-21481","updatedAt":"2026-09-11T05:18:36Z"},{"type":"conference","title":"SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models","authors":[{"name":"Yiyang Gu"},{"name":"Junwei Yang"},{"name":"Junyu Luo"},{"name":"Ye Yuan"},{"name":"Bin Feng"},{"name":"Yingce Xia"},{"name":"Shufang Xie"},{"name":"Kaili Liu"},{"name":"Bohan Wu"},{"name":"Qi Shi"},{"name":"Haoran Li"},{"name":"Beier Xiao"},{"name":"Zhiping Xiao"},{"name":"Xiao Luo"},{"name":"Weizhi Zhang"},{"name":"Philip S. Yu"},{"name":"Zequn Liu"},{"name":"Ming Zhang"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）","Peking University Zhongguancun Academy IDEA"],"relationType":"official-output","publishedAt":"2026-05-19","year":2026,"venue":"ACL 2026","status":"published","topics":["cs.CL"],"abstract":"Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Most benchmarks are manually curated or domain-generic, limiting scalability and alignment with real scientific use cases. In this paper, we propose a new framework named SciCustom to address the problem. It enables the custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs. SciCustom first organizes scientific knowledge into ontology-grounded knowledge units with controlled granularity and trains a tagger to map large-scale data instances into this knowledge space. Given a custom requirement, relevant knowledge units are identified via voting-based multi-model consensus. These units enable relevance-aware benchmark retrieval via binary search, followed by proxy subset selection and data-grounded benchmark generation for efficient evaluation. Experiments in chemistry and healthcare demonstrate that SciCustom reveals fine-grained differences in LLM scientific capabilities that standard benchmarks overlook, while requiring neither expert annotation nor synthetic question generation. This work provides a scalable and application-aware foundation for benchmarking scientific capabilities in LLMs. The source code is available at https://github.com/yjwtheonly/SciCustom.","identifiers":{"arxiv":"2605.19357","doi":"10.48550/arXiv.2605.19357"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.19357"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.19357"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_y4wg2eznr5f9qipfzk688str4knm80bl"},{"label":"Code","url":"https://github.com/yjwtheonly/SciCustom"},{"label":"HTML","url":"https://arxiv.org/html/2605.19357v1"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.19357v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_y4wg2eznr5f9qipfzk688str4knm80bl"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_y4wg2eznr5f9qipfzk688str4knm80bl"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Peking University Zhongguancun Academy IDEA","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.19357v1"}],"sources":["arXiv","北京中关村学院官网","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-19357","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Semantic-Enriched Latent Visual Reasoning","authors":[{"name":"Tianrun Xu"},{"name":"Yue Sun"},{"name":"Qixun Wang"},{"name":"Jingyi Lu"},{"name":"Yuan Wang"},{"name":"Tianren Zhang"},{"name":"Longteng Guo"},{"name":"Fengyun Rao"},{"name":"Jing Lyu"},{"name":"Feng Chen"},{"name":"Jing Liu"}],"institutions":["zgca"],"rawAffiliations":["Tianrun Xu Affiliation: Department of Automation, Tsinghua University, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: WeChat Vision, Tencent Inc, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-05-19","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches largely rely on visual supervision and produce latent representations that lack sufficient semantic richness, limiting their ability to support diverse region-level reasoning tasks. In this work, we introduce Semantic-Enriched Latent Visual Reasoning (SLVR), a two-stage learning framework that enriches latent representations with attribute-level visual semantics and aligns them with diverse reasoning objectives. In the first stage, SLVR learns semantically enriched region-centric latents under fine-grained attribute supervision. In the second stage, we design Multi-query Group Relative Policy Optimization (M-GRPO) to align latent representations across multiple queries grounded in the same region. To support this framework, we construct SLV-Set, comprising approximately 400K region-level attribute annotations and 800K multi-query question answering samples, and introduce SV-QA, a benchmark that evaluates latent reasoning under semantic variation. Experiments demonstrate that SLVR improves the robustness and semantic consistency of latent visual reasoning compared to existing baselines.","identifiers":{"arxiv":"2605.19342","doi":"10.48550/arXiv.2605.19342"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.19342"},{"label":"HTML","url":"https://arxiv.org/html/2605.19342v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.19342"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.19342v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Tianrun Xu Affiliation: Department of Automation, Tsinghua University, Beijing, China Affiliation: Zhongguancun Academy, Beijing, China Affiliation: WeChat Vision, Tencent Inc, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.19342v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-19342","updatedAt":"2026-09-11T05:18:36Z"},{"type":"preprint","title":"VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction","authors":[{"name":"K. S. Zhu"},{"name":"Yiwen Tang"},{"name":"Yifan Yang"},{"name":"Renrui Zhang"},{"name":"Bohan Zeng"},{"name":"Ziyu Guo"},{"name":"Ruichuan An"},{"name":"Zhou Liu"},{"name":"Qizhi Chen"},{"name":"Delin Qu"},{"name":"Jaehong Yoon"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Wentao Zhang † † thanks: Corresponding Author. Affiliation: Peking University Affiliation: Zhongguancun Academy Affiliation: Beijing Key Lab of Data Intel. & Security (PKU)"],"relationType":"affiliation","publishedAt":"2026-05-14","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environments in a single forward pass. However, despite their strong performance in static scene perception, these models remain limited in responding to dynamic human instructions, which restricts their use in interactive applications. Existing editing methods typically rely on a 2D-lifting strategy, where individual views are edited independently and then lifted back into 3D space. This indirect pipeline often leads to blurry textures and inconsistent geometry, as 2D editors lack the spatial awareness required to preserve structure across viewpoints. To address these limitations, we propose VGGT-Edit, a feed-forward framework for text-conditioned native 3D scene editing. VGGT-Edit introduces depth-synchronized text injection to align semantic guidance with the backbone's spatial poses, ensuring stable instruction grounding. This semantic signal is then processed by a residual transformation head, which directly predicts 3D geometric displacements to deform the scene while preserving background stability. To ensure high-fidelity results, we supervise the framework with a multi-term objective function that enforces geometric accuracy and cross-view consistency. We also construct the DeltaScene Dataset, a large-scale dataset generated through an automated pipeline with 3D agreement filtering to ensure ground-truth quality. Experiments show that VGGT-Edit substantially outperforms 2D-lifting baselines, producing sharper object details, stronger multi-view consistency, and near-instant inference speed. The project page is https://chriszkxxx.github.io/VGGT-Edit/.","identifiers":{"arxiv":"2605.15186","doi":"10.48550/arXiv.2605.15186"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.15186"},{"label":"HTML","url":"https://arxiv.org/html/2605.15186v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.15186"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.15186v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Wentao Zhang † † thanks: Corresponding Author. Affiliation: Peking University Affiliation: Zhongguancun Academy Affiliation: Beijing Key Lab of Data Intel. & Security (PKU)","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.15186v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-15186","updatedAt":"2026-09-10T05:23:07Z"},{"type":"preprint","title":"Distance-Matrix Wasserstein Statistics for Scalable Gromov--Wasserstein Learning","authors":[{"name":"Ao Xu"},{"name":"Tieru Wu"}],"institutions":["zgca"],"rawAffiliations":["Ao Xu xuao24@mails.jlu.edu.cn Affiliation: School of Artificial Intelligence Affiliation: Jilin University Affiliation: No. 2699, Qianjin Street, Chaoyang District Affiliation: Changchun 130012, China Affiliation: and Affiliation: Zhongguancun Academy Affiliation: Daniufang 2nd Ring Road, Haidian District Affiliation: Beijing 100094, China"],"relationType":"affiliation","publishedAt":"2026-05-14","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic optimal transport problem and is difficult to estimate at scale. We propose \\emph{Distance-Matrix Wasserstein} (DMW), a hierarchy of Wasserstein statistics comparing laws of random finite distance matrices. Rather than optimizing a global point-level alignment, DMW samples $n$ points from each space, records their pairwise distances, and transports the resulting matrix laws. We prove that DMW is a relaxation and lower bound of GW, and establish a reverse approximation inequality: the GW--DMW gap is controlled by the Wasserstein error of approximating each original measure with $n$ samples. Hence population DMW converges to GW as sampled subspaces become dense. We further give finite-sample bounds, including intrinsic-dimensional rates that depend on the data manifold rather than the ambient matrix dimension $\\binom n2$. For scalable computation, we introduce sliced and multi-scale DMW; for $p=1$, the sliced multi-scale dissimilarity yields positive-definite exponential kernels. Experiments on synthetic metric spaces, scalability benchmarks, graph classification, and two-sample testing validate the theory and demonstrate an interpretable GW-style proxy for structural comparison.","identifiers":{"arxiv":"2605.14981","doi":"10.48550/arXiv.2605.14981"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.14981"},{"label":"HTML","url":"https://arxiv.org/html/2605.14981v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.14981"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.14981v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Ao Xu xuao24@mails.jlu.edu.cn Affiliation: School of Artificial Intelligence Affiliation: Jilin University Affiliation: No. 2699, Qianjin Street, Chaoyang District Affiliation: Changchun 130012, China Affiliation: and Affiliation: Zhongguancun Academy Affiliation: Daniufang 2nd Ring Road, Haidian District Affiliation: Beijing 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.14981v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-14981","updatedAt":"2026-08-13T11:57:17Z"},{"type":"conference","title":"IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation","authors":[{"name":"Shijie Lian"},{"name":"Bin Yu"},{"name":"Xiaopeng Lin"},{"name":"Zhaolong Shen"},{"name":"Laurence Tianruo Yang"},{"name":"Yurun Jin"},{"name":"Haishan Liu"},{"name":"Changti Wu"},{"name":"Hang Yuan"},{"name":"Cong Huang"},{"name":"Kai Chen"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）"],"relationType":"official-output","publishedAt":"2026-05-14","year":2026,"venue":"EMNLP 2026","status":"published","topics":["cs.RO","cs.AI","cs.CL","cs.CV"],"abstract":"Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines","identifiers":{"arxiv":"2605.14712","doi":"10.48550/arXiv.2605.14712"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.14712"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.14712"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_jsml7c2h0u3f1xzkexdvtf9i5fpjwpaq"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.14712v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_jsml7c2h0u3f1xzkexdvtf9i5fpjwpaq"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_jsml7c2h0u3f1xzkexdvtf9i5fpjwpaq"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2605-14712","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding","authors":[{"name":"Xun Fang"},{"name":"Yunchen Li"},{"name":"Hang Yuan"},{"name":"Zhou Yu"}],"institutions":["zgca"],"rawAffiliations":["Hang Yuan Affiliation: East China Normal University Affiliation: Beijing Zhongguancun Academy Email: 52274404018@stu.ecnu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-14","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Discrete diffusion language models improve generation efficiency through parallel token prediction, but standard $X_0$ prediction methods introduce factorization errors by approximating the clean token posterior with independent token-wise distributions. This paper proposes Factorization-Error-Free Discrete Diffusion Language Modeling (FeF-DLLM), which replaces independent clean-token prediction with an exact prefix-conditioned factorization of the clean posterior to better preserve token dependencies. To reduce the sequential cost introduced by prefix conditioning, FeF-DLLM further incorporates speculative decoding within diffusion denoising, accelerating inference while maintaining the parallel prediction and re-masking properties of DLLMs. Theoretically, we prove that FeF-DLLM generates from the true joint distribution and derive its expected acceleration ratio. Experiments on GSM8K, MATH, HumanEval, and MBPP demonstrate that our method improves accuracy by an average of 5.04 percentage points while achieving an average inference speedup of $3.86\\times$.","identifiers":{"arxiv":"2605.14305","doi":"10.48550/arXiv.2605.14305"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.14305"},{"label":"HTML","url":"https://arxiv.org/html/2605.14305v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.14305"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.14305v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hang Yuan Affiliation: East China Normal University Affiliation: Beijing Zhongguancun Academy Email: 52274404018@stu.ecnu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.14305v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-14305","updatedAt":"2026-09-10T05:23:07Z"},{"type":"preprint","title":"FrameSkip: Learning from Fewer but More Informative Frames in VLA Training","authors":[{"name":"Bin Yu"},{"name":"Shijie Lian"},{"name":"Xiaopeng Lin"},{"name":"Zhaolong Shen"},{"name":"Yuliang Wei"},{"name":"Changti Wu"},{"name":"Hang Yuan"},{"name":"Haishan Liu"},{"name":"Bailing Wang"},{"name":"Cong Huang"},{"name":"Kai Chen"}],"institutions":["zgca","zgci"],"rawAffiliations":["Bin Yu Affiliation: Harbin Institute of Technology Affiliation: Zhongguancun Academy","Xiaopeng Lin Affiliation: Zhongguancun Institute of Artificial Intelligence Affiliation: The Hong Kong University of Science and Technology (Guangzhou)"],"relationType":"affiliation","publishedAt":"2026-05-13","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Vision-Language-Action (VLA) policies are commonly trained from dense robot demonstration trajectories, often collected through teleoperation, by sampling every recorded frame as if it provided equally useful supervision. We argue that this convention creates a temporal supervision imbalance: long low-change segments dominate the training stream, while manipulation-critical transitions such as alignment, contact, grasping, and release appear only sparsely. We introduce FrameSkip, a data-layer frame selection framework that scores trajectory frames using action variation, visual-action coherence, task-progress priors, and gripper-transition preservation, then remaps training samples toward high-importance frames under a target retention ratio. Because FrameSkip operates only in the dataloader, it leaves the VLA architecture, action head, training objective, and inference procedure unchanged. Across RoboCasa-GR1, SimplerEnv, and LIBERO, FrameSkip improves the success-retention trade-off over full-frame training and simpler frame selection variants, achieving a macro-average success rate of 76.15% across the three benchmarks compared with 66.50% for full-frame training while using a compressed trajectory view that retains 20% of unique frames in the main setting.","identifiers":{"arxiv":"2605.13757","doi":"10.48550/arXiv.2605.13757"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.13757"},{"label":"HTML","url":"https://arxiv.org/html/2605.13757v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.13757"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.13757v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Bin Yu Affiliation: Harbin Institute of Technology Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.13757v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Xiaopeng Lin Affiliation: Zhongguancun Institute of Artificial Intelligence Affiliation: The Hong Kong University of Science and Technology (Guangzhou)","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.13757v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-13757","updatedAt":"2026-09-10T05:23:07Z"},{"type":"conference","title":"PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding","authors":[{"name":"Yunhe Han"},{"name":"Yunqi Gao"},{"name":"Bing Hu"},{"name":"Mahdi Boloursaz Mashhadi"},{"name":"Yitong Duan"},{"name":"Pei Xiao"},{"name":"Yanfeng Zhang"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-05-13","year":2026,"venue":"ICML 2026","status":"published","topics":["cs.DC"],"abstract":"Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness, and privacy enhancement. However, existing collaborative inference frameworks with speculative decoding are constrained by (i) sequential token generation and communication with low resource utilization, and (ii) inflexible cloud non-autoregressive verification (NAV) triggering that induces premature verification or costly rollbacks. In this paper, we propose PipeSD, an efficient cloud-edge collaborative pipeline inference framework with speculative decoding. PipeSD overlaps token generation and communication by a token-batch pipeline scheduling mechanism optimized by dynamic programming, and improves verification flexibility through a dual-threshold NAV triggering mechanism with a lightweight Bayesian optimization autotuner. We implement PipeSD using llama-cpp-python, PyTorch, and FastAPI, and evaluate it on a real-world cloud-edge testbed with two draft-target model pairs across four scenarios. Results show that PipeSD consistently outperforms state-of-the-art baselines, achieving 1.16x-2.16x speedup and reducing energy consumption by 14.3%-25.3%.","identifiers":{"arxiv":"2605.13319","doi":"10.48550/arXiv.2605.13319"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.13319"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.13319"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_6typbbqd9r7y99a7hktif3zbkaa02k5t"},{"label":"Code","url":"https://github.com/Ghanyunhe/PipeSD"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.13319v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_6typbbqd9r7y99a7hktif3zbkaa02k5t"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2605-13319","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"LeanSearch v2: Global Premise Retrieval for Lean 4 Theorem Proving","authors":[{"name":"Guoxiong Gao"},{"name":"Zeming Sun"},{"name":"Jiedong Jiang"},{"name":"Yutong Wang"},{"name":"Jingda Xu"},{"name":"Peihao Wu"},{"name":"Bryan Dai"},{"name":"Bin Dong"}],"institutions":["zgca"],"rawAffiliations":["Guoxiong Gao Zeming Sun Jiedong Jiang Yutong Wang Jingda Xu Peihao Wu Bryan Dai Bin Dong School of Mathematical Sciences, Peking University IQuest Research Research Institute for Mathematical Sciences, Kyoto University Westlake Institute for Advanced Study, Westlake University Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University Center for Machine Learning Research, Peking University Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University Zhongguancun Academy Corresponding author: dongbin@math.pku.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-13","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Proving theorems in Lean 4 often requires identifying a scattered set of library lemmas whose joint use enables a concise proof -- a task we call global premise retrieval. Existing tools address adjacent problems: semantic search engines find individual declarations matching a query, while premise-selection systems predict useful lemmas one tactic step at a time. Neither recovers the full premise set an entire theorem requires. We present LeanSearch v2, a two-mode retrieval system for this task. Its standard mode applies a hierarchy-informalized Mathlib corpus with an embedding-reranker pipeline, achieving state-of-the-art single-query retrieval without domain-specific fine-tuning (nDCG@10 of 0.62 vs. 0.53 for the next-best system). Its reasoning mode builds on standard mode as its retrieval substrate, targeting global premise retrieval through iterative sketch-retrieve-reflect cycles. On a 69-query benchmark of research-level Mathlib theorems, reasoning mode recovers 46.1% of ground-truth premise groups within 10 retrieved candidates, outperforming strong reasoning retrieval systems (38.0%) and premise-selection baselines (9.3%) on the same benchmark. In a controlled downstream evaluation with a fixed prover loop, replacing alternative retrievers with LeanSearch v2 yields the highest proof success (20% vs. 16% for the next-best system and 4% without retrieval), confirming that retrieval quality propagates to proof generation. We have open-sourced all code, data, and benchmarks. Code and data: https://github.com/frenzymath/LeanSearch-v2 . The standard mode is publicly available with API access at https://leansearch.net/ .","identifiers":{"arxiv":"2605.13137","doi":"10.48550/arXiv.2605.13137"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.13137"},{"label":"HTML","url":"https://arxiv.org/html/2605.13137v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.13137"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.13137v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Guoxiong Gao Zeming Sun Jiedong Jiang Yutong Wang Jingda Xu Peihao Wu Bryan Dai Bin Dong School of Mathematical Sciences, Peking University IQuest Research Research Institute for Mathematical Sciences, Kyoto University Westlake Institute for Advanced Study, Westlake University Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University Center for Machine Learning Research, Peking University Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University Zhongguancun Academy Corresponding author: dongbin@math.pku.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.13137v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-13137","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models","authors":[{"name":"Zixing Lei"},{"name":"Changxing Liu"},{"name":"Yu Xiong"},{"name":"Moulin Xiong"},{"name":"Yudan Ding"},{"name":"Zhipeng Zhang"},{"name":"Weixin Li"},{"name":"Siheng Chen"}],"institutions":["zgca"],"rawAffiliations":["Zixing Lei Thanks: Equal contribution Affiliation: Shanghai Jiao Tong University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-05-13","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of extended closed-loop planning and diverse physical operations. We therefore propose VLAs-as-Tools, a strategy that distributes this burden across a high-level vision language model (VLM) agent for temporal reasoning and a family of specialized VLA tools for diverse local physical operations. The VLM handles scene analysis, global planning, and recovery, while each VLA tool executes a bounded subtask. To tightly couple agent planning with VLA tool execution in long-horizon tasks, we introduce a VLA tool-family interface that exposes explicit tool selection and in-execution progress feedback, enabling efficient event-triggered agent replanning without continuous agent polling. To obtain diverse specialized VLA tools that faithfully follow agent invocations, we further propose Tool-Aligned Post-Training (TAPT), which constructs invocation-aligned training units for instruction following and adopts tool-family residual adapters for efficient tool specialization. Experiments show that VLAs-as-Tools improves the success rate of $π_{0.5}$ by 4.8 points on LIBERO-Long and 23.1 points on RoboTwin, and further enhances invocation fidelity by 15.0 points as measured by Non-biased Rate. Code will be released.","identifiers":{"arxiv":"2605.13119","doi":"10.48550/arXiv.2605.13119"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.13119"},{"label":"HTML","url":"https://arxiv.org/html/2605.13119v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.13119"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.13119v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zixing Lei Thanks: Equal contribution Affiliation: Shanghai Jiao Tong University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.13119v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-13119","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation","authors":[{"name":"Jiayu Chen"},{"name":"Junbei Tang"},{"name":"Wenbiao Zhao"},{"name":"Maoliang Li"},{"name":"Jiayi Luo"},{"name":"Zihao Zheng"},{"name":"Jiawei Yang"},{"name":"Guojie Luo"},{"name":"Xiang Chen"}],"institutions":["zgca"],"rawAffiliations":["Jiayi Luo Affiliation: Beihang University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-05-13","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulated errors. Existing KVCache strategies usually apply unified historical-frame retention, implicitly assuming homogeneous historical dependencies across attention heads. We revisit historical-frame attention and reveal three distinct head types: Anchor Heads require broad long-range context, Wave Heads exhibit periodic temporal dependencies, and Veil Heads focus on initial and adjacent frames. Based on this finding, we propose Pyramid Forcing, a head-aware pyramidal KVCache framework that identifies head types offline, assigns behavior-specific cache policies, and supports heterogeneous cache lengths via efficient ragged-cache attention. Experiments on Self Forcing and Causal Forcing show that Pyramid Forcing consistently improves long-horizon generation quality on VBench-Long, increasing the 60-second Self Forcing score from 77.87 to 81.21 while enhancing motion dynamics, visual fidelity, and semantic consistency. Project: https://if-lab-pku.github.io/Pyramid-Forcing/.","identifiers":{"arxiv":"2605.13111","doi":"10.48550/arXiv.2605.13111"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.13111"},{"label":"HTML","url":"https://arxiv.org/html/2605.13111v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.13111"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.13111v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiayi Luo Affiliation: Beihang University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.13111v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-13111","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective","authors":[{"name":"Feng Zhang"},{"name":"Xinhong Ma"},{"name":"Ziqiang Dong"},{"name":"Xi Leng"},{"name":"Jianfei Zhao"},{"name":"Xin Sun"},{"name":"Yang Yang"},{"name":"Guanjun Jiang"}],"institutions":["zgca"],"rawAffiliations":["Jianfei Zhao Affiliation: Beijing Institute of Technology Affiliation: Zhongguancun Academy {bit_zhangfeng, zhqingan, sunxin}@bit.edu.cn{xinhong.mxh, ziqiang.dzq, chris.yang, guanj.jianggj}@alibaba-inc.comxileng@link.cuhk.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-13","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admits an equivalent discriminative reformulation, in which policy optimization maximizes the expected score gap between verified positive and negative rollouts. This reformulation reveals two objective-level limitations: likelihood-misaligned surrogate scores, in which clipped ratio-based scores are optimized rather than the sequence likelihoods that govern generation, and score-insensitive credit assignment, in which rollout-level credit does not reflect the current score gaps between positive and negative rollouts. To address these limitations, we propose ConSPO, a Contrastive Sequence-level Policy Optimization method that uses length-normalized sequence log-probabilities as rollout scores and contrasts verified positive rollouts against negative distractors within the same group. ConSPO optimizes a group-wise InfoNCE-style objective to adaptively strengthen updates for poorly separated positives and high-scoring negatives, together with a curriculum-scheduled margin that preserves separation pressure as training progresses. Experiments across diverse settings show that ConSPO outperforms strong baselines on challenging reasoning benchmarks. Code will be released upon paper acceptance.","identifiers":{"arxiv":"2605.12969","doi":"10.48550/arXiv.2605.12969"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.12969"},{"label":"HTML","url":"https://arxiv.org/html/2605.12969v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.12969"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.12969v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jianfei Zhao Affiliation: Beijing Institute of Technology Affiliation: Zhongguancun Academy {bit_zhangfeng, zhqingan, sunxin}@bit.edu.cn{xinhong.mxh, ziqiang.dzq, chris.yang, guanj.jianggj}@alibaba-inc.comxileng@link.cuhk.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.12969v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-12969","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning","authors":[{"name":"Zijun Shen"},{"name":"Sihan Yang"},{"name":"Ruichuan An"},{"name":"Ziyu Guo"},{"name":"Hao Liang"},{"name":"Ming Lü"},{"name":"Renrui Zhang"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Hao Liang Affiliation: Peking University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-05-11","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token-level alignment via supervised fine-tuning, which fails to fully capture the potential synergy between comprehension and creation. In this work, we propose Sync-R1, an end-to-end reinforcement learning framework that jointly optimizes personalized understanding and generation within a single, explicit reasoning loop. Through this unified feedback process, Sync-R1 enables personalized comprehension to guide content creation, while the resulting generation quality reciprocally refines understanding within an integrated reward landscape. To efficiently orchestrate this dual-task synergy, we introduce Sync-GRPO, a reinforcement learning method utilizing an ensemble reward system. Furthermore, we propose Dynamic Group Scaling (DGS), which adaptively filters low-potential trajectories to reduce gradient variance and accelerate convergence. To better reflect real-world complexity, we introduce UnifyBench++, featuring denser textual descriptions and richer user contexts. Experimental results demonstrate that Sync-R1 achieves state-of-the-art performance, showcasing superior cross-task reasoning and robust personalization without requiring complex cold-start procedures. The code and the UnifyBench++ dataset will be released at: https://github.com/arctanxarc/UniCTokens.","identifiers":{"arxiv":"2605.10445","doi":"10.48550/arXiv.2605.10445"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.10445"},{"label":"HTML","url":"https://arxiv.org/html/2605.10445v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.10445"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.10445v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hao Liang Affiliation: Peking University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.10445v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-10445","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"Accelerating Locality-Driven Integration in Quantum Chemistry with Block-Structured Matrix Multiplication","authors":[{"name":"Xinran Wei"},{"name":"Yan Pan"},{"name":"Fusong Ju"},{"name":"Zehao Zhou"},{"name":"Yihong Zhang"},{"name":"Lin Huang"},{"name":"Jianwei Zhu"},{"name":"Jia Zhang"},{"name":"Huanhuan Xia"},{"name":"Bin Shao"},{"name":"Tao Qin"}],"institutions":["zgca","zgci"],"rawAffiliations":["Xinran Wei 1 2 5 , Yan Pan 1 , Fusong Ju 5 1 6 , Zehao Zhou 5 , Yihong Zhang 5 , Lin Huang 4 , Jianwei Zhu 5 , Jia Zhang 4 , Huanhuan Xia 1 , Bin Shao 1 2 , Tao Qin 1 Affiliation: 1 Zhongguancun Academy , Beijing, China Affiliation: 2 Zhongguancun Institute of Artificial Intelligence , Beijing, China Affiliation: 4 IQuest Research , Beijing, China"],"relationType":"affiliation","publishedAt":"2026-05-11","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Locality-driven integration is a pervasive computational pattern in quantum chemistry, arising whenever spatially localized basis functions interact through numerical quadrature or integral screening. The dominant matrix multiplications in these tasks exhibit dynamic, structured sparsity driven by spatial locality, posing significant challenges for both dense batched kernels and generic sparse formats on GPUs. We present KerneLDI, a GPU-oriented framework that addresses this regime by co-designing data layout, screening logic, and matrix-computation operators to realize block-structured matrix multiplication for locality-driven integration. KerneLDI reorganizes operand matrices into a unified block-filtered representation that retains only spatially relevant blocks, and executes the resulting contractions with customized dense block multipliers that adapt proven dense-matmul optimizations to retained block pairs. We develop and evaluate KerneLDI on exchange--correlation (EXC) integration in Kohn--Sham density functional theory, a representative and computationally critical instance of this pattern. Across diverse molecular systems, KerneLDI preserves numerical accuracy while delivering up to 10$\\times$ speedup for EXC evaluation over a dense GPU baseline, scales favorably with increasing system size and multi-GPU parallelism, accelerates end-to-end self-consistent field calculations, and yields nearly 6$\\times$ throughput improvement for ab initio molecular dynamics.","identifiers":{"arxiv":"2605.10363","doi":"10.48550/arXiv.2605.10363"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.10363"},{"label":"HTML","url":"https://arxiv.org/html/2605.10363v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.10363"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.10363v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xinran Wei 1 2 5 , Yan Pan 1 , Fusong Ju 5 1 6 , Zehao Zhou 5 , Yihong Zhang 5 , Lin Huang 4 , Jianwei Zhu 5 , Jia Zhang 4 , Huanhuan Xia 1 , Bin Shao 1 2 , Tao Qin 1 Affiliation: 1 Zhongguancun Academy , Beijing, China Affiliation: 2 Zhongguancun Institute of Artificial Intelligence , Beijing, China Affiliation: 4 IQuest Research , Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.10363v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Xinran Wei 1 2 5 , Yan Pan 1 , Fusong Ju 5 1 6 , Zehao Zhou 5 , Yihong Zhang 5 , Lin Huang 4 , Jianwei Zhu 5 , Jia Zhang 4 , Huanhuan Xia 1 , Bin Shao 1 2 , Tao Qin 1 Affiliation: 1 Zhongguancun Academy , Beijing, China Affiliation: 2 Zhongguancun Institute of Artificial Intelligence , Beijing, China Affiliation: 4 IQuest Research , Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.10363v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-10363","updatedAt":"2026-08-22T02:17:52Z"},{"type":"preprint","title":"FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries","authors":[{"name":"Qijie You"},{"name":"Hao Liang"},{"name":"Mingrui Chen"},{"name":"Bohan Zeng"},{"name":"Meiyi Qiang"},{"name":"Zhenhao Wong"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Qijie You 1 Hao Liang 2,4 1 1 footnotemark: 1 Mingrui Chen 3 Bohan Zeng 2 Meiyi Qiang 2 Zhenhao Wong 2 Wentao Zhang 2,4 1 University of Science and Technology Beijing 2 Peking University 3 Institute of Automation, Chinese Academy of Sciences 4 Zhongguancun Academy Thanks: Equal contribution. Thanks: Project leader. Thanks: Corresponding author."],"relationType":"affiliation","publishedAt":"2026-05-11","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"As video becomes increasingly central to information dissemination and multimodal large language models (MLLMs) continue to advance, evaluating video retrieval has become increasingly important. In realistic search scenarios, this requires matching short user queries to long-form content using both visual and auditory evidence. Yet existing retrieval benchmarks are still dominated by short clips, single modalities, and caption-based evaluation. We introduce FLARE, a full-modality long-video audiovisual retrieval benchmark with user-simulated queries. Built from 399 carefully screened Video-MME videos (10--60 min, 225.4 h) to ensure source quality and diversity, FLARE contains 87,697 clips annotated with vision, audio, and unified audiovisual captions, together with 274,933 user-style queries. Cross-modal queries are further filtered by a hard bimodal constraint, requiring retrieval to fail under either modality alone but succeed when both are combined. FLARE evaluates models under two regimes, caption-based and query-based retrieval, across vision, audio, and unified audiovisual settings. Experiments with 15 representative retrievers show that user-style queries substantially change model behavior, strong caption-based performance does not always transfer to query-based retrieval, and audio--language alignment remains a key bottleneck for unified audiovisual retrieval. Our code and data are released at https://flarebench.github.io/","identifiers":{"arxiv":"2605.10228","doi":"10.48550/arXiv.2605.10228"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.10228"},{"label":"HTML","url":"https://arxiv.org/html/2605.10228v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.10228"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.10228v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Qijie You 1 Hao Liang 2,4 1 1 footnotemark: 1 Mingrui Chen 3 Bohan Zeng 2 Meiyi Qiang 2 Zhenhao Wong 2 Wentao Zhang 2,4 1 University of Science and Technology Beijing 2 Peking University 3 Institute of Automation, Chinese Academy of Sciences 4 Zhongguancun Academy Thanks: Equal contribution. Thanks: Project leader. Thanks: Corresponding author.","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.10228v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-10228","updatedAt":"2026-08-21T02:28:00Z"},{"type":"preprint","title":"SkillMaster: Toward Autonomous Skill Mastery in LLM Agents","authors":[{"name":"Min Yang"},{"name":"Jinghua Piao"},{"name":"Xu Xia"},{"name":"Xiaochong Lan"},{"name":"Jiaju Chen"},{"name":"Yongshun Gong"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Min Yang Affiliation: Shandong University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-05-09","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed by external teachers, hand-designed rules, or auxiliary modules. As a result, skills remain external resources to be invoked, rather than capabilities that agents can develop, adapt, and internalize through experience. To endow LLM agents with autonomous skill mastery, we propose SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving. This capability is achieved through three key designs. First, we train agents through trajectory-informed skill review, teaching agents to propose, update, or retain skills based on evidence from completed episodes. Second, each candidate skill edit is designed to be evaluated by its counterfactual utility on related probe tasks, providing a direct learning signal for training skill-editing decisions. Third, we introduce DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training across task solving and skill management. Experiments on ALFWorld and WebShop show that SkillMaster improves the overall success rate over state-of-the-art baselines by 8.8% and 9.3%, respectively, achieving the best performance among all compared methods. Further analysis reveals a marked shift in agent capability: agents trained with SkillMaster can identify skill failures, refine procedural knowledge from trajectory evidence, and transfer improvements to future tasks with limited skill-bank edits. Overall, SkillMaster moves LLM agents beyond mere skill use toward self-improving agents capable of developing, adapting, and applying their own skill repertoires.","identifiers":{"arxiv":"2605.08693","doi":"10.48550/arXiv.2605.08693"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.08693"},{"label":"HTML","url":"https://arxiv.org/html/2605.08693v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.08693"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.08693v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Min Yang Affiliation: Shandong University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.08693v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-08693","updatedAt":"2026-08-21T02:28:00Z"},{"type":"preprint","title":"SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation","authors":[{"name":"Jie Sun"},{"name":"Mao Zheng"},{"name":"Mingyang Song"},{"name":"Qiyong Zhong"},{"name":"Yilin Cheng"},{"name":"Bichuan Feng"},{"name":"Pengfei Liu"},{"name":"Junfeng Fang"},{"name":"Xiang Wang"}],"institutions":["zgca"],"rawAffiliations":["Yilin Cheng Affiliation: Zhongguancun Academy[4pt] sunjie2019@mail.ustc.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-08","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and student predictions are comparable token by token, an assumption that fails whenever the two models tokenize the same text differently. Under heterogeneous tokenizers, exact shared-token matching silently discards a large fraction of the teacher signal at precisely the positions where vocabularies disagree. We propose \\textbf{\\underline{Sim}ple \\underline{C}ross-\\underline{T}okenizer OPD (SimCT)}, which restores this signal by enlarging the supervision space: alongside shared tokens, SimCT compares teacher and student over short multi-token continuations that both tokenizers can realize, leaving the OPD loss form itself unchanged. We show that these units are the finest jointly tokenizable supervision interface, and that coarser alternatives remove teacher-student distinctions that are useful for on-policy learning. Across three heterogeneous teacher-student pairs on mathematical reasoning and code-generation benchmarks, SimCT shows consistent gains over shared-vocabulary OPD and representative cross-tokenizer baselines, with ablations confirming that the improvements come from recovering supervision discarded by exact shared-token matching. Code is available at \\href{https://github.com/sunjie279/SimCT-}{https://github.com/sunjie279/SimCT-}.","identifiers":{"arxiv":"2605.07711","doi":"10.48550/arXiv.2605.07711"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.07711"},{"label":"HTML","url":"https://arxiv.org/html/2605.07711v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.07711"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.07711v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yilin Cheng Affiliation: Zhongguancun Academy[4pt] sunjie2019@mail.ustc.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.07711v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-07711","updatedAt":"2026-08-21T02:28:00Z"},{"type":"preprint","title":"TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos","authors":[{"name":"Hengyi Feng"},{"name":"Hao Liang"},{"name":"Mingrui Chen"},{"name":"Bohan Zeng"},{"name":"Meiyi Qiang"},{"name":"Zhengyang Zhao"},{"name":"Zimo Meng"},{"name":"Zeang Sheng"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Bohan Zeng Affiliation: Peking University Affiliation: Zhongguancun Academy https://heinz217.github.io/TraceAV-Bench-Page"],"relationType":"affiliation","publishedAt":"2026-05-08","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit reasoning chain that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step semi-automated pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Moreover, we find that robustness to multimodal hallucination is largely decoupled from general multimodal reasoning performance. We anticipate that TraceAV-Bench will stimulate further research toward OmniLLMs that can reason coherently and faithfully over long-form audio-visual content.","identifiers":{"arxiv":"2605.07593","doi":"10.48550/arXiv.2605.07593"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.07593"},{"label":"HTML","url":"https://arxiv.org/html/2605.07593v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.07593"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.07593v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Bohan Zeng Affiliation: Peking University Affiliation: Zhongguancun Academy https://heinz217.github.io/TraceAV-Bench-Page","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.07593v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-07593","updatedAt":"2026-08-21T02:28:00Z"},{"type":"preprint","title":"Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training","authors":[{"name":"Chen Wang"},{"name":"Hexuan Deng"},{"name":"Yining Zhang"},{"name":"Yuchen Zhang"},{"name":"Jionghao Bai"},{"name":"Zhaochun Li"},{"name":"Ge Lan"},{"name":"Yue Wang"}],"institutions":["zgca"],"rawAffiliations":["Chen Wang 1,2,∗ Hexuan Deng 2,3 Yining Zhang 2,4 Yuchen Zhang 2,5 Jionghao Bai 2,6 Zhaochun Li 2,7 Ge Lan 1,† Yue Wang 2,† 1 College of Software, Nankai University 2 Zhongguancun Academy 3 Harbin Institute of Technology 4 Institute of Automation, Chinese Academy of Sciences 5 East China Normal University 6 Zhejiang University 7 Beijing Institute of Technology"],"relationType":"affiliation","publishedAt":"2026-05-08","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Reinforcement learning with verifiable rewards improves LLM reasoning but often induces overthinking, where models generate unnecessarily long reasoning traces. Existing methods mainly rely on length penalties or early-exit strategies; however, the former may degrade accuracy and induce underthinking, whereas the latter assumes that substantial portions of reasoning traces can be safely truncated. To obtain a compression signal without these limitations, we revisit the training dynamics of existing compression methods. We observe that the length--accuracy correlation is initially negative but continually increases during compression, indicating that shorter responses are initially more likely to be correct but gradually lose this property as the policy moves toward underthinking. Based on this observation, we formalize overthinking: a negative correlation indicates an overthinking regime, while a positive one indicates underthinking. When overthinking, the shortest correct responses are shorter than the group-average response length in expectation, making them natural compression targets already present in on-policy rollouts. We therefore propose \\emph{Implicit Compression Regularization} (ICR), an on-policy regularization method whose compression signal comes from a virtual shorter distribution induced by the shortest correct responses in rollout groups, guiding the policy toward concise yet correct trajectories. Training dynamics show that ICR maintains a better length--accuracy correlation during compression, indicating that short responses remain better aligned with correctness instead of drifting toward underthinking. Experiments on three reasoning backbones and multiple mathematical and knowledge-intensive benchmarks show that ICR consistently shortens responses while preserving or improving accuracy, achieving a stronger accuracy--length Pareto frontier.","identifiers":{"arxiv":"2605.07316","doi":"10.48550/arXiv.2605.07316"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.07316"},{"label":"HTML","url":"https://arxiv.org/html/2605.07316v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.07316"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.07316v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chen Wang 1,2,∗ Hexuan Deng 2,3 Yining Zhang 2,4 Yuchen Zhang 2,5 Jionghao Bai 2,6 Zhaochun Li 2,7 Ge Lan 1,† Yue Wang 2,† 1 College of Software, Nankai University 2 Zhongguancun Academy 3 Harbin Institute of Technology 4 Institute of Automation, Chinese Academy of Sciences 5 East China Normal University 6 Zhejiang University 7 Beijing Institute of Technology","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.07316v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-07316","updatedAt":"2026-08-21T02:28:00Z"},{"type":"preprint","title":"LDFE: Laplacian Decoupled Feature Enhancement block for dual-stream CNN-based RGB-IR object detection","authors":[{"name":"Wenhao Dong"},{"name":"Xiaoyan Luo"},{"name":"Linlin Yang"},{"name":"Haodong Zhu"},{"name":"Xiaorong Shi"},{"name":"Guo G"},{"name":"B Zhang"}],"institutions":["zgca"],"rawAffiliations":["Haodong Zhu Affiliation: The School of Artificial Intelligence, Beihang University, Beijing, 100191, China Affiliation: Zhongguancun Academy, Beijing, 100094, China"],"relationType":"affiliation","publishedAt":"2026-05-05","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"","identifiers":{"arxiv":"2607.08076","doi":"10.48550/arXiv.2607.08076","publishedDoi":"10.1016/j.patcog.2026.113935"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2607.08076"},{"label":"HTML","url":"https://arxiv.org/html/2607.08076v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2607.08076"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2607.08076v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haodong Zhu Affiliation: The School of Artificial Intelligence, Beihang University, Beijing, 100191, China Affiliation: Zhongguancun Academy, Beijing, 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2607.08076v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2607-08076","updatedAt":"2026-08-13T13:11:50Z"},{"type":"preprint","title":"Embody4D: A Generalist Data Engine for Embodied 4D World Modeling","authors":[{"name":"Peiyan Tu"},{"name":"Hanxin Zhu"},{"name":"Jingwen Sun"},{"name":"Shaojie Ren"},{"name":"Cong Wang"},{"name":"Yuyan Xu"},{"name":"Jiayi Luo"},{"name":"Xiaoqian Cheng"},{"name":"Z H Chen"}],"institutions":["zgca"],"rawAffiliations":["Jingwen Sun Affiliation: Beijing Zhongguancun Academy Affiliation: University of Science and Technology of China"],"relationType":"affiliation","publishedAt":"2026-05-03","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However, existing robot data are typically captured from fixed or sparse viewpoints, providing only partial and view-dependent observations, which limits multi-view perception and generalization across viewpoints. Given the difficulty of collecting additional viewpoints in real-world settings, we propose Embody4D, a dedicated video-to-video world model for embodied scenarios to bridge this observation gap by transforming a monocular robot video into novel-view videos from flexible target camera viewpoints. First, to tackle training data scarcity, we introduce a 3D-aware compositional synthesis pipeline to curate a heterogeneous dataset compositing cross-embodiment robotic arms with diverse backgrounds, promoting broad generalization. Second, to enforce geometric stability, we devise a latent confidence-aware expert modulation strategy, which estimates the reliability of warped latent priors and adaptively routes regions to copy, repair, or inpaint experts for spatiotemporally consistent 4D generation. Finally, to enhance the fidelity of the manipulation, we incorporate an interaction-aware attention mechanism that explicitly attends to the robotic interaction regions. Extensive experiments show that Embody4D achieves state-of-the-art performance on visual evaluation benchmarks, while both simulated and real-world robotic experiments further demonstrate its effectiveness as a robust data engine for synthesizing high-fidelity, view-consistent videos that empower downstream robotic planning and learning.","identifiers":{"arxiv":"2605.01799","doi":"10.48550/arXiv.2605.01799"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.01799"},{"label":"HTML","url":"https://arxiv.org/html/2605.01799v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.01799"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.01799v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jingwen Sun Affiliation: Beijing Zhongguancun Academy Affiliation: University of Science and Technology of China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.01799v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-01799","updatedAt":"2026-08-20T02:19:27Z"},{"type":"preprint","title":"Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay","authors":[{"name":"Jin Tong"},{"name":"Guang Liang"},{"name":"Peilin Sun"},{"name":"Jianxin Wu"}],"institutions":["zgca"],"rawAffiliations":["Guang Liang Affiliation: State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China Affiliation: School of Artificial Intelligence, Nanjing University, Nanjing 210023, China Affiliation: Zhongguancun Academy, Beijing 100094, China {tongj, liangg, sunpl}@lamda.nju.edu.cn, wujx2001@nju.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-02","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Low-bit quantization is a practical route for efficiently deploying vision Transformers, yet activation outliers complicate fully quantized deployment. Existing methods either handle quantization post-training or suppress large activations during training; however, aggressively restricting outliers in vision models can lead to a poorer trade-off between full-precision and quantized accuracy. We argue that rather than simply suppressing outliers, the training objective should control the structural amplification that makes them harmful. To this end, we introduce Colinearity-Decay (CD), a structural regularizer for ordered matrix pairs within Transformer blocks. CD penalizes detrimental cross-matrix alignment and mitigates extreme activations without altering the architecture or task loss. Applied as a decoupled update, CD is non-invasive and introduces minimal training overhead. Across ImageNet-1K pre-training, COCO detection, and downstream fine-tuning, CD consistently boosts quantized accuracy across multiple pipelines while preserving, or even improving, full-precision performance. Ultimately, our results demonstrate that structural regularization effectively prepares vision Transformers for low-bit deployment with zero inference-time overhead.","identifiers":{"arxiv":"2605.01330","doi":"10.48550/arXiv.2605.01330"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.01330"},{"label":"HTML","url":"https://arxiv.org/html/2605.01330v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.01330"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.01330v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Guang Liang Affiliation: State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China Affiliation: School of Artificial Intelligence, Nanjing University, Nanjing 210023, China Affiliation: Zhongguancun Academy, Beijing 100094, China {tongj, liangg, sunpl}@lamda.nju.edu.cn, wujx2001@nju.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.01330v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-01330","updatedAt":"2026-08-20T02:19:27Z"},{"type":"preprint","title":"LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks","authors":[{"name":"Jiayong Wan"},{"name":"Jiawei Chen"},{"name":"Zhaoxia Yin"},{"name":"Liu Shuyuan"},{"name":"Hang Su"}],"institutions":["zgca"],"rawAffiliations":["Jiawei Chen Affiliation: East China Normal University, Shanghai, China Affiliation: Beijing Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-05-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL"],"abstract":"Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context reward hacking (ICRH), a phenomenon where LLMs iteratively optimize their behavior to maximize proxy objectives, inadvertently producing harmful side effects. Existing defense methods are insufficient to address this risk, as ICRH arises not from adversarial inputs but from the model's own over-optimization. To mitigate this issue, we propose \\textbf{LLM-based Constraint Optimization (LCO)}, a framework that effectively reduces ICRH without model fine-tuning. LCO consists of two modules: \\textit{self-thought module}, which guides the LLM to proactively deliberate and integrate potential safety constraints before execution; and \\textit{evolutionary sampling module}, which employs LLM-based crossover and mutation to constrain the model's actions within a safe solution space while maintaining task performance. Experimental results demonstrate that LCO substantially alleviates ICRH in both output-refine and policy-refine scenarios. In particular, on the tweet engagement optimization task, LCO achieves a 39% reduction in the Toxicity Growth Rate (TGR) on GPT-4, while on the policy optimization benchmark, it reduces the ICRH Occurrence Rate by 15.23%, demonstrating safety improvement without sacrificing task performance.","identifiers":{"arxiv":"2605.27375","doi":"10.48550/arXiv.2605.27375"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.27375"},{"label":"HTML","url":"https://arxiv.org/html/2605.27375v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.27375"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.27375v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiawei Chen Affiliation: East China Normal University, Shanghai, China Affiliation: Beijing Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.27375v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-27375","updatedAt":"2026-09-13T05:26:44Z"},{"type":"preprint","title":"Targeted Remasking: Replacing Token Editing with Token-to-Mask Refinement in Discrete Diffusion Language Models","authors":[{"name":"Lin Yao"}],"institutions":["zgca"],"rawAffiliations":["Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-05-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL","cs.AI"],"abstract":"Discrete masked diffusion language models such as LLaDA generate text through iterative denoising, where mask tokens are progressively replaced with predicted tokens. LLaDA2.1 introduced a Token-to-Token (T2T) editing mechanism that accelerates generation by directly replacing committed tokens suspected of being incorrect. However, we identify fundamental limitations of T2T editing: it couples error detection with replacement, pollutes the generation context with potentially incorrect tokens, and introduces a train-inference noise mismatch where systematic model-generated errors differ from the random perturbations seen during training. We propose Token-to-Mask (T2M) remasking, a training-free, drop-in replacement for T2T editing that resets suspected erroneous tokens back to the mask state, allowing the diffusion process to re-predict them under cleaner context. We design and empirically validate three complementary error detection strategies -- probability-based, trigger-mirrored, and temporal-difference-based -- and provide a unified theoretical analysis showing that T2M remasking purifies the generation context, converts systematic inference errors back to the model's native mask noise type, and enables delayed commitment for joint multi-position optimization. Comprehensive experiments across 12 benchmarks spanning knowledge, reasoning, mathematics, coding, and instruction following show that T2M generally improves performance on tasks requiring precise token-level output, with the largest gain on mathematics (+5.92% on CMATH). Error analysis on CMATH reveals that the dominant failure mode is last-mile token corruption -- where correct reasoning produces a corrupted final answer -- and that T2M repairs 59.4% of such cases.","identifiers":{"arxiv":"2605.26436","doi":"10.48550/arXiv.2605.26436"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.26436"},{"label":"HTML","url":"https://arxiv.org/html/2605.26436v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.26436"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.26436v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.26436v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-26436","updatedAt":"2026-09-12T05:10:53Z"},{"type":"preprint","title":"Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision","authors":[{"name":"Hongyu Gu"},{"name":"Jingwen Fu"}],"institutions":["zgca"],"rawAffiliations":["Jingwen Fu † † thanks: Corresponding author. Affiliation: Zhongguancun Academy Affiliation: Beijing, China Email: jwfu99@gmail.com"],"relationType":"affiliation","publishedAt":"2026-05-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI"],"abstract":"Superposition allows Transformers to reason in depth, carrying an entire reasoning frontier in parallel through a bounded-depth forward pass instead of unrolling serial chain-of-thought tokens. While Zhu et al. (2025) hand-crafted an equal-weight breadth-first frontier in a single residual stream for graph reachability, it remained open whether gradient descent could ever find this target amidst permutation-symmetric saddles. We close this gap on Reachability-by-Superposition over Erdős-Rényi graphs by isolating architectural and supervisional contributions. Architecturally, we identify a Möbius attractor: under $S_n$-symmetry in the tree regime, layerwise dynamics reduce to a 1D Möbius map whose zero set is a codimension-one manifold of global optima containing the equal-weight superposition state. On the supervision side, we identify Cascade Supervision: a loss class whose backward pass simultaneously delivers (A) selectivity bootstrap, (B) gradient persistence across depth, and (C) per-step discrimination (e.g., \\mathcal{L}_{sup} and \\mathcal{L}_{node}). End-to-end supervision fails condition (B) and is provably insufficient: internal gradients at layer c decay as (np)^{-(D-c-2)/2} in the graph fan-out and stall before the manifold is reached. Our thesis: Möbius attractor + Cascade Supervision = emergence of superposition reasoning. The parameter-free decay law predicts a final-step cosine of 0.35 vs. 0.71 (end-to-end vs. cascade) at depth D=3; experiments confirm 0.37 vs. 0.69, matching within 0.02 at every step.","identifiers":{"arxiv":"2605.18820","doi":"10.48550/arXiv.2605.18820"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.18820"},{"label":"HTML","url":"https://arxiv.org/html/2605.18820v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.18820"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.18820v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jingwen Fu † † thanks: Corresponding author. Affiliation: Zhongguancun Academy Affiliation: Beijing, China Email: jwfu99@gmail.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.18820v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-18820","updatedAt":"2026-09-11T05:18:36Z"},{"type":"preprint","title":"Position-Aware Drafting for Inference Acceleration in LLM-Based Generative List-Wise Recommendation","authors":[{"name":"Jiaju Chen"},{"name":"Chongming Gao"},{"name":"Chenxiao Fan"},{"name":"Haoyan Liu"},{"name":"Qingpeng Cai"},{"name":"Peng Jiang"},{"name":"Xiangnan He"}],"institutions":["zgca"],"rawAffiliations":["Jiaju Chen Affiliation: University of Science and Technology of China, Hefei, China Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-04-30","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Large language model (LLM)-based generative list-wise recommendation has advanced rapidly, but decoding remains sequential and thus latency-prone. To accelerate inference without changing the target distribution, speculative decoding (SD) uses a small draft model to propose several next tokens at once and a target LLM to verify and accept the longest prefix, skipping multiple steps per round. In generative recommendation, however, each item is represented by multiple semantic-ID tokens, often with separators, and current drafts typically treat these tokens uniformly. This overlooks two practical facts: (i) a token's semantics depend on its within-item slot, and (ii) uncertainty tends to increase with speculation depth. Without modeling these effects, SD's speedups can be limited. We introduce PAD-Rec, Position-Aware Drafting for generative Recommendation, a lightweight module that augments the draft model with two complementary signals. Item position embeddings explicitly encode the within-item slot of each token, strengthening structural awareness. Step position embeddings encode the draft step, allowing the model to adapt to depth-dependent uncertainty and improve proposal quality. To harmonize these signals with base features, we add simple gates: a learnable coefficient for item slots and a context-driven gate for draft steps. The module is trainable, easy to integrate with standard draft models, and adds negligible inference overhead. Extensive experiments on four real-world datasets show up to 3.1x wall-clock speedup and about 5% average wall-clock speedup gain over strong SD baselines, while largely preserving recommendation quality.","identifiers":{"arxiv":"2604.27747","doi":"10.48550/arXiv.2604.27747"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.27747"},{"label":"HTML","url":"https://arxiv.org/html/2604.27747v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.27747"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.27747v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiaju Chen Affiliation: University of Science and Technology of China, Hefei, China Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.27747v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-27747","updatedAt":"2026-08-20T02:19:27Z"},{"type":"preprint","title":"Generative structure search for efficient and diverse discovery of molecular and crystal structures","authors":[{"name":"Yifang Qin"},{"name":"Yu Shi"},{"name":"Junfu Tan"},{"name":"Chang Liu"},{"name":"Ming Zhang"},{"name":"Ziheng Lu"}],"institutions":["zgca"],"rawAffiliations":["Yu Shi Email: shiyu@bza.edu.cn Affiliation: Zhongguancun Academy, Beijing, 100094, China"],"relationType":"affiliation","publishedAt":"2026-04-30","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Predicting stable and metastable structures is central to molecular and materials discovery, but remains limited by the cost of searching high-dimensional energy landscapes. Deep generative models offer efficient structure sampling, yet their outputs remain shaped by training data and can underexplore minima that are rare but physically relevant. We introduce generative structure search (GSS), a unified framework that formulates diffusion-based generation and random structure search (RSS) as limiting regimes of a common sampling process driven by learned score fields and physical forces. Coupling these drivers lets GSS use data priors to accelerate sampling while retaining energy-guided exploration of local minima. Across molecular and crystalline systems, GSS recovers diverse metastable structures with more than tenfold lower sampling cost than RSS for broad coverage and remains effective for compositions outside the training distribution. The results establish a physically grounded generative search strategy for discovering structures beyond the reach of data-driven sampling alone.","identifiers":{"arxiv":"2604.27636","doi":"10.48550/arXiv.2604.27636"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.27636"},{"label":"HTML","url":"https://arxiv.org/html/2604.27636v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.27636"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.27636v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yu Shi Email: shiyu@bza.edu.cn Affiliation: Zhongguancun Academy, Beijing, 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.27636v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-27636","updatedAt":"2026-08-20T02:19:27Z"},{"type":"preprint","title":"Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO","authors":[{"name":"Yu Tian"},{"name":"Jiawei Chen"},{"name":"Lifan Zheng"},{"name":"Mingxiang Tao"},{"name":"Xinyi Zeng"},{"name":"Zhaoxia Yin"},{"name":"Hang Su"},{"name":"Xian Sun"}],"institutions":["zgca"],"rawAffiliations":["† † affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-30","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"We introduce Skills-Coach, a novel automated framework designed to significantly enhance the self-evolution of skills within Large Language Model (LLM)-based agents. Addressing the current fragmentation of the skill ecosystem, Skills-Coach explores the boundaries of skill capabilities, thereby facilitating the comprehensive competency coverage essential for intelligent applications. The framework comprises four core modules: a Diverse Task Generation Module that systematically creates a comprehensive test suite for various skills; a Lightweight Optimization Module dedicated to optimizing skill prompts and their corresponding code; a Comparative Execution Module facilitating the execution and evaluation of both original and optimized skills; and a Traceable Evaluation Module, which rigorously evaluates performance against specified criteria. Skills-Coach offers flexible execution options through its virtual and real modes. To validate its efficacy, we introduce Skill-X, a comprehensive benchmark dataset consisting of 48 diverse skills. Experimental results demonstrate that Skills-Coach achieves significant performance improvements in skill capability across a wide range of categories, highlighting its potential to advance the development of more robust and adaptable LLM-based agents.","identifiers":{"arxiv":"2604.27488","doi":"10.48550/arXiv.2604.27488"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.27488"},{"label":"HTML","url":"https://arxiv.org/html/2604.27488v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.27488"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.27488v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"† † affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.27488v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-27488","updatedAt":"2026-08-20T02:19:27Z"},{"id":"doi-10-1007-s11390-026-5948-8","type":"article","title":"Data Preparation for Large Language Models","authors":[{"name":"Hao Liang","institutions":["zgca"]},{"name":"Zhen Hao Wong"},{"name":"Rui-Tong Liu"},{"name":"Yu-Han Wang"},{"name":"Mei-Yi Qiang"},{"name":"Zheng-Yang Zhao","institutions":["zgca"]},{"name":"Cheng-Yu Shen"},{"name":"Cong-Hui He"},{"name":"Wen-Tao Zhang","institutions":["zgca"]},{"name":"Bin Cui"}],"institutions":["zgca"],"rawAffiliations":["Beijing Zhongguancun Academy, Beijing 100871, China"],"relationType":"affiliation","publishedAt":"2026-04-30","year":2026,"venue":"Journal of Computer Science and Technology","status":"published","topics":["Large Language Models","Data Management","Data Preparation","Survey"],"abstract":"Large language models have demonstrated strong generalization across domains, largely supported by massive amounts of high-quality training data. This survey organizes data preparation algorithms and workflows into pre-training, continual pre-training and post-training, and reviews widely used datasets and associated preparation methods.","identifiers":{"doi":"10.1007/s11390-026-5948-8"},"links":[{"label":"DOI","url":"https://doi.org/10.1007/s11390-026-5948-8"},{"label":"PDF","url":"https://jcst.ict.ac.cn/cn/article/pdf/preview/10.1007/s11390-026-5948-8.pdf"},{"label":"Project","url":"https://github.com/haolpku/Awesome-LLM-Data-Preparation"}],"versions":[{"label":"Publisher version","url":"https://www.sciopen.com/article/10.1007/s11390-026-5948-8"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beijing Zhongguancun Academy, Beijing 100871, China","source":"JCST publisher PDF","sourceUrl":"https://jcst.ict.ac.cn/cn/article/pdf/preview/10.1007/s11390-026-5948-8.pdf"}],"sources":["Crossref","Publisher"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"preprint","title":"FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards","authors":[{"name":"Zhixin Han"},{"name":"Yanzhi Zhang"},{"name":"Chuyang Wei"},{"name":"Maohang Gao"},{"name":"Xiawei Yue"},{"name":"Kefei Chen"},{"name":"Yu Zhuang"},{"name":"Haoxiang Guan"},{"name":"Jiyan He"},{"name":"Jian Li"},{"name":"Yitong Duan"},{"name":"Yu Shi"},{"name":"Mengting Hu"},{"name":"Shuxin Zheng"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-04-29","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI","cs.LG"],"abstract":"Live future prediction refers to the task of making predictions about real-world events before they unfold. This task is increasingly studied using large language model-based agent systems, and it is important for building agents that can continually learn from the real world. It can provide a large number of prediction questions grounded in diverse real-world events, while preventing answer leakage. To leverage the advantages of future prediction, we present FutureWorld, a live agentic reinforcement learning environment that closes the training loop between prediction, outcome realization, and parameter updates. Specifically, we modify and extend verl-tool, resulting in a new framework that we call verl-tool-future. Unlike standard reinforcement learning training frameworks that rely on immediate rewards, verl-tool-future stores prediction-time rollouts, backfills rewards after real-world outcomes become available, and then replays the completed trajectories for policy update. Across three open-source agents, successive FutureWorld training rounds lead to consistent improvements in prediction accuracy, probabilistic scoring, and calibration, demonstrating that delayed real-world outcome feedback can serve as an effective reinforcement learning signal.","identifiers":{"arxiv":"2604.26733","doi":"10.48550/arXiv.2604.26733"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.26733"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.26733"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_8t639n2lzd7p6v1xwxost2goxugu6xoa"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.26733v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_8t639n2lzd7p6v1xwxost2goxugu6xoa"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2604-26733","updatedAt":"2026-08-08T11:59:38Z"},{"id":"rebalance-iclr-2026","type":"conference","title":"Efficient Reasoning with Balanced Thinking","authors":[{"name":"Yulin Li"},{"name":"Tengyao Tu","institutions":["zgca"]},{"name":"Li Ding"},{"name":"Junjie Wang"},{"name":"Huiling Zhen"},{"name":"Yixin Chen"},{"name":"Yong Li","institutions":["zgca"]},{"name":"Zhuotao Tian"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-24","year":2026,"venue":"ICLR 2026","status":"published","topics":["Large Reasoning Models","Efficient Inference","Reasoning","Model Steering"],"abstract":"ReBalance is a training-free framework for efficient reasoning with balanced thinking. It uses confidence dynamics to identify overthinking and underthinking, then dynamically steers reasoning trajectories to reduce redundant tokens while improving accuracy.","identifiers":{},"links":[{"label":"Project","url":"https://rebalance-ai.github.io/"},{"label":"Code","url":"https://github.com/hkust-nlp/ReBalance"},{"label":"Paper","url":"https://huggingface.co/papers/2510.05340"}],"versions":[{"label":"ICLR project page","url":"https://rebalance-ai.github.io/"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Official project page","sourceUrl":"https://rebalance-ai.github.io/"}],"sources":["Official project page","OpenReview"],"updatedAt":"2026-08-08T00:00:00Z"},{"id":"arxiv-2505-19558","type":"conference","title":"PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives","authors":[{"name":"Zhaowei Zhang"},{"name":"Xiaobo Wang"},{"name":"Minghua Yi"},{"name":"Mengmeng Wang"},{"name":"Fengshuo Bai"},{"name":"Zilong Zheng","institutions":["zgca"]},{"name":"Yipeng Kang"},{"name":"Yaodong Yang","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-24","year":2026,"venue":"ICLR 2026","status":"published","topics":["Large Language Models","Political Consensus","AI Evaluation","Social Choice"],"abstract":"PoliCon evaluates whether large language models can draft consensus resolutions from divergent political party positions. It contains 2,225 European Parliament deliberation records and uses a social-choice-based evaluation framework to test multiple consensus objectives.","identifiers":{"arxiv":"2505.19558","doi":"10.48550/arXiv.2505.19558"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2505.19558"},{"label":"PDF","url":"https://arxiv.org/pdf/2505.19558"},{"label":"Project","url":"https://zowiezhang.github.io/PoliCon/"}],"versions":[{"label":"arXiv v3","url":"https://arxiv.org/abs/2505.19558v3"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"ICLR paper PDF","sourceUrl":"https://arxiv.org/pdf/2505.19558"}],"sources":["arXiv","OpenReview"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"conference","title":"Quotient-Space Diffusion Models","authors":[{"name":"Yixian Xu"},{"name":"Yusong Wang"},{"name":"Shengjie Luo"},{"name":"Kaiyuan Gao"},{"name":"Tianyu He"},{"name":"Di He"},{"name":"Chang Liu"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-04-23","year":2026,"venue":"ICLR 2026","status":"published","topics":["cs.LG","cs.AI","q-bio.QM","stat.ML"],"abstract":"Diffusion-based generative models have reformed generative AI, and also enabled new capabilities in the science domain, e.g., fast generation of 3D structures of molecules. In such tasks, there is often a symmetry in the system, identifying elements that can be converted by certain transformations as equivalent. Equivariant diffusion models guarantee a symmetric distribution, but miss the opportunity to make learning easier, while alignment-based simplification attempts fail to preserve the target distribution. In this work, we develop quotient-space diffusion models, a principled generative framework to fully handle and leverage symmetry. By viewing the intrinsic generation process on the quotient space, the exact construction that removes symmetry redundancy, the framework simplifies learning by allowing model output to have an arbitrary intra-equivalence-class movement, while generating the correct symmetric target distribution with guarantee. We instantiate the framework for molecular structure generation which follows $\\mathrm{SE}(3)$ (rigid-body movement) symmetry. It improves the performance over equivariant diffusion models and outperforms alignment-based methods universally for small molecules and proteins, representing a new framework that surpasses previous symmetry treatments in generative models.","identifiers":{"arxiv":"2604.21809","doi":"10.48550/arXiv.2604.21809"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.21809"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.21809"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_j8oft0n01fkw77lwg26q7682its89suf"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.21809v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_j8oft0n01fkw77lwg26q7682its89suf"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2604-21809","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Geometry-Guided 3D Visual Token Pruning for Video-Language Models","authors":[{"name":"Han Li"},{"name":"Zehao Huang"},{"name":"Jiahui Fu"},{"name":"Naiyan Wang"},{"name":"Si Liu"}],"institutions":["zgca"],"rawAffiliations":["Han Li Affiliation: Beihang University Affiliation: Zhongguancun Academy Email: lihan0620@buaa.edu.cn"],"relationType":"affiliation","publishedAt":"2026-04-20","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as 3D spatial videos composed of image sequences with depth and camera pose information, enabling pre-trained video-language models to perform 3D reasoning tasks. However, the large number of visual tokens in spatial videos remains a major bottleneck for efficient inference and context management. Existing pruning methods overlook the view consistency of spatial videos and the spatial diversity of the remaining tokens, which prevents them from effectively removing inter-frame redundancy and preserving scene completeness. In this paper, we propose Geo3DPruner, a Geometry-Guided 3D Visual Token Pruning framework. Geo3DPruner first models cross-frame relevance through geometry-aware global attention, and then performs a two-stage pruning process. The intra-voxel stage selects representative multi-view features within each voxel, while the inter-voxel stage preserves spatial diversity by selecting a globally distributed subset of voxels. Extensive experiments on multiple 3D scene understanding benchmarks demonstrate that Geo3DPruner retains over 90% of the original performance while pruning 90% of visual tokens, significantly outperforming existing text-guided and vision-guided pruning methods.","identifiers":{"arxiv":"2604.18260","doi":"10.48550/arXiv.2604.18260"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.18260"},{"label":"HTML","url":"https://arxiv.org/html/2604.18260v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.18260"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.18260v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Han Li Affiliation: Beihang University Affiliation: Zhongguancun Academy Email: lihan0620@buaa.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.18260v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-18260","updatedAt":"2026-08-13T11:57:17Z"},{"type":"preprint","title":"Harnessing Pre-Resolution Signals for Future Prediction Agents","authors":[{"name":"Chuyang Wei"},{"name":"Maohang Gao"},{"name":"Zhixin Han"},{"name":"Kefei Chen"},{"name":"Yu Zhuang"},{"name":"Haoxiang Guan"},{"name":"Yanzhi Zhang"},{"name":"Yilin Cheng"},{"name":"Xiren Zhou"},{"name":"Huanhuan Chen"},{"name":"Jian Li"},{"name":"Jiyan He"},{"name":"Yu Shi"},{"name":"Yitong Duan"},{"name":"Shuxin Zheng"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-04-17","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI"],"abstract":"Many high-stakes decisions depend on forecasts made before outcomes are known. In this future prediction setting, the central challenge is that public evidence evolves over time, while the main supervision signal arrives only after resolution: the realized outcome mainly assesses final correctness, offering only coarse guidance on what to track, what to verify, and which judgments to leave uncertain along the way. Our key observation is that revisiting the same unresolved question over time creates informative temporal contrasts across evolving evidence and repeated forecasts, exposing what earlier attempts missed before resolution and yielding a diagnostic signal we call the pre-resolution signal. We instantiate this idea in Milkyway, a future prediction agent with a persistent future prediction harness, an editable external state that stores reusable procedural guidance across revisits to the same unresolved question. As the same unresolved question is revisited, Milkyway extracts pre-resolution signals from evolving evidence and repeated forecasts, uses them to update the harness, and improves later forecasts on that question before resolution. After resolution, the realized outcome serves as a post-resolution check of provisional updates. On the FutureX and FutureWorld benchmarks, Milkyway achieves strong performance against competitive baselines, and a mechanism study suggests that the gains stem from harness evolution driven by pre-resolution signals rather than repeated prediction alone.","identifiers":{"arxiv":"2604.15719","doi":"10.48550/arXiv.2604.15719"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.15719"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.15719"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_8t639n2lzd7p6v1xwxost2goxugu6xoa"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.15719v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_8t639n2lzd7p6v1xwxost2goxugu6xoa"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2604-15719","updatedAt":"2026-08-08T11:59:38Z"},{"type":"conference","title":"CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors","authors":[{"name":"Hang Su"},{"name":"Zequn Liu"},{"name":"Chen Hu"},{"name":"Xuesong Lu"},{"name":"Yingce Xia"},{"name":"Zhen Liu"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面","Hang Su Affiliation: East China Normal University, China. Affiliation: Zhongguancun Academy, China. Email: s-sh25@bza.edu.cn"],"relationType":"official-output","publishedAt":"2026-04-16","year":2026,"venue":"ACL 2026","status":"published","topics":["cs.CL"],"abstract":"While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradigms predominantly rely on lexical-level similarity or manual heuristics, often lacking sufficient data-driven validation. We address this by mining Community-Individual Preference Divergence (CIPD), where individual choices override consensus, to distill six key personalization factors as evaluative dimensions. Accordingly, we introduce CoPA, a benchmark with 1,985 user profiles for fine-grained, factor-level assessment. By quantifying the alignment between model outputs and user-specific cognitive preferences inferred from interaction patterns, CoPA provides a more comprehensive and discriminative standard for evaluating personalized QA than generic metrics. The code is available at https://github.com/bjzgcai/CoPA.","identifiers":{"arxiv":"2604.14773","doi":"10.48550/arXiv.2604.14773"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.14773"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.14773"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_dogy7yn9wqq8w3vdfu0gv0gnzmz5nhtm"},{"label":"Code","url":"https://github.com/bjzgcai/CoPA"},{"label":"HTML","url":"https://arxiv.org/html/2604.14773v1"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.14773v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_dogy7yn9wqq8w3vdfu0gv0gnzmz5nhtm"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Hang Su Affiliation: East China Normal University, China. Affiliation: Zhongguancun Academy, China. Email: s-sh25@bza.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.14773v1"}],"sources":["arXiv","北京中关村学院官网","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-14773","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Learning Shared Sentiment Prototypes for Adaptive Multimodal Sentiment Analysis","authors":[{"name":"Chen Su"},{"name":"Yuanhe Tian"},{"name":"Yan Song"}],"institutions":["zgca"],"rawAffiliations":["Yuanhe Tian Affiliation: Zhongguancun Academy email: yhtian94@gmail.com"],"relationType":"affiliation","publishedAt":"2026-04-07","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Multimodal sentiment analysis (MSA) aims to predict human sentiment from textual, acoustic, and visual information in videos. Recent studies improve multimodal fusion by modeling modality interaction and assigning different modality weights. However, they usually compress diverse sentiment cues into a single compact representation before sentiment reasoning. This early aggregation makes it difficult to preserve the internal structure of sentiment evidence, where different cues may complement, conflict with, or differ in reliability from each other. In addition, modality importance is often determined only once during fusion, so later reasoning cannot further adjust modality contributions. To address these issues, we propose PRISM, a framework that unifies structured affective extraction and adaptive modality evaluation. PRISM organizes multimodal evidence in a shared prototype space, which supports structured cross-modal comparison and adaptive fusion. It further applies dynamic modality reweighting during reasoning, allowing modality contributions to be continuously refined as semantic interactions become deeper. Experiments on three benchmark datasets show that PRISM outperforms representative baselines.","identifiers":{"arxiv":"2604.05873","doi":"10.48550/arXiv.2604.05873"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.05873"},{"label":"HTML","url":"https://arxiv.org/html/2604.05873v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.05873"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.05873v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuanhe Tian Affiliation: Zhongguancun Academy email: yhtian94@gmail.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.05873v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-05873","updatedAt":"2026-09-06T06:18:47Z"},{"type":"preprint","title":"BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination","authors":[{"name":"Xingyu Peng"},{"name":"Chen Gao"},{"name":"Liankai Jin"},{"name":"A.B. Li"},{"name":"Si Liu"}],"institutions":["zgca"],"rawAffiliations":["Xingyu Peng Note: Equal Contribution. email: pengxyai@buaa.edu.cn Affiliation: Beihang University , Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-07","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Bimanual manipulation, i.e., the coordinated use of two robotic arms to complete tasks, is essential for achieving human-level dexterity in robotics. Recent simulation benchmarks, e.g., RoboTwin and RLBench2, have advanced data-driven learning for bimanual manipulation. However, existing tasks are short-horizon and only loosely coordinated, failing to capture the spatial-temporal coupling inherent in real-world bimanual behaviors. To address this gap, we introduce BiCoord, a benchmark for long-horizon and tightly coordinated bimanual manipulation. Specifically, BiCoord comprises diverse tasks that require continuous inter-arm dependency and dynamic role exchange across multiple sub-goals. Also, we propose a suite of quantitative metrics that evaluate coordination from temporal, spatial, and spatial-temporal perspectives, enabling systematic measurement of bimanual cooperation. Experimental results show that representative manipulation policies, e.g., DP, RDT, Pi0, and OpenVLA-OFT, struggle with long-duration and highly coupled tasks, revealing fundamental challenges in achieving long-horizon and tight coordination tasks. We hope BiCoord can serve as a foundation for studying long-horizon cooperative manipulation and inspire future research on coordination-aware robotic learning. All datasets, codes and supplements could be found at https://buaa-colalab.github.io/BiCoord/.","identifiers":{"arxiv":"2604.05831","doi":"10.48550/arXiv.2604.05831"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.05831"},{"label":"HTML","url":"https://arxiv.org/html/2604.05831v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.05831"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.05831v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xingyu Peng Note: Equal Contribution. email: pengxyai@buaa.edu.cn Affiliation: Beihang University , Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.05831v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-05831","updatedAt":"2026-08-13T11:18:16Z"},{"type":"conference","title":"AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery","authors":[{"name":"Yu Li"},{"name":"Chenyang Shao"},{"name":"Xinyang Liu"},{"name":"Ruotong Zhao"},{"name":"Peijie Liu"},{"name":"Hongyuan Su"},{"name":"Zhibin Chen"},{"name":"Qinglong Yang"},{"name":"Anjie Xu"},{"name":"Yi Fang"},{"name":"Qingbin Zeng"},{"name":"Tianxing Li"},{"name":"Jingbo Xu"},{"name":"Fengli Xu"},{"name":"Yong Li"},{"name":"Tie-Yan Liu"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-04-07","year":2026,"venue":"CVPR 2026","status":"published","topics":["cs.CL","cs.CE"],"abstract":"Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.","identifiers":{"arxiv":"2604.05550","doi":"10.48550/arXiv.2604.05550"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.05550"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.05550"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_d0qhgfwk5u8c45svd8cmydi5eh526adr"},{"label":"Code","url":"https://github.com/tsinghua-fib-lab/AutoSOTA"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.05550v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_d0qhgfwk5u8c45svd8cmydi5eh526adr"},{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_38df3gx56izdt48elmvdwcldz7upshy1"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2604-05550","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"AgentEconomist: An End-to-end Agentic System Translating Economic Intuitions into Executable Computational Experiments","authors":[{"name":"Jiaju Chen"},{"name":"Jinghua Piao"},{"name":"Xia Xu"},{"name":"Songwei Li"},{"name":"Tong Xia"},{"name":"Xiangnan He"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Jiaju Chen Note: Equal contribution. email: cjj01@mail.ustc.edu.cn Affiliation: Zhongguancun Academy , Beijing , China Affiliation: University of Science and Technology of China , Hefei , China"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.HC","cs.AI"],"abstract":"A long-standing challenge in economics lies not in the lack of intuition, but in the difficulty of translating intuitive insights into verifiable research. To address this challenge, we introduce AgentEconomist, an end-to-end interactive system designed to translate abstract intuitions into executable computational experiments. Grounded in a domain-specific knowledge base covering over 13,000 high-quality academic papers, the system employs a modular multi-stage architecture. Specifically, the Idea Development Stage generates literature-grounded hypotheses, the Experimental Design Stage configures simulator-aligned experimental parameters and protocols, and the Experimental Execution Stage runs experiments and returns structured analyses. Together, these stages form a human-in-the-loop, iterative workflow that translates economic intuitions into executable computational experiments. Through extensive experiments involving human expert evaluation and large language models (LLMs) as judges, we show that the system generates research ideas with stronger literature grounding and higher novelty and insight than state-of-the-art generic LLMs. Overall, AgentEconomist adopts a human-AI collaboration paradigm that enables researchers to focus on high-level intuitions, while delegating the labor-intensive processes of translation and computational execution to agents.","identifiers":{"arxiv":"2604.27725","doi":"10.48550/arXiv.2604.27725"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.27725"},{"label":"HTML","url":"https://arxiv.org/html/2604.27725v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.27725"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.27725v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiaju Chen Note: Equal contribution. email: cjj01@mail.ustc.edu.cn Affiliation: Zhongguancun Academy , Beijing , China Affiliation: University of Science and Technology of China , Hefei , China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.27725v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-27725","updatedAt":"2026-09-09T05:41:51Z"},{"type":"preprint","title":"DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models","authors":[{"name":"JiYang Wang"},{"name":"Jiawei Chen"},{"name":"Mengqi Xiao"},{"name":"Yu Cheng"},{"name":"Yangfu Li"},{"name":"Zhaoxia Yin"}],"institutions":["zgca"],"rawAffiliations":["Jiawei Chen ∗ Affiliation: Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.AI"],"abstract":"Object level hallucination remains a central reliability challenge for vision language models (VLMs), particularly in binary object existence verification. Existing benchmarks emphasize aggregate accuracy but rarely disentangle whether errors stem from perceptual limitations or from the influence of contextual textual priors, leaving underlying failure mechanisms ambiguous. We introduce DO-Bench, a controlled diagnostic benchmark that isolates these sources through structured multimodal interventions. Rather than evaluating models in unconstrained settings, DO-Bench probes two complementary dimensions: the Prior Override dimension progressively strengthens contextual textual priors while holding visual evidence constant to assess resistance to prior pressure, and the Perception-Limited dimension incrementally enhances visual evidence from full-scene context to localized object crops to measure perceptual grounding strength. This paired design enables attribution of errors to prior suppression, perceptual insufficiency, or their interaction. We further define two diagnostic metrics, PriorRobust and PerceptionAbility, to quantify these behaviors consistently. Evaluations across diverse open- and closed-source VLMs reveal systematic differences in prior sensitivity and perceptual reliability, demonstrating that object hallucination reflects heterogeneous, mechanism dependent failure patterns beyond aggregate accuracy.","identifiers":{"arxiv":"2604.22822","doi":"10.48550/arXiv.2604.22822"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.22822"},{"label":"HTML","url":"https://arxiv.org/html/2604.22822v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.22822"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.22822v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiawei Chen ∗ Affiliation: Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.22822v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-22822","updatedAt":"2026-09-09T05:41:51Z"},{"type":"preprint","title":"Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Language Models","authors":[{"name":"Lin Yao"}],"institutions":["zgca"],"rawAffiliations":["Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL"],"abstract":"Diffusion language models (dLLMs) generate text through iterative denoising, filling multiple masked positions at each step. Positions filled in the same step are predicted without conditioning on one another's newly filled values and can therefore be mutually inconsistent; once retained, these inconsistencies become context for later predictions. We introduce \\emph{Token-to-Mask} (T2M), a training-free inference-time correction method that identifies low-confidence positions using the model's probability of the current token, remasks them, and reconstructs them in later denoising steps. On dLLMs equipped with correction mechanisms, a single T2M configuration transfers across tasks and models without retuning and broadly improves task metrics over each model's native correction mechanism. In controlled experiments, we decompose correction methods into a detector that identifies suspicious tokens and an action that determines how to revise them. Holding the detector fixed, remasking yields higher task metrics than replacement; across the tested detector--action combinations, current-token-probability detection paired with remasking performs best. Compared with direct editing, T2M converts additional inference compute into performance gains more effectively and, on most tasks, retains a sequential-step advantage over autoregressive token-by-token decoding.","identifiers":{"arxiv":"2604.18738","doi":"10.48550/arXiv.2604.18738"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.18738"},{"label":"HTML","url":"https://arxiv.org/html/2604.18738v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.18738"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.18738v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Lin Yao Affiliation: School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China Affiliation: Zhongguancun Academy, Beijing, 100097, China Email: lin.yao@sjtu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.18738v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-18738","updatedAt":"2026-09-09T05:41:51Z"},{"type":"preprint","title":"Spike-driven Large Language Model","authors":[{"name":"Han Xu"},{"name":"Xuerui Qiu"},{"name":"Baiyu Chen"},{"name":"Xinhao Luo"},{"name":"Xingrun Xing"},{"name":"Jiahong Zhang"},{"name":"Bo Lei"},{"name":"Tiejun Huang"},{"name":"Bo Xu"},{"name":"Guoqi Li"}],"institutions":["zgca"],"rawAffiliations":["Xingrun Xing Address: Zhongguancun Academy, Beijing 100094, China"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.NE","cs.AI"],"abstract":"Current Large Language Models (LLMs) are primarily based on large-scale dense matrix multiplications. Inspired by the brain's information processing mechanism, we explore the fundamental question: how to effectively integrate the brain's spiking-driven characteristics into LLM inference. Spiking Neural Networks (SNNs) possess spike-driven characteristics, and some works have attempted to combine SNNs with Transformers. However, achieving spike-driven LLMs with billions of parameters, relying solely on sparse additions, remains a challenge in the SNN field. To address the issues of limited representational capacity and sparsity in existing spike encoding schemes at the LLM level, we propose SDLLM, a spike-driven large language model that eliminates dense matrix multiplications through sparse addition operations. Specifically, we use the plug-and-play gamma-SQP two-step spike encoding method to ensure that the quantization process aligns with the model's semantic space, mitigating representation degradation caused by binary spikes. Furthermore, we introduce bidirectional encoding under symmetric quantization and membrane potential clipping mechanisms, leading to spike trains with no or low firing counts dominating, significantly reducing the model's spike firing rate, while halving the number of time steps. Experimental results show that SDLLM not only significantly reduces inference costs but also achieves state-of-the-art task performance under the spike-based paradigm. For example, compared to previous spike-based LLMs, SDLLM reduces energy consumption by 7x and improves accuracy by 4.2%. Our model provides inspiration for the architecture design of the next generation of event-driven neuromorphic chips.","identifiers":{"arxiv":"2604.16475","doi":"10.48550/arXiv.2604.16475"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.16475"},{"label":"HTML","url":"https://arxiv.org/html/2604.16475v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.16475"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.16475v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xingrun Xing Address: Zhongguancun Academy, Beijing 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.16475v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-16475","updatedAt":"2026-09-09T05:41:51Z"},{"type":"preprint","title":"Paper2Data: Large-Scale LLM Extraction and Metadata Structuring of Global Urban Data from Scientific Literature","authors":[{"name":"Runwen You"},{"name":"Tong Xia"},{"name":"Jingzhi Wang"},{"name":"Jiankun Zhang"},{"name":"Tengyao Tu"},{"name":"Jinghua Piao"},{"name":"Yi Chang"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Runwen You Affiliation: Jilin University , Zhongguancun Academy , Beijing, China"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.IR","cs.AI"],"abstract":"Urban data support a wide range of applications across multiple disciplines. However, at the global scale, there is no unified platform for urban data discovery. As a result, researchers often have to manually search through websites or scientific literature to identify relevant datasets. To address this problem, we curate an open urban data discovery portal, \\textit{UrbanDataMiner}, which supports dataset-level search and filtering over more than 60{,}000 urban datasets extracted from over 15{,}000 Nature-affiliated publications. \\textit{UrbanDataMiner} is enabled by \\textit{Paper2Data}, a novel large-scale LLM-driven pipeline that automatically identifies dataset mentions in scientific papers and structures them using a unified urban data metadata schema. Human-annotated evaluation demonstrates that \\textit{Paper2Data} achieves high recall (approximately 90\\%) in dataset identification and high field-level precision (above 80\\%). In addition, \\textit{UrbanDataMiner} can retrieve over 9\\% of datasets that are not easily discoverable through general-purpose search engines such as Google. Overall, our work provides the first large-scale, literature-derived infrastructure for urban data discovery and enables more systematic and reusable data-driven research across disciplines. Our code and data are publicly available\\footnote{https://github.com/Yourunwen/Paper2Data}.","identifiers":{"arxiv":"2604.16317","doi":"10.48550/arXiv.2604.16317"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.16317"},{"label":"HTML","url":"https://arxiv.org/html/2604.16317v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.16317"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.16317v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Runwen You Affiliation: Jilin University , Zhongguancun Academy , Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.16317v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-16317","updatedAt":"2026-09-08T05:43:11Z"},{"type":"preprint","title":"Targeted Exploration via Unified Entropy Control for Reinforcement Learning","authors":[{"name":"Chen Wang"},{"name":"Lai Wei"},{"name":"Yanzhi Zhang"},{"name":"Chenyang Shao"},{"name":"Zedong Dan"},{"name":"Weiran Huang"},{"name":"Ge Lan"},{"name":"Yue Wang"}],"institutions":["zgca"],"rawAffiliations":["Chen Wang Affiliation: College of Software, Nankai University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI"],"abstract":"Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used Group Relative Policy Optimization (GRPO) consistently suffers from entropy collapse, causing the policy to converge prematurely and lose diversity. Existing exploration methods introduce additional bias or variance during exploration, making it difficult to maintain optimization stability. We propose Unified Entropy Control for Reinforcement Learning (UEC-RL), a framework that provides targeted mechanisms for exploration and stabilization. UEC-RL activates more exploration on difficult prompts to search for potential and valuable reasoning trajectories. In parallel, a stabilizer prevents entropy from growing uncontrollably, thereby keeping training stable as the model consolidates reliable behaviors. Together, these components expand the search space when needed while maintaining robust optimization throughout training. Experiments on both LLM and VLM reasoning tasks show consistent gains over RL baselines on both Pass@1 and Pass@$k$. On Geometry3K, UEC-RL achieves a 37.9\\% relative improvement over GRPO, indicating that it sustains effective exploration without compromising convergence and underscoring UEC-RL as a key for scaling RL-based reasoning in large models. Our code is available at https://github.com/597358816/UEC-RL.","identifiers":{"arxiv":"2604.14646","doi":"10.48550/arXiv.2604.14646"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.14646"},{"label":"HTML","url":"https://arxiv.org/html/2604.14646v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.14646"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.14646v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chen Wang Affiliation: College of Software, Nankai University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.14646v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-14646","updatedAt":"2026-09-08T05:43:11Z"},{"type":"preprint","title":"WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models","authors":[{"name":"Hongjin Chen"},{"name":"Shangyun Jiang"},{"name":"Tonghua Su"},{"name":"Chen Gao"},{"name":"Xinlei Chen"},{"name":"Yong Li"},{"name":"Zhibo Chen"}],"institutions":["zgca"],"rawAffiliations":["Hongjin Chen 1,5,∗ , Shangyun Jiang 2,∗ , Tonghua Su 1,† , Chen Gao 3,5,† , Xinlei Chen 3 , Yong Li 3 , Zhibo Chen 4,5,† Affiliation: 1 Harbin Institute of Technology 2 Shandong University 3 Tsinghua University 4 University of Science and Technology of China 5 Zhongguancun Academy Affiliation:"],"relationType":"affiliation","publishedAt":"2026-04-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI","cs.CV","cs.RO"],"abstract":"Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predictors, while world models support look-ahead reasoning by imagining future views. Yet predicting a reliable trajectory from a single egocentric observation remains challenging. Current VLMs often generate unstable trajectories, and world models, though able to synthesize plausible futures, do not directly provide the grounded signals needed for navigation learning. This raises a central question: how can generated futures be turned into supervision for grounded trajectory prediction? We present WorldMAP, a teacher--student framework that converts world-model-generated futures into persistent semantic-spatial structure and planning-derived supervision. Its world-model-driven teacher builds semantic-spatial memory from generated videos, grounds task-relevant targets and obstacles, and produces trajectory pseudo-labels through explicit planning. A lightweight student with a multi-hypothesis trajectory head is then trained to predict navigation trajectories directly from vision-language inputs. On Target-Bench, WorldMAP achieves the best ADE and FDE among compared methods, reducing ADE by 18.0% and FDE by 42.1% relative to the best competing baseline, while lifting a small open-source VLM to DTW performance competitive with proprietary models. More broadly, the results suggest that, in embodied navigation, the value of world models may lie less in supplying action-ready imagined evidence than in synthesizing structured supervision for navigation learning.","identifiers":{"arxiv":"2604.07957","doi":"10.48550/arXiv.2604.07957"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.07957"},{"label":"HTML","url":"https://arxiv.org/html/2604.07957v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.07957"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.07957v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hongjin Chen 1,5,∗ , Shangyun Jiang 2,∗ , Tonghua Su 1,† , Chen Gao 3,5,† , Xinlei Chen 3 , Yong Li 3 , Zhibo Chen 4,5,† Affiliation: 1 Harbin Institute of Technology 2 Shandong University 3 Tsinghua University 4 University of Science and Technology of China 5 Zhongguancun Academy Affiliation:","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.07957v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-07957","updatedAt":"2026-09-07T05:45:08Z"},{"type":"conference","title":"GaussianPile: A Unified Sparse Gaussian Splatting Framework for Slice-based Volumetric Reconstruction","authors":[{"name":"Di Kong"},{"name":"Yikai Wang"},{"name":"Wenjie Guo"},{"name":"Yifan Bu"},{"name":"Boya Zhang"},{"name":"Yuexin Duan"},{"name":"Xiawei Yue"},{"name":"Wenbiao Du"},{"name":"Yiman Zhong"},{"name":"Yuwen Chen"},{"name":"Cheng Ma"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-03-21","year":2026,"venue":"CVPR 2026","status":"published","topics":["cs.CV"],"abstract":"Slice-based volumetric imaging is widely applied and it demands representations that compress aggressively while preserving internal structure for analysis. We introduce GaussianPile, unifying 3D Gaussian splatting with an imaging system-aware focus model to address this challenge. Our proposed method introduces three key innovations: (i) a slice-aware piling strategy that positions anisotropic 3D Gaussians to model through-slice contributions, (ii) a differentiable projection operator that encodes the finite-thickness point spread function of the imaging acquisition system, and (iii) a compact encoding and joint optimization pipeline that simultaneously reconstructs and compresses the Gaussian sets. Our CUDA-based design retains the compression and real-time rendering efficiency of Gaussian primitives while preserving high-frequency internal volumetric detail. Experiments on microscopy and ultrasound datasets demonstrate that our method reduces storage and reconstruction cost, sustains diagnostic fidelity, and enables fast 2D visualization, along with 3D voxelization. In practice, it delivers high-quality results in as few as 3 minutes, up to 11x faster than NeRF-based approaches, and achieves consistent 16x compression over voxel grids, offering a practical path to deployable compression and exploration of slice-based volumetric datasets.","identifiers":{"arxiv":"2603.20611","doi":"10.48550/arXiv.2603.20611"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.20611"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.20611"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_o1swjxxaebea2eanmwv6830gh976rgp9"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.20611v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_o1swjxxaebea2eanmwv6830gh976rgp9"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2603-20611","updatedAt":"2026-08-08T11:59:38Z"},{"type":"conference","title":"Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA","authors":[{"name":"Tutian Tang"},{"name":"Xingyu Ji"},{"name":"Wanli Xing"},{"name":"Ce Hao"},{"name":"Wenqiang Xu"},{"name":"Lin Shao"},{"name":"Cewu Lu"},{"name":"Qiaojun Yu"},{"name":"Jiangmiao Pang"},{"name":"Kaifeng Zhang"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面","Ce Hao Affiliation: Zhongguancun Academy"],"relationType":"official-output","publishedAt":"2026-03-09","year":2026,"venue":"IROS 2026","status":"published","topics":["cs.RO"],"abstract":"While Vision-Language-Action (VLA) models have demonstrated remarkable success in robotic manipulation, their application has largely been confined to low-degree-of-freedom end-effectors performing simple, vision-guided pick-and-place tasks. Extending these models to human-like, bimanual dexterous manipulation-specifically contact-rich in-hand operations-introduces critical challenges in high-fidelity data acquisition, multi-skill learning, and multimodal sensory fusion. In this paper, we propose an integrated framework to address these bottlenecks, built upon two components. First, we introduce IMCopilot (In-hand Manipulation Copilot), a suite of reinforcement learning-trained atomic skills that plays a dual role: it acts as a shared-autonomy assistant to simplify teleoperation data collection, and it serves as a callable low-level execution primitive for the VLA. Second, we present MoDE-VLA (Mixture-of-Dexterous-Experts VLA), an architecture that seamlessly integrates heterogeneous force and tactile modalities into a pretrained VLA backbone. By utilizing a residual injection mechanism, MoDE-VLA enables contact-aware refinement without degrading the model's pretrained knowledge. We validate our approach on four tasks of escalating complexity, demonstrating doubled success rate improvement over the baseline in dexterous contact-rich tasks.","identifiers":{"arxiv":"2603.08122","doi":"10.48550/arXiv.2603.08122"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.08122"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.08122"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_zcnezi3d5262bwox2y8rzs3etx81racl"},{"label":"HTML","url":"https://arxiv.org/html/2603.08122v1"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.08122v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_zcnezi3d5262bwox2y8rzs3etx81racl"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Ce Hao Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.08122v1"}],"sources":["arXiv","北京中关村学院官网","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-08122","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"D3LM: A Discrete DNA Diffusion Language Model for Bidirectional DNA Understanding and Generation","authors":[{"name":"Zhao Yang"},{"name":"Hengchang Liu"},{"name":"Chuan Cao"},{"name":"Bing Su"}],"institutions":["zgca"],"rawAffiliations":["Chuan Cao Affiliation: Zhongguancun Academy Email: chuancao@bza.edu.cn"],"relationType":"affiliation","publishedAt":"2026-03-02","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Early DNA foundation models adopted BERT-style training, achieving good performance on DNA understanding tasks but lacking generative capabilities. Recent autoregressive models enable DNA generation, but employ left-to-right causal modeling that is suboptimal for DNA where regulatory relationships are inherently bidirectional. We present D3LM (\\textbf{D}iscrete \\textbf{D}NA \\textbf{D}iffusion \\textbf{L}anguage \\textbf{M}odel), which unifies bidirectional representation learning and DNA generation through masked diffusion. D3LM directly adopts the Nucleotide Transformer (NT) v2 architecture but reformulates the training objective as masked diffusion in discrete DNA space, enabling both bidirectional understanding and generation capabilities within a single model. Compared to NT v2 of the same size, D3LM achieves improved performance on understanding tasks. Notably, on regulatory element generation, D3LM achieves an SFID of 10.92, closely approaching real DNA sequences (7.85) and substantially outperforming the previous best result of 29.16 from autoregressive models. Our work suggests diffusion language models as a promising paradigm for unified DNA foundation models. We further present the first systematic study of masked diffusion models in the DNA domain, investigating practical design choices such as tokenization schemes and sampling strategies, thereby providing empirical insights and a solid foundation for future research. D3LM has been released at https://huggingface.co/collections/Hengchang-Liu/d3lm.","identifiers":{"arxiv":"2603.01780","doi":"10.48550/arXiv.2603.01780"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.01780"},{"label":"HTML","url":"https://arxiv.org/html/2603.01780v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.01780"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.01780v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chuan Cao Affiliation: Zhongguancun Academy Email: chuancao@bza.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.01780v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-01780","updatedAt":"2026-08-13T10:41:06Z"},{"type":"preprint","title":"RiskProp: Collision-Anchored Self-Supervised Risk Propagation for Early Accident Anticipation","authors":[{"name":"Yiyang Zou"},{"name":"Tianhao Zhao"},{"name":"Peilun Xiao"},{"name":"Hongyu Jin"},{"name":"Longyu Qi"},{"name":"Yuxuan Li"},{"name":"Liyin Liang"},{"name":"Yifeng Qian"},{"name":"Chunbo Lai"},{"name":"Yutian Lin"},{"name":"Zhihui Li"},{"name":"Yu Wu"}],"institutions":["zgca"],"rawAffiliations":["Tianhao Zhao 1 1 footnotemark: 1 Affiliation: School of Computer Science, Wuhan University Affiliation: Zhongguancun Academy, Beijing, China Email: happytianhao@whu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"Accident anticipation aims to predict impending collisions from dashcam videos and trigger early alerts. Existing methods rely on binary supervision with manually annotated \"anomaly onset\" frames, which are subjective and inconsistent, leading to inaccurate risk estimation. In contrast, we propose RiskProp, a novel collision-anchored self-supervised risk propagation paradigm for early accident anticipation, which removes the need for anomaly onset annotations and leverages only the reliably annotated collision frame. RiskProp models temporal risk evolution through two observation-driven losses: first, since future frames contain more definitive evidence of an impending accident, we introduce a future-frame regularization loss that uses the model's next-frame prediction as a soft target to supervise the current frame, enabling backward propagation of risk signals; second, inspired by the empirical trend of rising risk before accidents, we design an adaptive monotonic constraint to encourage a non-decreasing progression over time. Experiments on CAP and Nexar demonstrate that RiskProp achieves state-of-the-art performance and produces smoother, more discriminative risk curves, improving both early anticipation and interpretability.","identifiers":{"arxiv":"2603.27165","doi":"10.48550/arXiv.2603.27165"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.27165"},{"label":"HTML","url":"https://arxiv.org/html/2603.27165v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.27165"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.27165v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Tianhao Zhao 1 1 footnotemark: 1 Affiliation: School of Computer Science, Wuhan University Affiliation: Zhongguancun Academy, Beijing, China Email: happytianhao@whu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.27165v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-27165","updatedAt":"2026-09-05T05:30:31Z"},{"type":"preprint","title":"Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions","authors":[{"name":"Shiqin Wang"},{"name":"Haoyang Chen"},{"name":"Huaizhou Huang"},{"name":"Yinkan He"},{"name":"Dongfang Sun"},{"name":"Xiaoqing Chen"},{"name":"Xingyu Liu"},{"name":"Zheng Wang"},{"name":"Kaiyan Zhao"}],"institutions":["zgca"],"rawAffiliations":["Shiqin Wang † † thanks: These authors contributed equally to this work. Affiliation: National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Affiliation: Zhongguancun Academy, Beijing, China. 100094 School of Computer Science, Wuhan University, Wuhan, China"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"The learning order of semantic classes significantly impacts unsupervised domain adaptation for semantic segmentation, especially under adverse weather conditions. Most existing curricula rely on handcrafted heuristics (e.g., fixed uncertainty metrics) and follow a static schedule, which fails to adapt to a model's evolving, high-dimensional training dynamics, leading to category bias. Inspired by Reinforcement Learning, we cast curriculum learning as a sequential decision problem and propose an autonomous class scheduler. This scheduler consists of two components: (i) a high-dimensional state encoder that maps the model's training status into a latent space and distills key features indicative of progress, and (ii) a category-fair policy-gradient objective that ensures balanced improvement across classes. Coupled with mixed source-target supervision, the learned class rankings direct the network's focus to the most informative classes at each stage, enabling more adaptive and dynamic learning. It is worth noting that our method achieves state-of-the-art performance on three widely used benchmarks (e.g., ACDC, Dark Zurich, and Nighttime Driving) and shows generalization ability in synthetic-to-real semantic segmentation.","identifiers":{"arxiv":"2603.24322","doi":"10.48550/arXiv.2603.24322"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.24322"},{"label":"HTML","url":"https://arxiv.org/html/2603.24322v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.24322"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.24322v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Shiqin Wang † † thanks: These authors contributed equally to this work. Affiliation: National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Affiliation: Zhongguancun Academy, Beijing, China. 100094 School of Computer Science, Wuhan University, Wuhan, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.24322v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-24322","updatedAt":"2026-09-05T05:30:31Z"},{"type":"preprint","title":"PEARL: Personalized Streaming Video Understanding Model","authors":[{"name":"Yuanhong Zheng"},{"name":"Ruichuan An"},{"name":"Xiaopeng Lin"},{"name":"Yuxing Liu"},{"name":"Sihan Yang"},{"name":"Huanyu Zhang"},{"name":"Haodong Li"},{"name":"Qintong Zhang"},{"name":"Renrui Zhang"},{"name":"Guopeng Li"},{"name":"Yifan Zhang"},{"name":"Yuheng Li"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Wentao Zhang Affiliation: Peking University Adobe CASIA Stepfun CUHK Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.AI","cs.IR"],"abstract":"Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real-world feedback, limiting their ability to provide the real-time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL-Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to personalized concepts at exact timestamps under two modes: (1) Frame-level, focusing on a specific person or object in discrete frames, and (2) a novel Video-level, focusing on personalized actions unfolding across continuous frames. PEARL-Bench comprises 132 unique videos and 2,173 fine-grained annotations with precise timestamps. Concept diversity and annotation quality are strictly ensured through a combined pipeline of automated generation and human verification. To tackle this challenging new setting, we further propose PEARL, a plug-and-play, training-free strategy that serves as a strong baseline. Extensive evaluations across 8 offline and online models demonstrate that PEARL achieves state-of-the-art performance. Notably, it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be a highly effective and robust strategy. We hope this work advances vision-language model (VLM) personalization and inspires further research into streaming personalized AI assistants. Code is available at https://github.com/Yuanhong-Zheng/PEARL.","identifiers":{"arxiv":"2603.20422","doi":"10.48550/arXiv.2603.20422"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.20422"},{"label":"HTML","url":"https://arxiv.org/html/2603.20422v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.20422"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.20422v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Wentao Zhang Affiliation: Peking University Adobe CASIA Stepfun CUHK Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.20422v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-20422","updatedAt":"2026-09-04T05:36:47Z"},{"type":"preprint","title":"Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization","authors":[{"name":"Zequn Liu"},{"name":"Kehan Wu"},{"name":"Shufang Xie"},{"name":"Zekun Guo"},{"name":"Wei Zhang"},{"name":"Tao Qin"},{"name":"Renhe Liu"},{"name":"Yingce Xia"}],"institutions":["zgca"],"rawAffiliations":["Shufang Xie Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["q-bio.BM","cs.AI","cs.LG"],"abstract":"Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are rarely documented at scale. To bridge this gap, we propose DESRO, a framework for deciphering scientific reasoning from outcomes. By analyzing shared patterns and key differences within grouped data, a large language model (LLM) can recover the underlying logic. We instantiate this framework in molecule optimization, a pivotal stage in drug discovery that traditionally relies on the iterative reasoning of medicinal chemists. Across 2.3 million molecular property records, our framework infers optimization rationales by grouping molecules with shared fragments, then using an LLM to analyze how structural variations correlate with property differences. Based on the derived data, we train a model that conducts molecule optimization through an interpretable reasoning process. DESRO achieves the highest success rates on 15 out of 18 tasks, spanning both single- and multi-property optimization of bioactivity and ADMET properties. The reasoning process enables robust generalization to out-of-distribution scenarios, including novel property combinations, unseen biological targets, and unseen properties defined solely by natural language descriptions. In retrospective case studies under strict temporal splits, the model autonomously reconstructs expert-level lead optimization trajectories. Additionally, our framework extends beyond molecule optimization to reaction ligand selection. Our results establish deciphering reasoning steps from outcome data as a viable paradigm for enabling scientific reasoning, providing a scalable approach to accelerate scientific discovery.","identifiers":{"arxiv":"2603.20262","doi":"10.48550/arXiv.2603.20262"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.20262"},{"label":"HTML","url":"https://arxiv.org/html/2603.20262v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.20262"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.20262v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Shufang Xie Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.20262v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-20262","updatedAt":"2026-09-04T05:36:47Z"},{"type":"preprint","title":"Semantic Audio-Visual Navigation in Continuous Environments","authors":[{"name":"Yichen Zeng"},{"name":"Hebaixu Wang"},{"name":"Meng Liu"},{"name":"Yu Zhou"},{"name":"Chen Gao"},{"name":"Kehan Chen"},{"name":"Gongping Huang"}],"institutions":["zgca"],"rawAffiliations":["Yichen Zeng Hebaixu Wang Meng Liu Yu Zhou Chen Gao Kehan Chen Gongping Huang Affiliation: Wuhan University Affiliation: Zhongguancun Academy Affiliation: Shandong Jianzhu University Affiliation: Nankai University Affiliation: Tsinghua University Affiliation: CASIA Affiliation: UCAS"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.SD"],"abstract":"Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and leading to spatially discontinuous observations. To establish a more realistic setting, we introduce Semantic Audio-Visual Navigation in Continuous Environments (SAVN-CE), where agents can move freely in 3D spaces and perceive temporally and spatially coherent audio-visual streams. In this setting, targets may intermittently become silent or stop emitting sound entirely, causing agents to lose goal information. To tackle this challenge, we propose MAGNet, a multimodal transformer-based model that jointly encodes spatial and semantic goal representations and integrates historical context with self-motion cues to enable memory-augmented goal reasoning. Comprehensive experiments demonstrate that MAGNet significantly outperforms state-of-the-art methods, achieving up to a 12.1\\% absolute improvement in success rate. These results also highlight its robustness to short-duration sounds and long-distance navigation scenarios. The code is available at https://github.com/yichenzeng24/SAVN-CE.","identifiers":{"arxiv":"2603.19660","doi":"10.48550/arXiv.2603.19660"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.19660"},{"label":"HTML","url":"https://arxiv.org/html/2603.19660v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.19660"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.19660v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yichen Zeng Hebaixu Wang Meng Liu Yu Zhou Chen Gao Kehan Chen Gongping Huang Affiliation: Wuhan University Affiliation: Zhongguancun Academy Affiliation: Shandong Jianzhu University Affiliation: Nankai University Affiliation: Tsinghua University Affiliation: CASIA Affiliation: UCAS","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.19660v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-19660","updatedAt":"2026-09-04T05:36:47Z"},{"type":"preprint","title":"Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering","authors":[{"name":"Jiayi Luo"},{"name":"Jiayu Chen"},{"name":"Jiankun Wang"},{"name":"Cong Wang"},{"name":"Hanxin Zhu"},{"name":"Qingyun Sun"},{"name":"Chen Gao"},{"name":"Zhibo Chen"},{"name":"Jianxin Li"}],"institutions":["zgca"],"rawAffiliations":["Jiayi Luo Affiliation: SKLCCSE, School of Computer Science and Engineering, Beihang University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, motivating sparse attention techniques for improving efficiency. However, existing training-free sparse attention methods for video generation still face two unresolved limitations: ignoring layer heterogeneity in attention pruning and ignoring query-key coupling in block partitioning, which hinder a better quality-speedup trade-off. In this work, we uncover a critical insight: attention sparsity is an intrinsic layer-wise property, with only minor variation across different inputs. Motivated by this observation, we propose SVOO, a training-free sparse attention framework for fast video generation via offline layer-wise sparsity profiling and online bidirectional co-clustering. Specifically, SVOO adopts a two-stage paradigm: (i) offline layer-wise sensitivity profiling to derive intrinsic per-layer pruning levels, and (ii) online block-wise sparse attention via a bidirectional co-clustering algorithm. Extensive experiments on seven widely used video generation models demonstrate that SVOO achieves a superior quality-speedup trade-off over state-of-the-art methods, delivering up to 1.93x speedup while maintaining a PSNR of up to 29 dB on Wan2.1. Code is available at: https://github.com/Mutual-Luo/SVOO.","identifiers":{"arxiv":"2603.18636","doi":"10.48550/arXiv.2603.18636"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.18636"},{"label":"HTML","url":"https://arxiv.org/html/2603.18636v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.18636"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.18636v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jiayi Luo Affiliation: SKLCCSE, School of Computer Science and Engineering, Beihang University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.18636v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-18636","updatedAt":"2026-09-04T05:36:47Z"},{"type":"preprint","title":"ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations","authors":[{"name":"Hankun Kang"},{"name":"Xin Miao"},{"name":"Jianhao Chen"},{"name":"Jintao Wen"},{"name":"Mayi Xu"},{"name":"Weiyu Zhang"},{"name":"Wenpeng Lu"},{"name":"Tieyun Qian"}],"institutions":["zgca"],"rawAffiliations":["Jianhao Chen Affiliation: School of Computer Science , Wuhan University , Wuhan , Hubei , China Affiliation: Zhongguancun Academy , Beijing , China email: chenjianhao@whu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL","cs.AI"],"abstract":"Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop evasive perturbations to disguise toxic content and evade detectors. Traditional detectors or methods are static over time and are inadequate in addressing these evolving evasion tactics. Thus, continual learning emerges as a logical approach to dynamically update detection ability against evolving perturbations. Nevertheless, disparities across perturbations hinder the detector's continual learning on perturbed text. More importantly, perturbation-induced noises distort semantics to degrade comprehension and also impair critical feature learning to render detection sensitive to perturbations. These amplify the challenge of continual learning against evolving perturbations. In this work, we present ContiGuard, the first framework tailored for continual learning of the detector on time-evolving perturbed text (termed continual toxicity detection) to enable the detector to continually update capability and maintain sustained resilience against evolving perturbations. Specifically, to boost the comprehension, we present an LLM-powered semantic enriching strategy, where we dynamically incorporate possible meaning and toxicity-related clues excavated by LLM into the perturbed text to improve the comprehension. To mitigate non-critical features and amplify critical ones, we propose a discriminability-driven feature learning strategy, where we strengthen discriminative features while suppressing the less-discriminative ones to shape a robust classification boundary for detection...","identifiers":{"arxiv":"2603.14843","doi":"10.48550/arXiv.2603.14843"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.14843"},{"label":"HTML","url":"https://arxiv.org/html/2603.14843v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.14843"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.14843v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jianhao Chen Affiliation: School of Computer Science , Wuhan University , Wuhan , Hubei , China Affiliation: Zhongguancun Academy , Beijing , China email: chenjianhao@whu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.14843v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-14843","updatedAt":"2026-09-03T05:34:18Z"},{"type":"preprint","title":"HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System","authors":[{"name":"Kailin Lyu"},{"name":"Kangyi Wu"},{"name":"Pengna Li"},{"name":"Xiuyu Hu"},{"name":"Qingyi Si"},{"name":"Cui Miao"},{"name":"Ning Yang"},{"name":"Zihang Wang"},{"name":"Long Xiao"},{"name":"Lianyu Hu"},{"name":"Jingyuan Sun"},{"name":"Ce Hao"}],"institutions":["zgca"],"rawAffiliations":["Kailin Lyu Affiliation: Kailin Lyu is with the Institute of Automation, Chinese Academy of Sciences, and the Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.RO"],"abstract":"LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs as navigators, which face challenges related to high token costs and potential data leakage risks. Recent efforts have attempted to address this by using open-source LLMs combined with a spatiotemporal CoT framework, but they still fall far short compared to closed-source models. In this work, we identify a critical issue, Navigation Amnesia, through a detailed analysis of the navigation process. This issue leads to navigation failures and amplifies the gap between open-source and closed-source methods. To address this, we propose HiMemVLN, which incorporates a Hierarchical Memory System into a multimodal large model to enhance visual perception recall and long-term localization, mitigating the amnesia issue and improving the agent's navigation performance. Extensive experiments in both simulated and real-world environments demonstrate that HiMemVLN achieves nearly twice the performance of the open-source state-of-the-art method. The code is available at https://github.com/lvkailin0118/HiMemVLN.","identifiers":{"arxiv":"2603.14807","doi":"10.48550/arXiv.2603.14807"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.14807"},{"label":"HTML","url":"https://arxiv.org/html/2603.14807v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.14807"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.14807v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Kailin Lyu Affiliation: Kailin Lyu is with the Institute of Automation, Chinese Academy of Sciences, and the Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.14807v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-14807","updatedAt":"2026-09-03T05:34:18Z"},{"type":"preprint","title":"LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos","authors":[{"name":"Rongyi Yu"},{"name":"Chenyuan Duan"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Wentao Zhang Affiliation: Peking University & Zhongguancun Academy , Beijing , China email: wentao.zhang@pku.edu.cn"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.IR"],"abstract":"Long video question answering (Long-Video QA) increasingly relies on agentic tool use to retrieve evidence from long videos. In realistic settings, this process often requires multi-hop retrieval, where agents must iteratively gather multiple discontinuous evidence clips. However, existing long-video benchmarks are largely static: they rarely enforce strict multi-hop retrieval and typically lack a standardized evidence-access interface, making it difficult to separate failures in retrieval planning from those in answer generation. To address this gap, we introduce LongVidSearch, a benchmark for evaluating agentic multi-hop evidence retrieval planning in long videos under standardized access constraints. LongVidSearch enforces retrieval necessity: a Hop-k question requires exactly k necessary evidence clips, and removing any single clip renders the question unsolvable. The benchmark contains 3,000 questions over 447 long videos (average length 26 minutes), covering four reasoning categories: State Mutation, Causal Inference, Global Summary, and Visual Tracking, with 2-hop, 3-hop, and 4-hop evidence requirements. To ensure fair and controlled evaluation, all agents interact with LongVidSearch through a unified tool interface, which fixes the retrieval backend and isolates the agent's ability to formulate queries and plan iterative retrieval. In addition to answer accuracy, we measure tool-call cost to analyze the accuracy-efficiency trade-off under identical access conditions. We evaluate VideoAgent-style QA agents with multiple backbone LLMs using three-judge majority voting. GPT-5 achieves the highest accuracy (42.43), outperforming Gemini 3 Pro (30.97) and GPT-4o (19.20), yet remaining below 50 %, highlighting the difficulty of multi-hop retrieval planning. With gold evidence clips, performance becomes near-perfect, confirming retrieval planning as the primary bottleneck.","identifiers":{"arxiv":"2603.14468","doi":"10.48550/arXiv.2603.14468"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.14468"},{"label":"HTML","url":"https://arxiv.org/html/2603.14468v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.14468"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.14468v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Wentao Zhang Affiliation: Peking University & Zhongguancun Academy , Beijing , China email: wentao.zhang@pku.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.14468v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-14468","updatedAt":"2026-09-03T05:34:18Z"},{"type":"preprint","title":"One-Eval: An Agentic System for Automated and Traceable LLM Evaluation","authors":[{"name":"Chengyu Shen"},{"name":"Yanheng Hou"},{"name":"Minghui Pan"},{"name":"Runming He"},{"name":"Zhen Hao Wong"},{"name":"Meiyi Qiang"},{"name":"Zhou Liu"},{"name":"Hao Liang"},{"name":"Peichao Lai"},{"name":"Zeang Sheng"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Hao Liang Affiliation: Peking University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL"],"abstract":"Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify appropriate benchmarks, reproduce heterogeneous evaluation codebases, configure dataset schema mappings, and interpret aggregated metrics. To address these challenges, we present One-Eval, an agentic evaluation system that converts natural-language evaluation requests into executable, traceable, and customizable evaluation workflows. One-Eval integrates (i) NL2Bench for intent structuring and personalized benchmark planning, (ii) BenchResolve for benchmark resolution, automatic dataset acquisition, and schema normalization to ensure executability, and (iii) Metrics \\& Reporting for task-aware metric selection and decision-oriented reporting beyond scalar scores. The system further incorporates human-in-the-loop checkpoints for review, editing, and rollback, while preserving sample evidence trails for debugging and auditability. Experiments show that One-Eval can execute end-to-end evaluations from diverse natural-language requests with minimal user effort, supporting more efficient and reproducible evaluation in industrial settings. Our framework is publicly available at https://github.com/OpenDCAI/One-Eval.","identifiers":{"arxiv":"2603.09821","doi":"10.48550/arXiv.2603.09821"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.09821"},{"label":"HTML","url":"https://arxiv.org/html/2603.09821v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.09821"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.09821v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Hao Liang Affiliation: Peking University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.09821v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-09821","updatedAt":"2026-09-02T05:36:27Z"},{"type":"preprint","title":"Memory-Guided View Refinement for Dynamic Human-in-the-loop EQA","authors":[{"name":"Xin Lu"},{"name":"Rui Li"},{"name":"Xun Huang"},{"name":"Weixin Li"},{"name":"Chuanqing Zhuang"},{"name":"Jiayuan Li"},{"name":"Zhengda Lu"},{"name":"Jun Xiao"},{"name":"Yunhong Wang"}],"institutions":["zgca"],"rawAffiliations":["Xin Lu Affiliation: University of Chinese Academy of Sciences, China Affiliation: Zhongguancun Academy, China"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV","cs.MM"],"abstract":"Embodied Question Answering (EQA) has traditionally been evaluated in temporally stable environments where visual evidence can be accumulated reliably. However, in dynamic, human-populated scenes, human activities and occlusions introduce significant perceptual non-stationarity: task-relevant cues are transient and view-dependent, while a store-then-retrieve strategy over-accumulates redundant evidence and increases inference cost. This setting exposes two practical challenges for EQA agents: resolving ambiguity caused by viewpoint-dependent occlusions, and maintaining compact yet up-to-date evidence for efficient inference. To enable systematic study of this setting, we introduce DynHiL-EQA, a human-in-the-loop EQA dataset with two subsets: a Dynamic subset featuring human activities and temporal changes, and a Static subset with temporally stable observations. To address the above challenges, we present DIVRR (Dynamic-Informed View Refinement and Relevance-guided Adaptive Memory Selection), a training-free framework that couples relevance-guided view refinement with selective memory admission. By verifying ambiguous observations before committing them and retaining only informative evidence, DIVRR improves robustness under occlusions while preserving fast inference with compact memory. Extensive experiments on DynHiL-EQA and the established HM-EQA dataset demonstrate that DIVRR consistently improves over existing baselines in both dynamic and static settings while maintaining high inference efficiency.","identifiers":{"arxiv":"2603.09541","doi":"10.48550/arXiv.2603.09541"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.09541"},{"label":"HTML","url":"https://arxiv.org/html/2603.09541v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.09541"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.09541v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Xin Lu Affiliation: University of Chinese Academy of Sciences, China Affiliation: Zhongguancun Academy, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.09541v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-09541","updatedAt":"2026-09-02T05:36:27Z"},{"type":"preprint","title":"AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models","authors":[{"name":"Hankun Kang"},{"name":"Di Lin"},{"name":"Zhirong Liao"},{"name":"Pengfei Bai"},{"name":"Xinyi Zeng"},{"name":"Jiawei Jiang"},{"name":"Yuanyuan Zhu"},{"name":"Tieyun Qian"}],"institutions":["zgca"],"rawAffiliations":["Tieyun Qian Affiliation: Wuhan University , Wuhan , Hubei , China Affiliation: Zhongguancun Academy , Beijing , China email: qty@whu.edu.cn"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL","cs.AI"],"abstract":"With the widespread adoption of Large Language Models (LLMs), respecting indigenous cultures becomes essential for models' culturally safety and responsible global applications. Existing studies separately consider cultural safety and cultural knowledge and neglect that the former should be grounded by the latter. This severely prevents LLMs from yielding culture-specific respectful responses. Consequently, adaptive cultural safety remains a formidable task. In this work, we propose to jointly model cultural safety and knowledge. First and foremost, cultural-safety and knowledge-paired data serve as the key prerequisite to conduct this research. However, the cultural diversity across regions and the subtlety of cultural differences pose significant challenges to the creation of such paired evaluation data. To address this issue, we propose a novel framework that integrates authoritative cultural knowledge descriptions curation, LLM-automated query generation, and heavy manual verification. Accordingly, we obtain a dataset named AdaCultureSafe containing 4.8K manually decomposed fine-grained cultural descriptions and the corresponding 48K manually verified safety- and knowledge-oriented queries. Upon the constructed dataset, we evaluate three families of popular LLMs on their cultural safety and knowledge proficiency, via which we make a critical discovery: no significant correlation exists between their cultural safety and knowledge proficiency. We then delve into the utility-related neuron activations within LLMs to investigate the potential cause of the absence of correlation, which can be attributed to the difference of the objectives of pre-training and post-alignment. We finally present a knowledge-grounded method, which significantly enhances cultural safety by enforcing the integration of knowledge into the LLM response generation process.","identifiers":{"arxiv":"2603.08275","doi":"10.48550/arXiv.2603.08275"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.08275"},{"label":"HTML","url":"https://arxiv.org/html/2603.08275v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.08275"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.08275v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Tieyun Qian Affiliation: Wuhan University , Wuhan , Hubei , China Affiliation: Zhongguancun Academy , Beijing , China email: qty@whu.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.08275v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-08275","updatedAt":"2026-09-02T05:36:27Z"},{"type":"preprint","title":"Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization","authors":[{"name":"Haodong Zhu"},{"name":"Yangyang Ren"},{"name":"Yanjing Li"},{"name":"Mingbao Lin"},{"name":"Linlin Yang"},{"name":"Xuhui Liu"},{"name":"Xiantong Zhen"},{"name":"Haiguang Liu"},{"name":"Baochang Zhang"}],"institutions":["zgca"],"rawAffiliations":["Haodong Zhu Affiliation: Beihang University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI"],"abstract":"Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the underlying sampling distribution, compromising theoretical rigor and convergence behavior. To address this limitation, we propose Dynamic Pruning Policy Optimization (DPPO), a framework that enables dynamic pruning while preserving unbiased gradient estimation through importance sampling-based correction. By incorporating mathematically derived rescaling factors, DPPO significantly accelerates GRPO training without altering the optimization objective of the full-batch baseline. Furthermore, to mitigate the data sparsity induced by pruning, we introduce Dense Prompt Packing, a window-based greedy strategy that maximizes valid token density and hardware utilization. Extensive experiments demonstrate that DPPO consistently accelerates training across diverse models and benchmarks. For instance, on Qwen3-4B trained on MATH, DPPO achieves 2.37$\\times$ training speedup and outperforms GRPO by 3.36% in average accuracy across six mathematical reasoning benchmarks.","identifiers":{"arxiv":"2603.04135","doi":"10.48550/arXiv.2603.04135"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.04135"},{"label":"HTML","url":"https://arxiv.org/html/2603.04135v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.04135"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.04135v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haodong Zhu Affiliation: Beihang University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.04135v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-04135","updatedAt":"2026-09-01T06:09:06Z"},{"type":"preprint","title":"Any2Any: Unified Arbitrary Modality Translation for Remote Sensing","authors":[{"name":"Haoyang Chen"},{"name":"Jing Zhang"},{"name":"Hebaixu Wang"},{"name":"Shiqin Wang"},{"name":"Pohsun Huang"},{"name":"Jiayuan Li"},{"name":"Haonan Guo"},{"name":"Di Wang"},{"name":"Zheng Wang"},{"name":"Bo Du"}],"institutions":["zgca"],"rawAffiliations":["Haoyang Chen Affiliation: Wuhan University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited generalization to unseen modality combinations. We formulate Any-to-Any translation as inference over a shared latent representation of the scene, where different modalities correspond to partial observations of the same underlying semantics. Based on this formulation, we propose Any2Any, a unified latent diffusion framework that projects heterogeneous inputs into a geometrically aligned latent space. Such structure performs anchored latent regression with a shared backbone, decoupling modality-specific representation learning from semantic mapping. Moreover, lightweight target-specific residual adapters are used to correct systematic latent mismatches without increasing inference complexity. To support learning under sparse but connected supervision, we introduce RST-1M, the first million-scale remote sensing dataset with paired observations across five sensing modalities, providing supervision anchors for any-to-any translation. Experiments across 14 translation tasks show that Any2Any consistently outperforms pairwise translation methods and exhibits strong zero-shot generalization to unseen modality pairs. Code and models are available at https://github.com/MiliLab/Any2Any.","identifiers":{"arxiv":"2603.04114","doi":"10.48550/arXiv.2603.04114"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.04114"},{"label":"HTML","url":"https://arxiv.org/html/2603.04114v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.04114"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.04114v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haoyang Chen Affiliation: Wuhan University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.04114v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-04114","updatedAt":"2026-09-01T06:09:06Z"},{"type":"preprint","title":"Swimming Under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion","authors":[{"name":"Xinyu Cui"},{"name":"Fei Han"},{"name":"Hang Xu"},{"name":"Yongcheng Zeng"},{"name":"Luoyang Sun"},{"name":"Ruizhi Zhang"},{"name":"Jian Zhao"},{"name":"Haifeng Zhang"},{"name":"Weikun Li"},{"name":"Hao Chen"},{"name":"Jun Wang"},{"name":"Dixia Fan"}],"institutions":["zgca"],"rawAffiliations":["Jian Zhao Affiliation: Zhongguancun Academy, Beijing, 100094, China"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.RO"],"abstract":"Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes forward thrust while minimizing destabilizing fluctuations. Our proposed framework, Accelerated Constrained Proximal Policy Optimization with a PID-regulated Lagrange multiplier (ACPPO-PID), enforces constraints with a PID-regulated Lagrange multiplier, accelerates learning via conditional asymmetric clipping, and stabilizes updates through cycle-wise geometric aggregation. Initialized with imitation learning and refined through on-hardware towing-tank experiments, ACPPO-PID produces control policies that transfer effectively to quadrupedal free-swimming trials. Results demonstrate improved thrust efficiency, reduced destabilizing forces, and faster convergence compared with state-of-the-art baselines, underscoring the importance of constraint-aware safe RL for robust and generalizable bio-inspired locomotion in complex fluid environments.","identifiers":{"arxiv":"2603.04073","doi":"10.48550/arXiv.2603.04073"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.04073"},{"label":"HTML","url":"https://arxiv.org/html/2603.04073v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.04073"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.04073v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jian Zhao Affiliation: Zhongguancun Academy, Beijing, 100094, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.04073v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-04073","updatedAt":"2026-09-01T06:09:06Z"},{"type":"preprint","title":"SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse","authors":[{"name":"Xuanran Zhai"},{"name":"Zekai Huang"},{"name":"Longyan Wu"},{"name":"Qianyou Zhao"},{"name":"Qiaojun Yu"},{"name":"Jieji Ren"},{"name":"Ce Hao"},{"name":"Harold Soh"}],"institutions":["zgca"],"rawAffiliations":["Ce Hao Affiliation: Beijing Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-03-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.RO"],"abstract":"Recent progress in vision-language-action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environments. However, mainstream bimanual VLA formulations largely overlook the critical challenge of combinatorial diversity. Different pairings of single-arm behaviors can induce qualitatively distinct task behaviors, yet existing models do not explicitly account for this structure. We argue that effective bimanual VLAs should support skill reuse - the ability to recombine previously learned single-arm skills across novel left-right pairings - thereby avoiding the need to separately learn every possible combination. Current VLA designs entangle skills across arms, preventing such recomposition and limiting scalability. To address this limitation, we propose SkillVLA, a framework explicitly designed to enable skill reuse in dual-arm manipulation. Extensive experiments demonstrate that SkillVLA substantially improves skill composition, increasing overall success rate from 0% to 51%, and achieves strong performance on cooperative and long-horizon tasks.","identifiers":{"arxiv":"2603.03836","doi":"10.48550/arXiv.2603.03836"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2603.03836"},{"label":"HTML","url":"https://arxiv.org/html/2603.03836v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2603.03836"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2603.03836v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Ce Hao Affiliation: Beijing Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2603.03836v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2603-03836","updatedAt":"2026-09-01T06:09:06Z"},{"id":"whitepaper-beijing-genai-2026","type":"whitepaper","title":"北京市生成式人工智能大模型产业发展与治理白皮书","authors":[{"name":"中国科学院信息工程研究所"},{"name":"北京中关村学院","institutions":["zgca"]},{"name":"中关村人工智能研究院","institutions":["zgci"]}],"institutions":["zgca","zgci"],"rawAffiliations":["中国科学院信息工程研究所联合北京中关村学院与中关村人工智能研究院共同编制"],"relationType":"official-output","publishedAt":"2026-02-28","year":2026,"venue":"第四届北京人工智能产业创新发展大会","status":"released","topics":["Generative AI","Large Models","AI Governance","Industry"],"abstract":"白皮书围绕北京市生成式人工智能大模型的发展态势、安全风险、测评方法、安全成熟度、产业发展与治理建议展开，并提出以安全成熟度指标体系促进产业规范发展。","identifiers":{},"links":[{"label":"Official","url":"https://www.iie.ac.cn/xwdt/kydt/202603/t20260303_8147538.html"}],"versions":[{"label":"Release announcement","url":"https://www.iie.ac.cn/xwdt/kydt/202603/t20260303_8147538.html"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院","source":"Institute of Information Engineering, CAS","sourceUrl":"https://www.iie.ac.cn/xwdt/kydt/202603/t20260303_8147538.html"},{"level":"official-listing","institution":"zgci","matchedText":"中关村人工智能研究院","source":"Institute of Information Engineering, CAS","sourceUrl":"https://www.iie.ac.cn/xwdt/kydt/202603/t20260303_8147538.html"}],"sources":["Official institutional release"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"conference","title":"Extending Sequence Length is Not All You Need: Effective Integration of Multimodal Signals for Gene Expression Prediction","authors":[{"name":"Zhao Yang"},{"name":"Yi Duan"},{"name":"Jiwei Zhu"},{"name":"Ying Ba"},{"name":"Chuan Cao"},{"name":"Bing Su"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-02-25","year":2026,"venue":"ICLR 2026","status":"published","topics":["cs.LG","q-bio.GN"],"abstract":"Gene expression prediction, which predicts mRNA expression levels from DNA sequences, presents significant challenges. Previous works often focus on extending input sequence length to locate distal enhancers, which may influence target genes from hundreds of kilobases away. Our work first reveals that for current models, long sequence modeling can decrease performance. Even carefully designed algorithms only mitigate the performance degradation caused by long sequences. Instead, we find that proximal multimodal epigenomic signals near target genes prove more essential. Hence we focus on how to better integrate these signals, which has been overlooked. We find that different signal types serve distinct biological roles, with some directly marking active regulatory elements while others reflect background chromatin patterns that may introduce confounding effects. Simple concatenation may lead models to develop spurious associations with these background patterns. To address this challenge, we propose Prism, a framework that learns multiple combinations of high-dimensional epigenomic features to represent distinct background chromatin states and uses backdoor adjustment to mitigate confounding effects. Our experimental results demonstrate that proper modeling of multimodal epigenomic signals achieves state-of-the-art performance using only short sequences for gene expression prediction.","identifiers":{"arxiv":"2602.21550","doi":"10.48550/arXiv.2602.21550"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.21550"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.21550"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_h5w6mf0s0gevg8ibfx0boc4vp99w5tc4"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.21550v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_h5w6mf0s0gevg8ibfx0boc4vp99w5tc4"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2602-21550","updatedAt":"2026-08-08T11:59:38Z"},{"type":"conference","title":"ChordEdit: One-Step Low-Energy Transport for Image Editing","authors":[{"name":"Liangsi Lu"},{"name":"Xuhang Chen"},{"name":"Minzhe Guo"},{"name":"Shichu Li"},{"name":"Jingchao Wang"},{"name":"Yang Shi"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-02-22","year":2026,"venue":"CVPR 2026","status":"published","topics":["cs.CV"],"abstract":"The advent of one-step text-to-image (T2I) models offers unprecedented synthesis speed. However, their application to text-guided image editing remains severely hampered, as forcing existing training-free editors into a single inference step fails. This failure manifests as severe object distortion and a critical loss of consistency in non-edited regions, resulting from the high-energy, erratic trajectories produced by naive vector arithmetic on the models' structured fields. To address this problem, we introduce ChordEdit, a model agnostic, training-free, and inversion-free method that facilitates high-fidelity one-step editing. We recast editing as a transport problem between the source and target distributions defined by the source and target text prompts. Leveraging dynamic optimal transport theory, we derive a principled, low-energy control strategy. This strategy yields a smoothed, variance-reduced editing field that is inherently stable, facilitating the field to be traversed in a single, large integration step. A theoretically grounded and experimentally validated approach allows ChordEdit to deliver fast, lightweight and precise edits, finally achieving true real-time editing on these challenging models.","identifiers":{"arxiv":"2602.19083","doi":"10.48550/arXiv.2602.19083"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.19083"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.19083"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_d0qhgfwk5u8c45svd8cmydi5eh526adr"},{"label":"Code","url":"https://github.com/tsinghua-fib-lab/AutoSOTA"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.19083v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_d0qhgfwk5u8c45svd8cmydi5eh526adr"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2602-19083","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model","authors":[{"name":"Jingwen Sun"},{"name":"Wenyao Zhang"},{"name":"Zekun Qi"},{"name":"Shaojie Ren"},{"name":"Zezhi Liu"},{"name":"Hanxin Zhu"},{"name":"Guangzhong Sun"},{"name":"Xin Jin"},{"name":"Zhibo Chen"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）"],"relationType":"official-output","publishedAt":"2026-02-10","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.RO","cs.CV"],"abstract":"Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods.","identifiers":{"arxiv":"2602.10098","doi":"10.48550/arXiv.2602.10098"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.10098"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.10098"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_l7phjpnrrbyvox1hdf4hw105uqmctm1v"},{"label":"Code","url":"https://github.com/ginwind/VLA-JEPA"},{"label":"Model","url":"https://huggingface.co/ginwind/VLA-JEPA"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.10098v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_l7phjpnrrbyvox1hdf4hw105uqmctm1v"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_l7phjpnrrbyvox1hdf4hw105uqmctm1v"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2602-10098","updatedAt":"2026-08-08T11:59:38Z"},{"type":"conference","title":"RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI","authors":[{"name":"Hongzhi Zang"},{"name":"Shu'ang Yu"},{"name":"Hao Lin"},{"name":"Tianxing Zhou"},{"name":"Zefang Huang"},{"name":"Zhen Guo"},{"name":"Xin Xu"},{"name":"Jiakai Zhou"},{"name":"Yuze Sheng"},{"name":"Shizhe Zhang"},{"name":"Feng Gao"},{"name":"Wenhao Tang"},{"name":"Yufeng Yue"},{"name":"Quanlu Zhang"},{"name":"Xinlei Chen"},{"name":"Chao Yu"},{"name":"Yu Wang"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2026-02-08","year":2026,"venue":"RSS 2026","status":"published","topics":["cs.RO"],"abstract":"Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, which makes scalable data collection, heterogeneous deployment, and long-horizon effective training difficult. These challenges suggest that real-world policy learning is not only an algorithmic issue but fundamentally a systems problem. We present USER, a Unified and extensible SystEm for Real-world online policy learning. USER treats physical robots as first-class hardware resources alongside GPUs through a unified hardware abstraction layer, enabling automatic discovery, management, and scheduling of heterogeneous robots. To address cloud-edge communication, USER introduces an adaptive communication plane with tunneling-based networking, distributed data channels for traffic localization, and streaming-multiprocessor-aware weight synchronization to regulate GPU-side overhead. On top of this infrastructure, USER organizes learning as a fully asynchronous framework with a persistent, cache-aware buffer, enabling efficient long-horizon experiments with robust crash recovery and reuse of historical data. In addition, USER provides extensible abstractions for rewards, algorithms, and policies, supporting online imitation or reinforcement learning of CNN/MLP, generative policies, and large vision-language-action (VLA) models within a unified pipeline. Results in both simulation and the real world show that USER enables multi-robot coordination, heterogeneous manipulators, edge-cloud collaboration with large models, and long-running asynchronous training, offering a unified and extensible systems foundation for real-world online policy learning.","identifiers":{"arxiv":"2602.07837","doi":"10.48550/arXiv.2602.07837"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.07837"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.07837"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_7o9560ekbye9ncxknflzf3nrwfhv18wd"},{"label":"Code","url":"https://github.com/RLinf/RLinf"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.07837v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_7o9560ekbye9ncxknflzf3nrwfhv18wd"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2602-07837","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Test-Time Computing for Referring Multimodal Large Language Models","authors":[{"name":"Mingrui Wu"},{"name":"Hao Chen"},{"name":"Jiayi Ji"},{"name":"Xiaoshuai Sun"},{"name":"Zhiyuan Liu"},{"name":"Liujuan Cao"},{"name":"Ming-Ming Cheng"},{"name":"Rongrong Ji"}],"institutions":["zgca"],"rawAffiliations":["Mingrui Wu Affiliation: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China Affiliation: Zhongguancun Academy, Beijing, China.100094"],"relationType":"affiliation","publishedAt":"2026-02-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"We propose ControlMLLM++, a novel test-time adaptation framework that injects learnable visual prompts into frozen multimodal large language models (MLLMs) to enable fine-grained region-based visual reasoning without any model retraining or fine-tuning. Leveraging the insight that cross-modal attention maps intrinsically encode semantic correspondences between textual tokens and visual regions, ControlMLLM++ optimizes a latent visual token modifier during inference via a task-specific energy function to steer model attention towards user-specified areas. To enhance optimization stability and mitigate language prompt biases, ControlMLLM++ incorporates an improved optimization strategy (Optim++) and a prompt debiasing mechanism (PromptDebias). Supporting diverse visual prompt types including bounding boxes, masks, scribbles, and points, our method demonstrates strong out-of-domain generalization and interpretability. The code is available at https://github.com/mrwu-mac/ControlMLLM.","identifiers":{"arxiv":"2602.19505","doi":"10.48550/arXiv.2602.19505"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.19505"},{"label":"HTML","url":"https://arxiv.org/html/2602.19505v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.19505"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.19505v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Mingrui Wu Affiliation: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China Affiliation: Zhongguancun Academy, Beijing, China.100094","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2602.19505v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2602-19505","updatedAt":"2026-08-31T06:36:19Z"},{"type":"preprint","title":"AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG","authors":[{"name":"Qijie You"},{"name":"Wenkai Yu"},{"name":"Wentao Zhang"}],"institutions":["zgca"],"rawAffiliations":["Wentao Zhang Affiliation: Peking University Affiliation: Zhongguancun Academy Affiliation: Beijing Key Laboratory of Data Intelligence and Security (Peking University) u202342615@xs.ustb.edu.cn, { ywk, wentao.zhang } @pku.edu.cn"],"relationType":"affiliation","publishedAt":"2026-02-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL"],"abstract":"With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop reasoning, which requires models to engage in deliberate thinking and multi-step interaction, serves as a critical testbed for assessing such capabilities. However, existing benchmarks typically provide only final questions and answers, while lacking the intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query. This limitation prevents researchers from analyzing at which step an agent fails and restricts more fine-grained evaluation of model capabilities. Moreover, most current benchmarks are manually constructed, which is both time-consuming and labor-intensive, while also limiting scalability and generalization. To address these challenges, we introduce AgenticRAGTracer, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation. Our benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks. Extensive experiments demonstrate that even the best large language models perform poorly on our dataset. For instance, GPT-5 attains merely 22.6\\% EM accuracy on the hardest portion of our dataset. Hop-aware diagnosis reveals that failures are primarily driven by distorted reasoning chains -- either collapsing prematurely or wandering into over-extension. This highlights a critical inability to allocate steps consistent with the task's logical structure, providing a diagnostic dimension missing in traditional evaluations. We believe our work will facilitate research in Agentic RAG and inspire further meaningful progress in this area. Our code and data are available at https://github.com/YqjMartin/AgenticRAGTracer.","identifiers":{"arxiv":"2602.19127","doi":"10.48550/arXiv.2602.19127"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.19127"},{"label":"HTML","url":"https://arxiv.org/html/2602.19127v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.19127"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.19127v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Wentao Zhang Affiliation: Peking University Affiliation: Zhongguancun Academy Affiliation: Beijing Key Laboratory of Data Intelligence and Security (Peking University) u202342615@xs.ustb.edu.cn, { ywk, wentao.zhang } @pku.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2602.19127v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2602-19127","updatedAt":"2026-08-30T06:21:20Z"},{"type":"preprint","title":"Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models","authors":[{"name":"Liangzhi Shi"},{"name":"Shuaihang Chen"},{"name":"Feng Gao"},{"name":"Yinuo Chen"},{"name":"Kang Chen"},{"name":"Tonghe Zhang"},{"name":"Hongzhi Zang"},{"name":"Jiakai Zhou"},{"name":"Weinan Zhang"},{"name":"Chao Yu"},{"name":"Yu Wang"}],"institutions":["zgca"],"rawAffiliations":["Shuaihang Chen Affiliation: Tsinghua University Affiliation: Harbin Institute of Technology Affiliation: Zhongguancun Academy Equal contribution. Project Leader. Corresponding Authors: yuchao@sz.tsinghua.edu.cn , yu-wang@mail.tsinghua.edu.cn"],"relationType":"affiliation","publishedAt":"2026-02-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.RO"],"abstract":"Simulation offers a scalable and low-cost way to enrich vision-language-action (VLA) training, reducing reliance on expensive real-robot demonstrations. However, most sim-real co-training methods rely on supervised fine-tuning (SFT), which treats simulation as a static source of demonstrations and does not exploit large-scale closed-loop interaction. Consequently, real-world gains and generalization are often limited. In this paper, we propose an RL-based sim-real Co-training (RL-Co) framework that leverages interactive simulation while preserving real-world capabilities. Our method follows a generic two-stage design: we first warm-start the policy with SFT on a mixture of real and simulated demonstrations, then fine-tune it with reinforcement learning in simulation while adding an auxiliary supervised loss on real-world data to anchor the policy and mitigate catastrophic forgetting. We evaluate our framework on four real-world tabletop manipulation tasks using two representative VLA architectures, OpenVLA and $π_{0.5}$, and observe consistent improvements over real-only fine-tuning and SFT-based co-training, including +24% real-world success on OpenVLA and +20% on $π_{0.5}$. Beyond higher success rates, RL co-training yields stronger generalization to unseen task variations and substantially improved real-world data efficiency, providing a practical and scalable pathway for leveraging simulation to enhance real-robot deployment.","identifiers":{"arxiv":"2602.12628","doi":"10.48550/arXiv.2602.12628"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.12628"},{"label":"HTML","url":"https://arxiv.org/html/2602.12628v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.12628"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.12628v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Shuaihang Chen Affiliation: Tsinghua University Affiliation: Harbin Institute of Technology Affiliation: Zhongguancun Academy Equal contribution. Project Leader. Corresponding Authors: yuchao@sz.tsinghua.edu.cn , yu-wang@mail.tsinghua.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2602.12628v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2602-12628","updatedAt":"2026-08-30T06:21:20Z"},{"type":"preprint","title":"RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents","authors":[{"name":"Haitian Zhong"},{"name":"Jixiu Zhai"},{"name":"Lei Song"},{"name":"Jiang Bian"},{"name":"Qiang Liu"},{"name":"Tieniu Tan"}],"institutions":["zgca"],"rawAffiliations":["Haitian Zhong Affiliation: New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences Affiliation: Microsoft Research Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-02-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI","cs.CL"],"abstract":"Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall when within-group reward variation is low (e.g., more rollouts in a group receive the all 0 or all 1 reward), making the group-normalized advantage uninformative and yielding vanishing updates. To address this problem, we propose RC-GRPO (Reward-Conditioned Group Relative Policy Optimization), which treats exploration as a controllable steering problem via discrete reward tokens. We first fine-tune a Reward-Conditioned Trajectory Policy (RCTP) on mixed-quality trajectories with reward goal special tokens (e.g., , ) injected into the prompts, enabling the model to learn how to generate distinct quality trajectories on demand. Then during RL, we sample diverse reward tokens within each GRPO group and condition rollouts on the sampled token to improve within-group diversity, improving advantage gains. On the Berkeley Function Calling Leaderboard v4 (BFCLv4) multi-turn benchmark, our method yields consistently improved performance than baselines, and the performance on Qwen-2.5-7B-Instruct even surpasses all closed-source API models.","identifiers":{"arxiv":"2602.03025","doi":"10.48550/arXiv.2602.03025"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.03025"},{"label":"HTML","url":"https://arxiv.org/html/2602.03025v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.03025"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.03025v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Haitian Zhong Affiliation: New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences Affiliation: Microsoft Research Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2602.03025v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2602-03025","updatedAt":"2026-08-29T07:58:11Z"},{"type":"dataset","title":"TransNormal-Synthetic","authors":[{"name":"longxiang-ai"}],"institutions":["zgca"],"rawAffiliations":["1 Zhejiang University, 2 Zhongguancun Academy"],"relationType":"official-output","publishedAt":"2026-01-31","year":2026,"venue":"Hugging Face Datasets","status":"released","topics":["Open Dataset"],"abstract":"Dataset released alongside TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation","identifiers":{},"links":[{"label":"Dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic"},{"label":"Project","url":"https://github.com/longxiang-ai/TransNormal"}],"versions":[{"label":"Hugging Face dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, 2 Zhongguancun Academy","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TransNormal"}],"sources":["Official GitHub project README","Hugging Face"],"id":"work-59f5f8c7d054","updatedAt":"2026-08-08T12:09:53Z"},{"type":"conference","title":"TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation","authors":[{"name":"Mingwei Li"},{"name":"Hehe Fan"},{"name":"Yi Yang"}],"institutions":["zgca"],"rawAffiliations":["1 Zhejiang University, 2 Zhongguancun Academy","Mingwei Li 1,2 Hehe Fan 1 Yi Yang 1 , {}^{1,\\text{\\faIcon{envelope}}} 1 Zhejiang University, Hangzhou, China 2 Zhongguancun Academy, Beijing, China {mingweili, hehefan, yangyics}@zju.edu.cn Project page: https://longxiang-ai.github.io/TransNormal","1 Zhejiang University, Hangzhou, China 2 Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-31","year":2026,"venue":"ICML 2026","status":"published","topics":["cs.CV"],"abstract":"Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical properties often lead to catastrophic failures in conventional depth and normal sensors, hindering the deployment of embodied AI in scientific environments. We propose TransNormal, a novel framework that adapts pre-trained diffusion priors for single-step normal regression. To handle the lack of texture in transparent surfaces, TransNormal integrates dense visual semantics from DINOv3 via a cross-attention mechanism, providing strong geometric cues. Furthermore, we employ a multi-task learning objective and wavelet-based regularization to ensure the preservation of fine-grained structural details. To support this task, we introduce TransNormal-Synthetic, a physics-based dataset with high-fidelity normal maps for transparent labware. Extensive experiments demonstrate that TransNormal significantly outperforms state-of-the-art methods: on the ClearGrasp benchmark, it reduces mean error by 24.4% and improves 11.25° accuracy by 22.8%; on ClearPose, it achieves a 15.2% reduction in mean error. The code and dataset will be made publicly available at https://longxiang-ai.github.io/TransNormal.","identifiers":{"arxiv":"2602.00839","doi":"10.48550/arXiv.2602.00839"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2602.00839"},{"label":"PDF","url":"https://arxiv.org/pdf/2602.00839"},{"label":"Code","url":"https://github.com/longxiang-ai/TransNormal"},{"label":"Dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic"},{"label":"Model","url":"https://huggingface.co/Longxiang-ai/TransNormal"},{"label":"HTML","url":"https://arxiv.org/html/2602.00839v1"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2602.00839v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, 2 Zhongguancun Academy","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TransNormal"},{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, Hangzhou, China 2 Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2602.00839v1"}],"sources":["arXiv","Official GitHub project README","arXiv HTML"],"id":"doi-10-48550-arxiv-2602-00839","updatedAt":"2026-08-08T12:09:53Z"},{"id":"arxiv-2601-03260","type":"preprint","title":"SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval","authors":[{"name":"Chenyang Shao","institutions":["zgca"]},{"name":"Fengli Xu","institutions":["zgca"]},{"name":"Yong Li","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Department of Electronic Engineering, BNRist, Tsinghua University; Zhongguancun Academy, Beijing, China","Fengli Xu 1 1 footnotemark: 1 Affiliation: Department of Electronic Engineering, BNRist, Tsinghua University Affiliation: Zhongguancun Academy, Beijing, China Email: ˜˜ fenglixu@tsinghua.edu.cn"],"relationType":"affiliation","publishedAt":"2026-01-07","year":2026,"venue":"arXiv","status":"preprint","topics":["AI Agents","Scientific Retrieval","Knowledge Networks","Benchmark","cs.CE","cs.CL"],"abstract":"AI agents have seen widespread adoption in information retrieval for scientific research, giving rise to tools such as Deep Research. However, existing retrieval agents mainly rely on keyword- or embedding-based methods. While effective at capturing content-level similarities, they struggle to understand complex relational networks among scientific papers, such as identifying corroborating or conflicting studies and tracing technological lineages. This fundamental limitation often results in fragmented knowledge structures, misinterpreted research sentiment, and ineffective modeling of collective scientific progress. To address this limitation, we introduce SciNet, the first Scientific Network relation-aware dataset for information retrieval agents. Built on a meta-database of 269 million papers across 7 disciplines and containing 8,940 carefully designed tasks, SciNet systematically captures three levels of relational understanding: ego-centric retrieval of papers with novel knowledge structures, pairwise identification of scholarly relationships, and path-wise reconstruction of scientific evolution. Extensive evaluation of three categories of retrieval agents shows that their accuracy on relation-aware tasks often falls below 20%, highlighting a fundamental shortcoming of current retrieval paradigms. Importantly, in a downstream literature review application, agents empowered with SciNet achieve a 25.3% improvement in review quality, highlighting the critical value of relation-aware retrieval for deepening scientific insights. We publicly release SciNet at https://github.com/tsinghua-fib-lab/SciNet to support future research.","identifiers":{"arxiv":"2601.03260","doi":"10.48550/arXiv.2601.03260"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.03260"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.03260"},{"label":"Code","url":"https://github.com/tsinghua-fib-lab/SciNet"},{"label":"HTML","url":"https://arxiv.org/html/2601.03260v1"}],"versions":[{"label":"arXiv v2","url":"https://arxiv.org/abs/2601.03260v2"},{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.03260v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"arXiv PDF","sourceUrl":"https://arxiv.org/pdf/2601.03260"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Fengli Xu 1 1 footnotemark: 1 Affiliation: Department of Electronic Engineering, BNRist, Tsinghua University Affiliation: Zhongguancun Academy, Beijing, China Email: ˜˜ fenglixu@tsinghua.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.03260v1"}],"sources":["arXiv","Official PDF","arXiv HTML"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"preprint","title":"OmniSyn unifies target-aware molecular generation and optimization within a synthesis-native LLM framework across the human proteome","authors":[{"name":"Zheng QIN"},{"name":"Yuzhang Li"},{"name":"Yueqing Zhang"},{"name":"Yunshuo Zhao"},{"name":"Huan Yee Koh"},{"name":"Zheng Wan"},{"name":"Haocheng Ren"},{"name":"Changying Huang"},{"name":"Yiming Shi"},{"name":"Zhenguo Wu"},{"name":"Yaosen Min","institutions":["zgci"]},{"name":"Jing Yang"},{"name":"Xiao He"},{"name":"Duanhua Cao"}],"institutions":["zgci"],"rawAffiliations":["Hong Kong University of Science & Technology;","Tongji University;","East China Normal University;","Monash University;","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Designing target-specific bioactive molecules with actionable synthesis routes for the human proteome holds enormous potential for expanding therapeutic discovery, but remains a challenge. Existing target-aware generative models often depend on protein structures and generate molecules before assessing synthetic feasibility. Here we present OmniSyn, a protein-sequence-conditioned Mixture-of-Experts (MoE) language model that couples a task-conditioned interaction module with a synthesis-action decoder to generate molecules with explicit synthesis traces, thereby unifying de novo ligand generation, synthesizability projection and hit-to-lead (H2L) optimization within synthesis-traceable chemical space. OmniSyn is pre-trained with self-distillation and post-trained with task-specific reinforcement learning (RL) to adapt expert routing and optimize molecular properties across design modes. On unseen protein targets from MolGenBench, a real-world drug-discovery benchmark, OmniSyn achieves state-of-the-art performance across de novo design and H2L optimization, including target-awareness and hit-rediscovery metrics, despite relying only on protein sequences rather than three-dimensional (3D) pocket structures. By embedding synthesis planning into the design process, OmniSyn shifts molecular generation from a generate-then-filter paradigm toward design-with-synthesis paradigm, transforming virtual predictions into experimentally actionable candidates. Independent AiZynthFinder evaluation yielded retrosynthetic success rates of 68.47% for de novo generation and 71.92% for H2L optimization, improving over the strongest baselines by 61.3% and 184.2%, respectively, and supporting the synthetic feasibility of OmniSyn-generated molecules. In synthesizability projection, OmniSyn further converts outputs from external generative models into close analogues with improved retrosynthetic feasibility while preserving molecular similarity. Having established strong benchmark performance and external retrosynthetic feasibility, we next applied OmniSyn at human-proteome scale, spanning more than 21,000 targets. Rapid sequence-conditioned sampling enabled the construction of, to our knowledge, the largest human-proteome-scale generative virtual library, comprising 2.7 billion target-specific molecules, each accompanied by model-derived synthesis traces and target-specific prioritization scores. By enabling scalable target-specific molecular design with synthesis-aware generation across the human proteome, OmniSyn opens new opportunities for exploring previously inaccessible therapeutic targets, including those lacking experimentally resolved structures.","identifiers":{"doi":"10.64898/2026.09.02.748775"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.09.02.748775"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.09.02.748775"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.09.02.748775"}],"sources":["Crossref"],"id":"doi-10-64898-2026-09-02-748775","updatedAt":"2026-09-08T04:58:29Z"},{"type":"preprint","title":"scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning","authors":[{"name":"Shengjie Wang","institutions":["zgca"]},{"name":"Zongyong Hu","institutions":["zgca"]},{"name":"Yunlong Bie","institutions":["zgca"]},{"name":"Qijin Yin","institutions":["zgca"]},{"name":"Hechang Chen"},{"name":"Qiuyi Li","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Jilin University & Zhongguancun Academy;","Xiamen University & Zhongguancun academy;","Beijing University of Posts and Telecommunications & Zhongguancun Academy;","Zhongguancun Academy;","Jilin University","School of Artificial Intelligence, Jilin University, Jilin, China","Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learning stable biological representations. This challenge is particularly relevant to single-cell RNA sequencing, where sparsity, incomplete gene detection, and technical variation can obscure the underlying biological state. Here, we introduce scRep, a compact latent-space self-distillation framework for single-cell representation learning. Rather than reconstructing raw expression values, scRep aligns differently perturbed views of the same cell through a momentum-updated teacher–student architecture, with self-distillation objectives at both the cell and gene levels. This representation-centered formulation encourages the model to capture biological information that remains stable across incomplete and perturbed transcriptomic observations. Using frozen representations without task-specific fine-tuning, scRep pretrained on approximately 2.8 million cells achieves the strongest overall performance across the evaluated frozen-representation benchmarks, demonstrating strong sample efficiency. A larger-scale scRep model pretrained on 30.72 million cells further demonstrates that the framework remains effective when scaled to a substantially larger and more diverse corpus. Beyond cell identity, scRep prioritizes established marker genes, recovers transcription factor–associated gene programs with cell-type-specific activity, and preserves continuous developmental structure that supports graph-based pseudotime inference. We further show that pretraining performance is closely associated with biological diversity: reducing redundant cells while improving cell-type coverage can match or exceed the performance of larger, less balanced training corpora. Together, these results establish latent-space self-distillation as an effective alternative to expression reconstruction for single-cell foundation modeling and suggest that efficient scaling depends not only on the number of cells, but also on the learning objective and the biological diversity of the pretraining corpus.","identifiers":{"doi":"10.64898/2026.08.31.747784"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.08.31.747784"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.08.31.747784"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.08.31.747784"}],"sources":["Crossref"],"id":"doi-10-64898-2026-08-31-747784","updatedAt":"2026-09-05T04:45:42Z"},{"type":"preprint","title":"IgGM2: An All-Atom Foundation Model for Adaptive Immune Receptor Design","authors":[{"name":"Jian Ma","institutions":["zgca"]},{"name":"Fandi Wu"},{"name":"Lin Yao","institutions":["zgca"]},{"name":"Jing Gao","institutions":["zgca"]},{"name":"Rubo Wang"},{"name":"Qifeng Li","institutions":["zgca"]},{"name":"Nianzu Yang"},{"name":"Songlin Jiang"},{"name":"Dawei Huang"},{"name":"Xiaoyong Pan"},{"name":"Yiheng Zhu","institutions":["zgca","zgci"]},{"name":"Tingjun Hou","institutions":["zgca","zgci"]},{"name":"Jianhua Yao"},{"name":"Junchi Yan"}],"institutions":["zgca","zgci"],"rawAffiliations":["Shanghai Jiao Tong University","Zhongguancun Academy","Tencent, AI for Life Sciences Lab, Shenzhen, China","Zhongguancun Institute of Artificial Intelligence","Zhejiang University"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Accurate immune receptor design requires modeling the coupled variation of aminoacid sequence, full-atom conformation, and target-binding geometry across antibodies, nanobodies, and T-cell receptors (TCRs). Existing methods often address only part of this problem, either by separating structure generation from sequence design, relying on fixed-backbone inverse folding, or focusing on a single receptor class. We introduce IgGM2 , a unified all-atom generative framework for immune receptor structure prediction and CDR sequence–structure co-design. IgGM2 follows a structure-to-design strategy: it first learns how immune receptors are positioned around fixed target structures, and then transfers this target-conditioned structural prior to CDR design. Unlike modular design pipelines, IgGM2 jointly generates CDR residue identities and full-atom receptor structures, allowing frame-work geometry to adapt to designed CDRs without separate inverse folding or external sidechain packing. Unlike continuous residue encodings based on virtualatom geometry, IgGM2 keeps sequence prediction explicit while using atom14 placeholders only for full-atom representation. On structure prediction benchmarks, IgGM2 better captures receptor–target spatial relationships than AlphaFold3 on FoldBench and achieves strong performance on TCR–pMHC modeling. On sequence design benchmarks, IgGM2 achieves competitive amino-acid recovery and improves Rosetta-based interface preference metrics, suggesting more favorable generated binding interfaces. These results support IgGM2 as a unified all-atom framework for adaptive immune receptor structure prediction and design.","identifiers":{"doi":"10.64898/2026.07.09.737510"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.07.09.737510"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.07.09.737510"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.07.09.737510"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.07.09.737510"}],"sources":["Crossref"],"id":"doi-10-64898-2026-07-09-737510","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"Functional Locality–Aligned Learning Reveals Structure–Function Causality in Enzyme Kinetics","authors":[{"name":"Hao Zhang"},{"name":"He Zhang","institutions":["zgca","zgci"]},{"name":"Miao Kang"},{"name":"Kaipeng Zhang"},{"name":"Tao Yang"},{"name":"Nanning Zheng"}],"institutions":["zgca","zgci"],"rawAffiliations":["State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University","Zhongguancun Academy","Zhongguancun Institute of Artificial Intelligence","Shanghai Artificial Intelligence Laboratory"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Accurate estimation of enzyme kinetic parameters is essential for enzyme engineering and industrial biocatalysis, yet their experimental measurement remains labor-intensive and costly. Although machine learning offers an efficient alternative, existing methods still struggle to generalize to unseen enzymes and substrates. In particular, current three-dimensional (3D) structure–aware approaches rely on whole-enzyme 3D geometric structures and neglect substrate 3D geometry, often yielding limited or even degraded performance. We identify a fundamental limitation underlying these methods: a mismatch between the structural representation learning scale and the functional locality scale, which weakens structure–function causality in enzyme kinetics. To address this issue, we introduce EnzymePlex , a functional locality–aligned framework that aligns inductive biases with the localized structural determinants of enzyme function by prioritizing catalytic pockets, integrating substrate 3D geometry, and modeling nuanced enzyme–substrate interplay under the guidance of pocket-level structural priors. EnzymePlex achieves state-of-the-art performance across multiple benchmarks and substantially improves generalization under stringent out-of-distribution evaluations. Beyond predictive accuracy, EnzymePlex learns mechanistically aligned representations, with attention enriched at catalytic pocket residues and substrate reaction centers despite receiving no explicit super-vision for either. Moreover, when applied to recently reported wet-lab data, EnzymePlex effectively prioritizes high-activity enzyme variants and identifies potent inhibitors, highlighting its potential to accelerate enzyme engineering and drug discovery.","identifiers":{"doi":"10.64898/2026.03.04.709726"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.03.04.709726"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.03.04.709726"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.03.04.709726"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.03.04.709726"}],"sources":["Crossref"],"id":"doi-10-64898-2026-03-04-709726","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"Ophiuchus-Ab: A Versatile Generative Foundation Model for Advanced Antibody-Based Immunotherapy","authors":[{"name":"Yiheng Zhu","institutions":["zgca","zgci"]},{"name":"Jian Ma","institutions":["zgca"]},{"name":"Mingze Yin"},{"name":"Jialu Wu"},{"name":"Lin Tang","institutions":["zgca"]},{"name":"Zhiyun Zhang"},{"name":"Qiuyi Li","institutions":["zgca","zgci"]},{"name":"Shikun Feng","institutions":["zgca","zgci"]},{"name":"Haiguang Liu","institutions":["zgca","zgci"]},{"name":"Tao Qin","institutions":["zgca","zgci"]},{"name":"Junchi Yan"},{"name":"Chang-Yu Hsieh"},{"name":"Tingjun Hou","institutions":["zgca","zgci"]}],"institutions":["zgca","zgci"],"rawAffiliations":["Zhongguancun Academy, Beijing, China","Zhongguancun Institute of Artificial Intelligence, Beijing, China","Shanghai Jiao Tong University, Shanghai, China","Zhejiang University, Hangzhou, China","Carnegie Mellon University, Pittsburgh, USA","Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Antibodies exhibit extraordinary specificity and diversity in antigen recognition and have become a central class of therapeutics across a wide range of diseases. Despite this clinical success, antibody design remains fundamentally challenging. Antibody function emerges from intricate and highly coupled interactions between heavy and light chains, which complicate sequence-function relationships and limit the rational design of developable antibodies. Here, we reveal that modeling antibody sequence space at the level of paired heavy and light chains is essential to faithfully capture inter-chain dependencies, enabling a deeper understanding of antibody function and facilitating antibody discovery. We present Ophiuchus-Ab, a generative foundation model pre-trained on largescale paired antibody repertoires within a diffusion language modeling framework, unifying antibody generation and representation learning in a single probabilistic formulation. This framework excels diverse antibody design tasks, including CDR infilling, antibody humanization, and light-chain pairing. Beyond generation, diffusion-based pre-training yields transferable representations that enable accurate prediction of antibody properties, including developability, binding affinity, and specificity, even in low-data regimes. Together, these results establish Ophiuchus-Ab as a versatile foundation model for modeling antibodies, providing a foundation for next-generation antibody-based immunotherapy.","identifiers":{"doi":"10.64898/2026.02.02.703197"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.02.02.703197"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.02.02.703197"},{"label":"DOI 10.5281/zenodo.18478480","url":"https://doi.org/10.5281/zenodo.18478480"},{"label":"DOI 10.5281/zenodo.18478479","url":"https://doi.org/10.5281/zenodo.18478479"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.02.02.703197"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.02.02.703197"},{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18478480"},{"level":"structured","institution":"zgci","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18478480"},{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18478479"},{"level":"structured","institution":"zgci","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18478479"}],"sources":["Crossref","DataCite"],"id":"doi-10-64898-2026-02-02-703197","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"GENERator-v2: Reconciling Coarse Tokenization with Single-Nucleotide Resolution in Genomic Language Modeling","authors":[{"name":"Qiuyi Li","institutions":["zgca","zgci"]},{"name":"Zhihao Zhan"},{"name":"Shikun Feng","institutions":["zgca","zgci"]},{"name":"Yiheng Zhu","institutions":["zgca","zgci"]},{"name":"Yuan He","institutions":["zgca","zgci"]},{"name":"Wei Wu"},{"name":"Zhenghang Shi","institutions":["zgca"]},{"name":"Shengjie Wang","institutions":["zgca","zgci"]},{"name":"Zongyong Hu","institutions":["zgca","zgci"]},{"name":"Zhao Yang","institutions":["zgca","zgci"]},{"name":"Jiaoyang Li","institutions":["zgca","zgci"]},{"name":"Jian Tang"},{"name":"Haiguang Liu","institutions":["zgca","zgci"]},{"name":"Tao Qin","institutions":["zgca","zgci"]}],"institutions":["zgca","zgci"],"rawAffiliations":["Zhongguancun Academy, Beijing, China","Zhongguancun Institute of Artificial Intelligence, Beijing, China","Mila - Québec AI Institute, Montréal, Canada","University of Montréal, Montréal, Canada","University of Science and Technology of China, Hefei, China","Beijing Youth AI Academy (Haidian), Beijing, China","The Affiliated High School of Peking University, Beijng, China","HEC Montréal, Montréal, Canada","CIFAR AI Chair, Canada"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"A bstract Genomic foundation models aim to learn general-purpose representations directly from DNA sequence, enabling sequence understanding, generation, and probabilistic reasoning across a wide range of biological tasks. Scaling such models to genomic lengths, however, remains challenging due to the tension between long-range context, nucleotide-level resolution, and practical computational efficiency. Architectural innovations have enabled increasingly long nominal inputs, but often struggle to translate additional context into meaningful performance gains, particularly in the presence of sparse functional signal along eukaryotic genomes. In this work, we revisit the design of long-context genomic foundation models from the perspective of training objective and data construction. We introduce Factorized Nucleotide Supervision (FNS), which reconciles efficient k -mer tokenization with single-nucleotide likelihoods through probability marginalization, and Genome Compression Pretraining (GCP), which reshapes the training distribution by concentrating on gene-centric and regulatory regions. Together, these techniques enable standard transformer-based models to perform functional in-context learning without sacrificing nucleotide-level fidelity or computational efficiency. Building on these ideas, we present a family of autoregressive genomic foundation models supporting contexts of up to 98k base pairs across eukaryotic and prokaryotic genomes. Across training-free evaluations and downstream fine-tuning benchmarks, our models consistently improve over prior approaches and match or exceed state-of-the-art baselines while enabling substantially more efficient inference. Together, these results demonstrate that aligning supervision and data regimes with the biological structure of genomic sequence provides a principled and effective path toward scalable and biologically faithful genomic language modeling. Models, data, and scripts for downstream analyses are publicly available at https://huggingface.co/GenerTeam .","identifiers":{"doi":"10.64898/2026.01.27.702015"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.01.27.702015"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.01.27.702015"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.01.27.702015"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.01.27.702015"}],"sources":["Crossref"],"id":"doi-10-64898-2026-01-27-702015","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"Phenotype-Guided In Silico Molecular Generation Using Large Language Models","authors":[{"name":"Qun Jiang","institutions":["zgca"]},{"name":"Xing Ye"},{"name":"Zekun Guo","institutions":["zgci"]},{"name":"Yingce Xia","institutions":["zgca"]},{"name":"Zequn Liu","institutions":["zgca"]},{"name":"Jiaying Xu","institutions":["zgci"]},{"name":"Peiran Jin","institutions":["zgca"]},{"name":"Fusong Ju","institutions":["zgca"]},{"name":"Huanhuan Xia","institutions":["zgci"]},{"name":"Shangya Feng"},{"name":"Rui Jiang"},{"name":"Haiguang Liu","institutions":["zgca"]},{"name":"Tao Qin","institutions":["zgca"]},{"name":"Pan Deng","institutions":["zgca"]},{"name":"Sida Shao"}],"institutions":["zgca","zgci"],"rawAffiliations":["Beijing Zhongguancun Academy, Beijing, China","Ministry of Education Key Laboratory of Bioinformatics, Bioinformatics Division at the Beijing National Research Center for Information Science and Technology, Center for Synthetic and Systems Biology, Department of Automation, Tsinghua University, Beijing, China","Zhejiang Key Laboratory of Precise Synthesis of Functional Molecules, Department of Chemistry, Westlake University, Hangzhou, China","School of Life Sciences, Westlake University, Hangzhou, China","Zhongguancun Institute of Artificial Intelligence, Beijing, China","School of Global College, Shanghai Jiao Tong University, Shanghai, China","Research Center for Industries of the Future, Westlake University, Hangzhou, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Complex diseases often emerge from coordinated, system-level cellular state changes that are difficult to address with target centric drug discovery. Phenotypic drug discovery offers a principled alternative but remains constrained by the cost and scalability of pathologically relevant assays. Here we present GEMGen, a large language model–based framework that performs in silico phenotypic drug discovery by generating small molecules directly from transcriptomic representations of cellular states. GEMGen encodes desired phenotypic transitions as text-based representations of up- and down-regulated gene sets, enabling transferable modeling across experimental platforms and data modalities. Trained on large scale chemical perturbation data, GEMGen robustly identifies phenotype-oriented compounds and mechanistically related but structurally distinct candidates across multiple benchmarks. Applied to signatures induced by genetic perturbations, GEMGen produces small molecules that phenocopy gene knockdown effects and identifies chemically novel inhibitors, including previously unreported KEAP1 inhibitors that activate NRF2 signaling. Extending this approach to a disease relevant model of fibrosis, GEMGen generates compounds that reverse profibrotic transcriptional programs and cellular phenotypes. These results establish a scalable framework for translating transcriptomic phenotypes into candidate therapeutic molecules, enabling systems-level exploration of vast chemical space and offering a complementary in silico counterpart to physical phenotypic drug screens.","identifiers":{"doi":"10.64898/2026.01.03.697483"},"links":[{"label":"DOI","url":"https://doi.org/10.64898/2026.01.03.697483"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.64898/2026.01.03.697483"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beijing Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.01.03.697483"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.64898/2026.01.03.697483"}],"sources":["Crossref"],"id":"doi-10-64898-2026-01-03-697483","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"Dynamic Graph Prior-Guided Latent Adaptation for Generalizable Single-Cell Perturbation Response Prediction","authors":[{"name":"Bie, Yunlong"},{"name":"Fan, Peishan"},{"name":"Yin, Qijin"},{"name":"Xiao, Li"}],"institutions":["zgca"],"rawAffiliations":["Beijing University of Posts and Telecommunications","Beijing Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"preprint","topics":[],"abstract":"DGP-LA is a dynamic graph prior-guided latent adaptation framework for generalizable single-cell drug perturbation response prediction. It combines a cell-state-conditioned transition model with perturbation-specific refinement of drug-target interaction and protein-protein interaction priors, topology-aware propagation, and support-aware latent adaptation. This software archive contains the DGP-LA implementation, final configurations, preprocessing and prior-construction utilities, split specifications, inference and evaluation code, and tests supporting the associated manuscript.","identifiers":{"doi":"10.5281/zenodo.22155397"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.22155397"}],"versions":[{"label":"DOI 10.5281/zenodo.22155398","url":"https://doi.org/10.5281/zenodo.22155398"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Beijing Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.22155397"},{"level":"structured","institution":"zgca","matchedText":"Beijing Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.22155398"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-22155397","updatedAt":"2026-08-29T07:14:00Z"},{"type":"dataset","title":"ProtChord: Model Checkpoint, Processed Data, and Evaluation Results","authors":[{"name":"Zhang, Wei"},{"name":"Guo, Zekun"},{"name":"Xia, Yingce"},{"name":"Jin, Peiran"},{"name":"Xie, Shufang"},{"name":"Liu, Zequn"},{"name":"Li, Xiang-Yang"},{"name":"Qin, Tao"}],"institutions":["zgca"],"rawAffiliations":["University of Science and Technology of China","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"released","topics":[],"abstract":"Research artifacts accompanying the paper\"ProtChord: Structure-Sequence Alignment for Protein-Centered Discovery\". This record contains the released ProtChord-RL model checkpoint, processed CrossDocked2020 data used for supervised fine-tuning, preference optimization, and evaluation, and raw molecular-generation outputs used in the experiments. Source code, inference scripts, evaluation scripts, and configuration files are available at:https://github.com/weizhang-ustc/ProtChord","identifiers":{"doi":"10.5281/zenodo.22093629"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.22093629"}],"versions":[{"label":"DOI 10.5281/zenodo.22093628","url":"https://doi.org/10.5281/zenodo.22093628"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.22093629"},{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.22093628"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-22093629","updatedAt":"2026-08-26T02:10:03Z"},{"type":"preprint","title":"Electron-density-informed supervision for physically consistent Hamiltonian prediction","authors":[{"name":"Wang, Yifei"},{"name":"Zhang, He"},{"name":"Wang, Jianji"},{"name":"Wei, Xinran"},{"name":"Ma, Yongqiang"},{"name":"Kang, Miao"},{"name":"Liu, Chang"},{"name":"Zheng, Nanning"}],"institutions":["zgca","zgci"],"rawAffiliations":["Xi'an Jiaotong University","Zhongguancun Academy","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"preprint","topics":["density functional theory","AI for science","deep learning"],"abstract":"This is the implementation of the paper \"Electron-density-informed supervision for physically consistent Hamiltonian prediction\".","identifiers":{"doi":"10.5281/zenodo.21257875"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.21257875"}],"versions":[{"label":"DOI 10.5281/zenodo.21257874","url":"https://doi.org/10.5281/zenodo.21257874"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.21257875"},{"level":"structured","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.21257875"},{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.21257874"},{"level":"structured","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.21257874"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-21257875","updatedAt":"2026-08-08T11:35:31Z"},{"type":"dataset","title":"PolyTamGen Data, Base Language Model, DPO Checkpoint, and Benchmark Results","authors":[{"name":"Wu, Kehan"},{"name":"Zhang, Wei"},{"name":"Guo, Han"},{"name":"Liu, Renhe"},{"name":"Shi, Yu"},{"name":"Jin, Peiran"},{"name":"Liu, Haiguang"},{"name":"Chen, Enhong"},{"name":"Qin, Tao"},{"name":"Liu, Zequn"},{"name":"Xia, Yingce"}],"institutions":["zgca"],"rawAffiliations":["University of Science and Technology of China, Hefei, Anhui, China; Zhongguancun Academy, Beijing, China","Global Health Drug Discovery Institute, Beijing, China","Zhongguancun Academy, Beijing, China","University of Science and Technology of China, Hefei, Anhui, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"released","topics":["PolyTamGen","polypharmacology","multi-target drug design","molecular generation","SMILES generation","protein-conditioned generation"],"abstract":"This record contains release artifacts for PolyTamGen, a sequence-based multi-target molecule generation framework for polypharmacological target-aware molecular design. The record contains four release files: polytamgen_data.tar.gzProcessed training, benchmark, and feature-store data assets used by the released PolyTamGen code. base-llm_NLM-1b.tarBase NatureLM/NLM language-model directory required for real PolyTamGen checkpoint inference. Extract this archive under source/pp2d/models/ so that it materializes source/pp2d/models/base-llm/NLM-1b/. polytamgen-dpo.pthFinal PolyTamGen-DPO checkpoint used for the main benchmark generation. This slim checkpoint omits optimizer/scheduler states and duplicated base-LLM weights; it requires base-llm_NLM-1b.tar. Place it at source/pp2d/models/polytamgen-dpo/polytamgen-dpo.pth. polytamgen_full_results.tar.gzBenchmark results bundle, including molecule-level scored results, raw-generation outputs, supplement materials, figures, and analysis inputs used to reproduce the reported paper metrics. PolyTamGen conditions a NatureLM SMILES generator on sets of protein targets. The released DPO workflow uses ESM3-derived protein features, a Q-Former-style protein adapter, set-wise target interaction, and direct preference optimization with docking/QED preferences. After downloading the companion source repository, materialize the archives under source/pp2d/ and follow the setup, inference, and paper-alignment instructions in the repository README files. DPO inference uses configs/polytamgen_dpo.yaml, models/base-llm/NLM-1b/, and models/polytamgen-dpo/polytamgen-dpo.pth. The record does not include SFT, ablation, adapter-ablation, or exploratory training checkpoints.","identifiers":{"doi":"10.5281/zenodo.20594128"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.20594128"}],"versions":[{"label":"DOI 10.5281/zenodo.20594127","url":"https://doi.org/10.5281/zenodo.20594127"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"University of Science and Technology of China, Hefei, Anhui, China; Zhongguancun Academy, Beijing, China","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.20594128"},{"level":"structured","institution":"zgca","matchedText":"University of Science and Technology of China, Hefei, Anhui, China; Zhongguancun Academy, Beijing, China","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.20594127"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-20594128","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"starVLA/starVLA: Framwork standard on VLM4A and WM4A","authors":[{"name":"Jinhui Ye"},{"name":"Axi404"},{"name":"Yilun Chen"},{"name":"MichaelYu"},{"name":"Fangjing Wang"},{"name":"Zeng Shuang"},{"name":"Qixiu Li"},{"name":"Microsoft Open Source"},{"name":"HENG"},{"name":"Shijie Lian"},{"name":"Bin Yu"},{"name":"GUO Weiyu"},{"name":"Zixuan WANG"},{"name":"Jiachen Shen"},{"name":"Danyal Ahmad"},{"name":"Cheng Yin"},{"name":"zhihelu-zero"},{"name":"yuxin chen"},{"name":"rakybond007"},{"name":"jrryzh(SII)"},{"name":"Zhijie-Song"},{"name":"Yi Yang (SII)"},{"name":"Yang Tian"},{"name":"Wendi Chen (SII)"},{"name":"Travor King"},{"name":"Sukai Huang"},{"name":"Senqiao Yang (杨森乔)"},{"name":"LiuRicky"}],"institutions":["zgca"],"rawAffiliations":["HKUST","Xi'an Jiaotong University","The Chinese University of Hong Kong","Fudan University","Robbyant","Tsinghua University/Microsoft","Microsoft","Huazhong University of Science and Technology","Harbin Institute of Technology & Zhongguancun Academy","Twolabs Inc","THUNLP","The Hong Kong University of Science and Technology","SII","Shanghai Jiao Tong University","Shanghai Jiao Tong University & Shanghai Innovation Institute"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"preprint","topics":[],"abstract":"StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing","identifiers":{"doi":"10.5281/zenodo.20336846"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.20336846"}],"versions":[{"label":"DOI 10.5281/zenodo.18264213","url":"https://doi.org/10.5281/zenodo.18264213"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Harbin Institute of Technology & Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.20336846"},{"level":"structured","institution":"zgca","matchedText":"Harbin Institute of Technology & Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18264213"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-20336846","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing","authors":[{"name":"Jinhui Ye"},{"name":"MichaelYu"},{"name":"Fangjing Wang"},{"name":"Axi404"},{"name":"Yilun Chen"},{"name":"Qixiu Li"},{"name":"Microsoft Open Source"},{"name":"Bin Yu"},{"name":"Chengyao Wang"},{"name":"Jiang Changjiu"},{"name":"Lipeng Wang"},{"name":"Senqiao Yang (杨森乔)"},{"name":"Zixuan WANG"},{"name":"jrryzh(SII)"}],"institutions":["zgca"],"rawAffiliations":["HKUST","Fudan University","SUSTech","Xi'an Jiaotong University","The Chinese University of Hong Kong","Tsinghua University/Microsoft","Microsoft","Harbin Institute of Technology & Zhongguancun Academy","SII"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Zenodo","status":"preprint","topics":[],"abstract":"StarVLA is a modular and flexible codebase for developing Vision-Language Model (VLM) to Vision-Language-Action (VLA) models. In StarVLA (also a pun on “start VLA” ), each functional component (model, data, trainer, config, evaluation, etc.) follows a top-down, intuitive separation and high cohesion and low coupling principle, which enabling plug-and-play design, rapid prototyping, and independent debugging.","identifiers":{"doi":"10.5281/zenodo.18264214"},"links":[{"label":"DOI","url":"https://zenodo.org/doi/10.5281/zenodo.18264214"}],"versions":[],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Harbin Institute of Technology & Zhongguancun Academy","source":"DataCite","sourceUrl":"https://zenodo.org/doi/10.5281/zenodo.18264214"}],"sources":["DataCite"],"id":"doi-10-5281-zenodo-18264214","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"MIDAS: Mutual Information Disentanglement With Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis","authors":[{"name":"Yuhua Wen"},{"name":"Yingying Zhou"},{"name":"Qifei Li"},{"name":"Yingming Gao"},{"name":"Zhengqi Wen"},{"name":"Jianhua Tao"},{"name":"Ya Li"}],"institutions":["zgca"],"rawAffiliations":["Yuhua Wen, Yingying Zhou, Qifei Li, Yingming Gao, Zhengqi Wen, Jianhua Tao, and Ya Li This work is supported by the National Key R&D Program of China under Grant No.2024YFB2808802. (Corresponding author: Ya Li.)Yuhua Wen is with the School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China, and also with the Zhongguancun Academy, Beijing 100094, China (e-mail: yuhuawen@bupt.edu.cn).Yingying Zhou, Qifei Li, Yingming Gao, and Ya Li are with the School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China (e-mail: yingyingzhou@bupt.edu.cn; liqifei@bupt.edu.cn; yingming.gao@bupt.edu.cn; yli01@bupt.edu.cn).Zhengqi Wen is with the Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China (e-mail: zqwen@tsinghua.edu.cn).Jianhua Tao is with"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that minimizes the mutual information between shared and exclusive spaces for stable disentanglement, while maximizing the mutual information among shared spaces across modalities to enhance semantic alignment. In addition, an uncertainty-aware fusion mechanism is introduced, where posterior variance is leveraged as a reliability indicator to adaptively weight latent features during fusion, ensuring robust integration even when modalities are incomplete. Extensive experiments on three widely used datasets show that MIDAS achieves strong and consistent performance gains over competitive baselines across a wide range of incomplete settings, demonstrating its effectiveness and robustness for incomplete data scenarios.","identifiers":{"arxiv":"2608.09986","doi":"10.48550/arXiv.2608.09986","publishedDoi":"10.1109/tpami.2026.3713694"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2608.09986"},{"label":"HTML","url":"https://arxiv.org/html/2608.09986v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2608.09986"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2608.09986v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuhua Wen, Yingying Zhou, Qifei Li, Yingming Gao, Zhengqi Wen, Jianhua Tao, and Ya Li This work is supported by the National Key R&D Program of China under Grant No.2024YFB2808802. (Corresponding author: Ya Li.)Yuhua Wen is with the School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China, and also with the Zhongguancun Academy, Beijing 100094, China (e-mail: yuhuawen@bupt.edu.cn).Yingying Zhou, Qifei Li, Yingming Gao, and Ya Li are with the School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China (e-mail: yingyingzhou@bupt.edu.cn; liqifei@bupt.edu.cn; yingming.gao@bupt.edu.cn; yli01@bupt.edu.cn).Zhengqi Wen is with the Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China (e-mail: zqwen@tsinghua.edu.cn).Jianhua Tao is with","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2608.09986v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2608-09986","updatedAt":"2026-08-16T02:48:42Z"},{"type":"preprint","title":"Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning","authors":[{"name":"Hui Li"},{"name":"Yulu Gao"},{"name":"Si Liu"},{"name":"Yuhang Wang"},{"name":"Bo Liu"},{"name":"Beipeng Mu"}],"institutions":["zgca"],"rawAffiliations":[],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements like lane centerlines and their topology. Existing lane topology reasoning methods typically follow a reasoning-by-detection paradigm, where lane topological relationships are primarily derived from lane detection results. In this paper, we propose an innovative method called Unified Modeling of Lane and Lane Topology (UniTopo), which represents the topological relationships between lanes as connected lanes, encompassing predecessor lanes, successor lanes, and their interconnections. This unified representation of lanes and lane topology allows us to simultaneously obtain both the positions and topological information of lanes within a shared perception pipeline, establishing a new paradigm for directly perceiving lane topology from original image features. We validate our method on the driving scene reasoning benchmark OpenLane-V2, which consists of two subsets, built based on Argoverse2 and nuScenes, respectively. Our method achieves TOPllof 30.1% and 31.8% on the two subsets, significantly surpassing the existing state-ofthe- art method T2SG by 6.0% and 8.6%.","identifiers":{"arxiv":"2605.08911","doi":"10.48550/arXiv.2605.08911","publishedDoi":"10.1109/tcsvt.2026.3690152"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2605.08911"},{"label":"HTML","url":"https://arxiv.org/html/2605.08911v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.08911"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.08911v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beipeng Mu Thanks: Han Li is with the School of Artificial Intelligence, Beihang University, Beijing, China, and also with Zhongguancun Academy, Beijing, China. E-mail: lihan0620@buaa.edu.cn. Thanks: Yulu Gao is with the Hangzhou International Innovation Institute, Beihang University, Hangzhou, China. E-mail: gyl97@buaa.edu.cn. Thanks: Si Liu is with the School of Artificial Intelligence, Beihang University, Beijing, China. E-mail: liusi@buaa.edu.cn. Thanks: Yuhang Wang, Bo Liu, and Beipeng Mu are with Meituan, Beijing, China. E-mails: {wangyuhang11, mubeipeng}@meituan.com, boliu1995@163.com. Thanks: The first two authors contributed equally to this work, and the corresponding author is Si Liu. Thanks: Digital Object Identifier 10.1109/TCSVT.2026.3690152","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.08911v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2605-08911","updatedAt":"2026-08-13T11:57:17Z"},{"type":"preprint","title":"Unleashing the Agility of Wheeled-Legged Robots for High-Dynamic Reflexive Obstacle Evasion","authors":[{"name":"Zhao, Yongen"},{"name":"Xu, Zihao"},{"name":"Lu, Wenzhi"},{"name":"Chu, Zhen"},{"name":"Hao, Ce"}],"institutions":["zgca"],"rawAffiliations":["School of Mechanical Engineering, Tianjin University, Tianjin, China","Beijing Zhongguancun Academy, Beijing, China","School of Computing, National University of Singapore, Singapore","DeepRobotics, Hangzhou, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["Robotics (cs.RO)","FOS: Computer and information sciences"],"abstract":"Wheeled-legged robots combine the energy efficiency of wheeled locomotion with the terrain adaptability of legged systems, making them promising platforms for agile mobility in complex and dynamic environments. However, enabling high-dynamic reflexive evasion against fast-moving obstacles remains challenging due to the hybrid morphology, mode coupling, and non-holonomic constraints of such platforms. In this work, we propose AWARE, Adaptive Wheeled-Legged Avoidance and Reflexive Evasion, a hierarchical reinforcement learning framework for high-dynamic obstacle avoidance in wheeled-legged robots. The proposed system naturally exhibits diverse emergent gaits and evasive behaviors, including forward lunge and lateral dodge, thereby leveraging the robot's hybrid morphology to enhance agility under highly dynamic threats. Extensive experiments in Isaac Lab simulation and real-world deployment on the M20 platform across diverse dynamic scenarios demonstrate that AWARE achieves robust and agile obstacle avoidance while revealing behaviorally distinct evasive strategies. These results highlight both the practical effectiveness of AWARE and the intrinsic reflexive agility of wheeled-legged robots.","identifiers":{"doi":"10.48550/arxiv.2604.23761"},"links":[{"label":"DOI","url":"https://arxiv.org/abs/2604.23761"}],"versions":[],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Beijing Zhongguancun Academy, Beijing, China","source":"DataCite","sourceUrl":"https://arxiv.org/abs/2604.23761"}],"sources":["DataCite"],"id":"doi-10-48550-arxiv-2604-23761","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?","authors":[{"name":"Meng, Lve"},{"name":"Zhao, Weilong"},{"name":"Zhang, Yanzhi"},{"name":"Guan, Haoxiang"},{"name":"He, Jiyan"}],"institutions":["zgca"],"rawAffiliations":["University of Science,Technology of China, Zhongguancun Academy","Université Paris Cité","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["Artificial Intelligence (cs.AI)","Commutative Algebra (math.AC)","Combinatorics (math.CO)","Category Theory (math.CT)","FOS: Computer and information sciences"],"abstract":"Large language models (LLMs) have recently achieved remarkable success in generating rigorous mathematical proofs, with \"AI for Math\" emerging as a vibrant field of research (Ju et al., 2026). While these models have mastered competition-level benchmarks like the International Mathematical Olympiad (Huang et al., 2025; Duan et al., 2025) and show promise in research applications through auto-formalization (Wang et al., 2025), their deployment via lightweight, natural-language pipelines for research problems remains underexplored. In this work, we demonstrate that next-generation models (e.g., Gemini 3 Pro, GPT-5.2 Pro), when integrated into a streamlined automated pipeline optimized for citation-based verification, can solve sophisticated research-grade problems. We evaluate our pipeline on two novel datasets: (1) the ICCM (2025) problem sets (comparable to the S.-T. Yau College Student Mathematics Contest) proposed by leading mathematicians (Shanghai Math Challenge, 2026), and (2) the \"First Proof\" problem set (Abouzaid et al., 2026), consisting of previously unpublished research questions. Our pipeline generated candidate proofs for all problems in the first two ICCM sets and the \"First Proof\" set. The solutions for the first two ICCM sets and Problem 4 of the \"First Proof\" set have been fully verified by our team. All generated proofs have been submitted to the official organization, and our generated results are publicly available at https://github.com/ml1301215/question_sets-test_results. We have open-sourced the code and developed a user-friendly UI for this workflow, accessible at https://github.com/ml1301215/research-math-assistant.","identifiers":{"doi":"10.48550/arxiv.2602.13695"},"links":[{"label":"DOI","url":"https://arxiv.org/abs/2602.13695"}],"versions":[],"evidence":[{"level":"structured","institution":"zgca","matchedText":"University of Science,Technology of China, Zhongguancun Academy","source":"DataCite","sourceUrl":"https://arxiv.org/abs/2602.13695"}],"sources":["DataCite"],"id":"doi-10-48550-arxiv-2602-13695","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence","authors":[{"name":"Zixing Lei"},{"name":"Genjia Liu"},{"name":"Yuanshuo Zhang"},{"name":"Qipeng Liu"},{"name":"Yuzhu Cai"},{"name":"Sixiang Chen"},{"name":"Jixian Wu"},{"name":"Yunhong Wang"},{"name":"Weixin Li"},{"name":"Chuan Wen"},{"name":"Bo Zhao"},{"name":"Shanghang Zhang"},{"name":"Wenzhao Lian"},{"name":"Siheng Chen"}],"institutions":["zgca"],"rawAffiliations":["Affiliation: Zhongguancun Academy Affiliation: School of Integrated Circuits, Shanghai Jiao Tong University Affiliation: School of Computer Science, Shanghai Jiao Tong University"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI","cs.RO"],"abstract":"The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this scaling capability remains severely bottlenecked by a reliance on labor-intensive manual oversight from intricate reward shaping to hyperparameter tuning across heterogeneous backends. Inspired by LLMs' success in software automation and science discovery, we introduce \\textsc{EmboCoach-Bench}, a benchmark evaluating the capacity of LLM agents to autonomously engineer embodied policies. Spanning 32 expert-curated RL and IL tasks, our framework posits executable code as the universal interface. We move beyond static generation to assess a dynamic closed-loop workflow, where agents leverage environment feedback to iteratively draft, debug, and optimize solutions, spanning improvements from physics-informed reward design to policy architectures such as diffusion policies. Extensive evaluations yield three critical insights: (1) autonomous agents can qualitatively surpass human-engineered baselines by 26.5\\% in average success rate; (2) agentic workflow with environment feedback effectively strengthens policy development and substantially narrows the performance gap between open-source and proprietary models; and (3) agents exhibit self-correction capabilities for pathological engineering cases, successfully resurrecting task performance from near-total failures through iterative simulation-in-the-loop debugging. Ultimately, this work establishes a foundation for self-evolving embodied intelligence, accelerating the paradigm shift from labor-intensive manual tuning to scalable, autonomous engineering in embodied AI field.","identifiers":{"arxiv":"2601.21570","doi":"10.48550/arXiv.2601.21570"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.21570"},{"label":"HTML","url":"https://arxiv.org/html/2601.21570v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.21570"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.21570v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Affiliation: Zhongguancun Academy Affiliation: School of Integrated Circuits, Shanghai Jiao Tong University Affiliation: School of Computer Science, Shanghai Jiao Tong University","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.21570v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-21570","updatedAt":"2026-08-28T12:20:54Z"},{"type":"preprint","title":"ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment","authors":[{"name":"Xiuyu Li"},{"name":"Jinkai Zhang"},{"name":"Mingyang Yi"},{"name":"Yu Li"},{"name":"Longqiang Wang"},{"name":"Yue Wang"},{"name":"Ju Fan"}],"institutions":["zgca"],"rawAffiliations":["Yue Wang Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.LG"],"abstract":"Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To address this, we propose a training-free inference method to sample directly from the optimal RL policy. The transition probability applied to Masked Language Modeling (MLM) consists of a reference policy model and an energy term. Based on this, our algorithm, Energy-Guided Test-Time Scaling (ETS), estimates the key energy term via online Monte Carlo, with a provable convergence rate. Moreover, to ensure practical efficiency, ETS leverages modern acceleration frameworks alongside tailored importance sampling estimators, substantially reducing inference latency while provably preserving sampling quality. Experiments on MLM (including autoregressive models and diffusion language models) across reasoning, coding, and science benchmarks show that our ETS consistently improves generation quality, validating its effectiveness and design. The code is available at https://github.com/sheriyuo/ETS.","identifiers":{"arxiv":"2601.21484","doi":"10.48550/arXiv.2601.21484"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.21484"},{"label":"HTML","url":"https://arxiv.org/html/2601.21484v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.21484"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.21484v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yue Wang Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.21484v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-21484","updatedAt":"2026-08-28T12:20:54Z"},{"type":"preprint","title":"GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents","authors":[{"name":"Yanxi Wang"},{"name":"Zhiling Zhang"},{"name":"Wenbo Zhou"},{"name":"Weiming Zhang"},{"name":"Jie Zhang"},{"name":"Qiannan Zhu"},{"name":"Yu Shi"},{"name":"Shuxin Zheng"},{"name":"Jiyan He"}],"institutions":["zgca"],"rawAffiliations":["Yanxi Wang † † thanks: Equal contribution. Affiliation: Beijing Normal University Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CR","cs.AI","cs.CV"],"abstract":"As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locations, and behavioral traces. While existing benchmarks primarily focus on task completion, grounding, or defenses against third-party attacks, current visual privacy datasets remain largely restricted to static natural images, limiting their ability to capture the contextual dependence and task relevance of privacy risks in GUI task trajectories. To bridge this gap, we introduce \\textbf{GUIGuard-Bench}, a first-step benchmark for studying privacy-preserving GUI agents in trajectory-based GUI workflows. GUIGuard-Bench contains 241 real GUI-agent trajectories with 4,080 screenshots across Android and PC environments. Each screenshot is annotated at the region level with privacy bounding boxes, semantic privacy categories, risk levels, and whether the private information is necessary for completing the task. Built on these annotations, GUIGuard-Bench supports three complementary evaluations: privacy recognition, offline planning fidelity under protected screenshots, and the utility impact of different protection strategies. Our results show that current models can often detect whether a screenshot contains private information, but they struggle with fine-grained localization, category recognition, risk assessment, and task-necessity judgment. We also find that closed-source models, exemplified by Claude Sonnet 4.6, can maintain largely consistent planner semantics in Android environments after privacy protection is applied. Our results highlight privacy recognition as a critical bottleneck for practical GUI agents. Project: https://futuresis.github.io/GUIGuard-page/","identifiers":{"arxiv":"2601.18842","doi":"10.48550/arXiv.2601.18842"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.18842"},{"label":"HTML","url":"https://arxiv.org/html/2601.18842v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.18842"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.18842v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yanxi Wang † † thanks: Equal contribution. Affiliation: Beijing Normal University Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.18842v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-18842","updatedAt":"2026-08-28T12:20:54Z"},{"type":"preprint","title":"Towards a Theoretical Understanding to the Generalization of RLHF","authors":[{"name":"Li, Zhaochun"},{"name":"Yi, Mingyang"},{"name":"Wang, Yue"},{"name":"Cui, Shisheng"},{"name":"Liu, Yong"}],"institutions":["zgca"],"rawAffiliations":["Beijing Institute of Technolegy","Zhongguancun Academy","Renmin University of China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["Machine Learning (cs.LG)","FOS: Computer and information sciences"],"abstract":"Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically effective, the theoretical generalization properties of these methods in high-dimensional settings remain to be explored. To this end, we build the generalization theory on RLHF of LLMs under the linear reward model, through the framework of algorithmic stability. In contrast to the existing works built upon the consistency of maximum likelihood estimations on reward model, our analysis is presented under an end-to-end learning framework, which is consistent with practice. Concretely, we prove that under a key \\textbf{feature coverage} condition, the empirical optima of policy model have a generalization bound of order $\\mathcal{O}(n^{-\\frac{1}{2}})$. Moreover, the results can be extrapolated to parameters obtained by gradient-based learning algorithms, i.e., Gradient Ascent (GA) and Stochastic Gradient Ascent (SGA). Thus, we argue that our results provide new theoretical evidence for the empirically observed generalization of LLMs after RLHF.","identifiers":{"doi":"10.48550/arxiv.2601.16403"},"links":[{"label":"DOI","url":"https://arxiv.org/abs/2601.16403"}],"versions":[],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy","source":"DataCite","sourceUrl":"https://arxiv.org/abs/2601.16403"}],"sources":["DataCite"],"id":"doi-10-48550-arxiv-2601-16403","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"AI for Mathematics: Progress, Challenges, and Prospects","authors":[{"name":"Haocheng Ju"},{"name":"Bin Dong"}],"institutions":["zgca"],"rawAffiliations":["Bin Dong Address: School of Mathematical Sciences, Peking University, Beijing 100871, China Address: Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University, Beijing 100871, China Address: Center for Machine Learning Research, Peking University, Beijing 100871, China Address: Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University, Dongguan 523000, China Address: Zhongguancun Academy, Beijing 100094, China Abstract AI for Mathematics (AI4Math) has emerged as a distinct field that leverages machine learning to navigate mathematical landscapes historically intractable for early symbolic systems. While mid-20th-century symbolic approaches successfully automated formal logic, they faced severe scalability limitations due to the combinatorial explosion of the search space. The recent integratio"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["math.HO"],"abstract":"AI for Mathematics (AI4Math) has emerged as a distinct field that leverages machine learning to navigate mathematical landscapes historically intractable for early symbolic systems. While mid-20th-century symbolic approaches successfully automated formal logic, they faced severe scalability limitations due to the combinatorial explosion of the search space. The recent integration of data-driven approaches has revitalized this pursuit. In this review, we provide a systematic overview of AI4Math, highlighting its primary focus on developing AI models to support mathematical research. Crucially, we emphasize that this is not merely the application of AI to mathematical activities; it also encompasses the development of stronger AI systems where the rigorous nature of mathematics serves as a premier testbed for advancing general reasoning capabilities. We categorize existing research into two complementary directions: problem-specific modeling, involving the design of specialized architectures for distinct mathematical tasks, and general-purpose modeling, focusing on foundation models capable of broader reasoning, retrieval, and exploratory workflows. We conclude by discussing key challenges and prospects, advocating for AI systems that go beyond facilitating formal correctness to enabling the discovery of meaningful results and unified theories, recognizing that the true value of a proof lies in the insights and tools it offers to the broader mathematical landscape.","identifiers":{"arxiv":"2601.13209","doi":"10.48550/arXiv.2601.13209"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.13209"},{"label":"HTML","url":"https://arxiv.org/html/2601.13209v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.13209"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.13209v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Bin Dong Address: School of Mathematical Sciences, Peking University, Beijing 100871, China Address: Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University, Beijing 100871, China Address: Center for Machine Learning Research, Peking University, Beijing 100871, China Address: Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University, Dongguan 523000, China Address: Zhongguancun Academy, Beijing 100094, China Abstract AI for Mathematics (AI4Math) has emerged as a distinct field that leverages machine learning to navigate mathematical landscapes historically intractable for early symbolic systems. While mid-20th-century symbolic approaches successfully automated formal logic, they faced severe scalability limitations due to the combinatorial explosion of the search space. The recent integratio","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.13209v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-13209","updatedAt":"2026-08-27T10:50:02Z"},{"type":"preprint","title":"STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories","authors":[{"name":"Qiuyu Tian"},{"name":"Yiding Li"},{"name":"Yingce Xia"},{"name":"Fengyi Chen"},{"name":"Youyong Kong"},{"name":"Fan Guo"},{"name":"Yuyao Li"},{"name":"Jinjing Shen"},{"name":"Zhijing Xie"},{"name":"Yiyun Luo"},{"name":"Xin Zhang"},{"name":"Zequn Liu"}],"institutions":["zgca"],"rawAffiliations":["Qiuyu Tian Yiding Li Fengyi Chen Zequn Liu Youyong Kong Fan Guo Yuyao Li Jinjing Shen Zhijing Xie Yiyun Luo Xin Zhang Southeast University, Nanjing, China Beijing Zhongguancun Academy, Beijing, China Nanjing Normal University, Nanjing, China ZhuiWen Technology Co., Ltd., Beijing, China qiuyutian@seu.edu.cn, liyiding@zhuiwen.net, fxc1494@g.rit.edu, liuzequn@bjzgca.edu.cn, kongyouyong@seu.edu.cn, guofan@zhuiwen.net, yuyaoli@zhuiwen.net, shenjinjing@zhuiwen.net, xiezhijing@zhuiwen.net, luoyiyun@zhuiwen.net, xinzhang@zhuiwen.net"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL","cs.AI"],"abstract":"Movie screenplays are a demanding testbed for long-form narrative understanding, as characters' goals, beliefs, knowledge, and relationships evolve continuously across scenes. However, existing benchmarks primarily evaluate isolated facts from the completed screenplay, leaving unassessed whether models can track the evolving state of characters as the story unfolds. We introduce STAGE, a benchmark over 151 English and Chinese full-length screenplays, built on a provenance-linked narrative backbone that recovers the state and epistemic access of each character at every point along its timeline. Three tasks derived from the backbone jointly probe whether models can maintain, explain, and act on evolving narrative state: Character Development Tracking updates a focal character's state between checkpoints, Cross-Scene Narrative Evolution Reasoning targets cross-scene state transitions, and In-Script Character Role-Playing requires responses bounded by the character's state and knowledge at a specified point. We identify three failure modes of current LLMs: silent forgetting under recursive state updating, limited cross-scene reasoning even when all relevant evidence is supplied, and a trade-off in role-playing where stylistic character fidelity and screenplay-grounded memory faithfulness are optimized by different memory-access strategies. STAGE thus provides a unified framework for diagnosing how current models fail to track, reason about, and enact story evolution.","identifiers":{"arxiv":"2601.08510","doi":"10.48550/arXiv.2601.08510"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.08510"},{"label":"HTML","url":"https://arxiv.org/html/2601.08510v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.08510"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.08510v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Qiuyu Tian Yiding Li Fengyi Chen Zequn Liu Youyong Kong Fan Guo Yuyao Li Jinjing Shen Zhijing Xie Yiyun Luo Xin Zhang Southeast University, Nanjing, China Beijing Zhongguancun Academy, Beijing, China Nanjing Normal University, Nanjing, China ZhuiWen Technology Co., Ltd., Beijing, China qiuyutian@seu.edu.cn, liyiding@zhuiwen.net, fxc1494@g.rit.edu, liuzequn@bjzgca.edu.cn, kongyouyong@seu.edu.cn, guofan@zhuiwen.net, yuyaoli@zhuiwen.net, shenjinjing@zhuiwen.net, xiezhijing@zhuiwen.net, luoyiyun@zhuiwen.net, xinzhang@zhuiwen.net","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.08510v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-08510","updatedAt":"2026-08-26T02:54:30Z"},{"type":"preprint","title":"Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions","authors":[{"name":"Yongqi Li"},{"name":"Hao Lang"},{"name":"Tieyun Qian"},{"name":"Yongbin Li"}],"institutions":["zgca"],"rawAffiliations":["Tieyun Qian Affiliation: School of Computer Science, Wuhan University Affiliation: Zhongguancun Academy {liyongqi,qty}@whu.edu.cn , {hao.lang,shuide.lyb}@alibaba-inc.com"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.CL","cs.AI","cs.LG"],"abstract":"Vision-language models are increasingly employed as multimodal conversational agents (MCAs) for diverse conversational tasks. Recently, reinforcement learning (RL) has been widely explored for adapting MCAs to various human-AI interaction scenarios. Despite showing great enhancement in generalization performance, fine-tuning MCAs via RL still faces challenges in handling the extremely large text token space. To address this, we learn a compact latent action space for RL fine-tuning instead. Specifically, we adopt the learning from observation mechanism to construct the codebook for the latent action space, where future observations are leveraged to estimate current latent actions that could further be used to reconstruct future observations. However, the scarcity of paired image-text data hinders learning a codebook with sufficient coverage. Thus, we leverage both paired image-text data and text-only data to construct the latent action space, using a cross-modal projector for transforming text embeddings into image-text embeddings. We initialize the cross-modal projector on paired image-text data, and further train it on massive text-only data with a novel cycle consistency loss to enhance its robustness. We show that our latent action based method outperforms competitive baselines on two conversation tasks across various RL algorithms.","identifiers":{"arxiv":"2601.07516","doi":"10.48550/arXiv.2601.07516"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.07516"},{"label":"HTML","url":"https://arxiv.org/html/2601.07516v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.07516"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.07516v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Tieyun Qian Affiliation: School of Computer Science, Wuhan University Affiliation: Zhongguancun Academy {liyongqi,qty}@whu.edu.cn , {hao.lang,shuide.lyb}@alibaba-inc.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.07516v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-07516","updatedAt":"2026-08-26T02:54:30Z"},{"type":"preprint","title":"Controllable LLM Reasoning via Sparse Autoencoder-Based Steering","authors":[{"name":"Yi Fang"},{"name":"Wenjie Wang"},{"name":"Mingfeng Xue"},{"name":"Boyi Deng"},{"name":"Fengli Xu"},{"name":"Dayiheng Liu"},{"name":"Fuli Feng"}],"institutions":["zgca"],"rawAffiliations":["Yi Fang † † thanks: Work done when Yi Fang and Boyi Deng were interns at Alibaba Group. Affiliation: University of Science and Technology of China Affiliation: Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":["cs.AI","cs.CL"],"abstract":"Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (\\eg backtracking, cross-verification) during the reasoning process, which improves their performance on complex tasks. Currently, reasoning strategies are autonomously selected by LRMs themselves. However, such autonomous selection often produces inefficient or even erroneous reasoning paths. To make reasoning more reliable and flexible, it is important to develop methods for controlling reasoning strategies. Existing methods struggle to control fine-grained reasoning strategies due to conceptual entanglement in LRMs' hidden states. To address this, we leverage Sparse Autoencoders (SAEs) to decompose strategy-entangled hidden states into a disentangled feature space. To identify the few strategy-specific features from the vast pool of SAE features, we propose SAE-Steering, an efficient two-stage feature identification pipeline. SAE-Steering first recalls features that amplify the logits of strategy-specific keywords, filtering out over 99\\% of features, and then ranks the remaining features by their control effectiveness. Using the identified strategy-specific features as control vectors, SAE-Steering outperforms existing methods by over 15\\% in control effectiveness. Furthermore, controlling reasoning strategies can redirect LRMs from erroneous paths to correct ones, achieving a 7\\% absolute accuracy improvement. Our code and data are available at https://github.com/Peter-Fy/SAE-Steering.","identifiers":{"arxiv":"2601.03595","doi":"10.48550/arXiv.2601.03595"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2601.03595"},{"label":"HTML","url":"https://arxiv.org/html/2601.03595v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2601.03595"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2601.03595v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yi Fang † † thanks: Work done when Yi Fang and Boyi Deng were interns at Alibaba Group. Affiliation: University of Science and Technology of China Affiliation: Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2601.03595v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2601-03595","updatedAt":"2026-08-26T02:54:30Z"},{"type":"preprint","title":"World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation","authors":[{"name":"Ziyuan Jiang"},{"name":"Kai Liu"},{"name":"Yuxin Qin"},{"name":"Shuai Tian"},{"name":"Yupeng Zheng"},{"name":"Mingcai Zhou"},{"name":"Zhengtao Zhang"},{"name":"Chao Yu"},{"name":"Haoran Li"},{"name":"Dongbin Zhao"}],"institutions":["zgca"],"rawAffiliations":["Chao Yu Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Department of Electronic Engineering, Tsinghua University, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Robotic manipulation policies are commonly initialized through imitation learning, but their performance is limited by the scarcity and narrow coverage of expert data. Reinforcement learning can refine polices to alleviate this limitation, yet real-robot training is costly and unsafe, while training in simulators suffers from the sim-to-real gap. Recent advances in generative models have demonstrated remarkable capabilities in real-world simulation, with diffusion models in particular excelling at generation. This raises the question of how diffusion model-based world models can be combined to enhance pre-trained policies in robotic manipulation. In this work, we propose World4RL, a framework that employs diffusion-based world models as high-fidelity simulators to refine pre-trained policies entirely in imagined environments for robotic manipulation. Unlike prior works that primarily employ world models for planning, our framework enables direct end-to-end policy optimization. World4RL is designed around two principles: pre-training a diffusion world model that captures diverse dynamics on multi-task datasets and refining policies entirely within a frozen world model to avoid online real-world interactions. We further design a two-hot action encoding scheme tailored for robotic manipulation and adopt diffusion backbones to improve modeling fidelity. Extensive simulation and real-world experiments demonstrate that World4RL provides high-fidelity environment modeling and enables consistent policy refinement, yielding significantly higher success rates compared to imitation learning and other baselines.","identifiers":{"arxiv":"2509.19080","doi":"10.48550/arXiv.2509.19080","publishedDoi":"10.1109/lra.2026.3728345"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2509.19080"},{"label":"HTML","url":"https://arxiv.org/html/2509.19080v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2509.19080"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2509.19080v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Chao Yu Affiliation: Zhongguancun Academy, Beijing, China Affiliation: Department of Electronic Engineering, Tsinghua University, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2509.19080v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2509-19080","updatedAt":"2026-08-30T07:02:10Z"},{"type":"article","title":"Style2Risk: A Closed-Loop Testing System for Safety Evaluation of Automated Vehicles Under Style-Conditioned Adversarial Interactions","authors":[{"name":"Jintao Lai"},{"name":"Meng Wang","institutions":["zgca"]},{"name":"Zhen Zhang"},{"name":"Yiming Guo","institutions":["zgca"]},{"name":"Shixingyue Hu"},{"name":"Chengyuan Ma"},{"name":"Sa Gao","institutions":["zgca"]},{"name":"Sijin Liu","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Department of Control Science and Engineering, Tongji University, Shanghai 201804, China","Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Shanghai 201804, China","Beijing Zhongguancun Academy, Beijing 100094, China","Key Laboratory of Road and Traffic Engineering of the Ministry of Education, Tongji University, Shanghai 201804, China","Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China","College of Traffic & Transportation, Chongqing Jiaotong University, No. 66, Xuefu Ave., Chongqing 402247, China","Department of Civil and Environmental Engineering, University of Wisconsin–Madison, 1415 Engineering Drive, Madison, WI 53706, USA"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Systems","status":"published","topics":[],"abstract":"Closed-loop safety evaluation requires traffic agents that expose planner weaknesses without producing arbitrary behavior. Existing adversarial tests often couple behavior generation, risk search, and failure counting, limiting control over how interactions are created and whether repeated events constitute new evidence. Style2Risk uses interaction style as an explicit control interface. Assertiveness and courtesy condition a Gaussian behavior prior, while a bounded residual directs the controlled participant toward the tested planner. Protocol screens retain interactions that satisfy declared action and trajectory constraints, and a style–interaction–event archive separates discovery breadth from repeated observations. We evaluate seven configurations and three planners in 31,500 WOMD–Waymax rollouts. Relative to strict no-style, Style2Risk increases protocol-valid safety-critical events from 110.0 to 258.0 per 500-scenario planner chunk and expands occupied archive cells from 28 to 33. A common trajectory evaluator that does not call the generating prior reports a 1.343 m reduction in joint displacement error, a 0.163 increase in interaction consistency, and a 0.037 reduction in collision rate. A separate 26-scenario experiment assigns style from pre-rollout information and evaluates online archive guidance. The results support style-conditioned control as a practical mechanism for targeted closed-loop testing and explicit evidence accounting.","identifiers":{"doi":"10.3390/systems14091144"},"links":[{"label":"DOI","url":"https://doi.org/10.3390/systems14091144"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.3390/systems14091144"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beijing Zhongguancun Academy, Beijing 100094, China","source":"Crossref","sourceUrl":"https://doi.org/10.3390/systems14091144"}],"sources":["Crossref"],"id":"doi-10-3390-systems14091144","updatedAt":"2026-09-16T05:02:31Z"},{"type":"preprint","title":"Bridging the Training–Application Gap in Kohn–Sham Hamiltonian Learning through Dual-Space Supervision","authors":[{"name":"Yifei Wang"},{"name":"He Zhang","institutions":["zgca","zgci"]},{"name":"Xinran Wei","institutions":["zgca","zgci"]},{"name":"Jingwen Fu","institutions":["zgca","zgci"]},{"name":"Yongqiang Ma"},{"name":"Miao Kang"},{"name":"Jianji Wang"},{"name":"Chang Liu","institutions":["zgca","zgci"]},{"name":"Nanning Zheng"}],"institutions":["zgca","zgci"],"rawAffiliations":["State Key Laboratory of Human-Machine Hybrid Augmented Intelligence","Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University","Zhongguancun Academy","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Predicting the Kohn–Sham Hamiltonian in density functional theory (DFT) via machine learning offers a promising pathway toward efficient electronic structure modeling. Existing models are typically trained by directly minimizing element-wise Hamiltonian matrix errors, yet a significant trainingapplication gap has been observed between Hamiltonian-space loss and errors in derived electronic properties, limiting their utility in practical applications. Noting that these properties are functions on the state space obtained by solving the Hamiltonian eigenvalue problem, this gap can be traced to a misalignment of the error metrics on the Hamiltonian space and its induced state space. The operator-tostate mapping is highly anisotropic, such that Hamiltonian errors of comparable numerical magnitude can induce markedly different downstream property errors. Motivated by the Hamiltonian–state variational duality rooted in Lieb’s convex formulation of DFT, we introduce Dual-Space Loss (DSLoss) to bridge this training–application gap. By leveraging error metrics in both spaces, DSLoss allows Hamiltonian optimization to account for errors in both the predicted operator and its induced state. Across molecular benchmarks, DSLoss consistently improves a broad range of derived properties over conventional element-wise supervision. On the challenging ∇2DFT benchmark, electron density and total energy errors are reduced by factors of 40 and 500, respectively, relative to the strongest competing method. The resulting model also extrapolates robustly to molecules containing up to twice as many heavy atoms as those encountered during training, achieving more than an order-of-magnitude lower errors for multiple properties and substantially outperforming direct density prediction models and endto-end property predictors. Moreover, the model enables stable molecular dynamics simulations with accurate energetics, forces, and vibrational dynamics. These results show that DSLoss addresses an unrecognized bottleneck limiting operator-prediction models across diverse downstream tasks, thereby advancing scalable and reliable electronic structure modeling.","identifiers":{"doi":"10.26434/chemrxiv.15009467/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.26434/chemrxiv.15009467/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.26434/chemrxiv.15009467/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.26434/chemrxiv.15009467/v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.26434/chemrxiv.15009467/v1"}],"sources":["Crossref"],"id":"doi-10-26434-chemrxiv-15009467-v1","updatedAt":"2026-09-27T05:30:19Z"},{"type":"preprint","title":"Accelerated Stable Structure Prediction of Li-Intercalated Bilayer Graphene Using a Data-Efficient Deep Learning Framework","authors":[{"name":"Haoran Li"},{"name":"Bilin Gui"},{"name":"He Zhang","institutions":["zgca","zgci"]},{"name":"Le Yang"},{"name":"Hao-Sen Chen"}],"institutions":["zgca","zgci"],"rawAffiliations":["Institute of Advanced Structure Technology","Beijing Institute of Technology","Zhongguancun Academy","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Replacing bulk graphite with few-layer graphene is a promising route toward improving both the storage capacity and rate capability of fast-charging Li-ion batteries. As the simplest layered graphene system, bilayer graphene (BLG) provides an ideal model for uncovering the structural origins of these improvements, which in turn requires accurate identification of its thermodynamically stable structures. However, predicting stable Li-intercalated structures requires exhaustive exploration of a vast configurational space using density functional theory (DFT), which rapidly becomes computationally prohibitive for large supercells. Here, we present DESSP, a data-efficient deep learning framework for accelerated stable structure prediction in Liintercalated BLG. DESSP combines a genetic-algorithm-based search pipeline with a universal machine-learning interatomic potential (MLIP) used as a surrogate for DFT, enabling broad yet efficient exploration of the potential energy surface. Representative structures sampled along the search trajectories are then selectively labeled with DFT to train a high-fidelity MLIP, while a distillation-based strategy further improves predictive accuracy. The resulting model achieves near-DFT accuracy, generalizes effectively to larger supercells, and reproduces the thermodynamic convex hull at substantially lower computational cost, providing a scalable framework for studying nanoscale layered electrodes in batteries.","identifiers":{"doi":"10.26434/chemrxiv.15005431/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.26434/chemrxiv.15005431/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.26434/chemrxiv.15005431/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.26434/chemrxiv.15005431/v1"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.26434/chemrxiv.15005431/v1"}],"sources":["Crossref"],"id":"doi-10-26434-chemrxiv-15005431-v1","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Adaptive TD-Lambda for Cooperative Multi-agent Reinforcement Learning","authors":[{"name":"Yue Deng","institutions":["zgca"]},{"name":"Zirui Wang"},{"name":"Yin Zhang"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy","Zhejiang University","Yue Deng Affiliation: Zhongguancun Academy, Beijing, China Email: dengyue@zgci.ac.cn"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence","status":"published","topics":[],"abstract":"TD(λ) in value-based MARL algorithms or the Temporal Difference critic learning in Actor-Critic-based (AC-based) algorithms synergistically integrate elements from Monte-Carlo simulation and Q function bootstrapping via dynamic programming, which effectively addresses the inherent bias-variance trade-off in value estimation. Based on that, some recent works link the adaptive λ value to the policy distribution in the single-agent reinforcement learning area. However, because of the large joint action space from multiple agents and the limited transition data in Multi-agent Reinforcement Learning, the policy distribution is infeasible to calculate statistically. To solve the policy distribution calculation problem in MARL settings, we employ a parametric likelihood-free density ratio estimator with two replay buffers instead of calculating statistically. The two replay buffers of different sizes respectively store the historical trajectories that represent the data distribution of the past and current policies. Based on the estimator, we assign Adaptive TD(λ), ATD(λ), values to state-action pairs based on their likelihood under the stationary distribution of the current policy. We apply the proposed method on two competitive baseline methods, QMIX for value-based algorithms, and MAPPO for AC-based algorithms, over SMAC benchmarks and Gfootball academy scenarios, and demonstrate consistently competitive or superior performance compared to other baseline approaches with static λ values.","identifiers":{"doi":"10.24963/ijcai.2026/10"},"links":[{"label":"DOI","url":"https://doi.org/10.24963/ijcai.2026/10"},{"label":"arXiv","url":"https://arxiv.org/abs/2605.11880"},{"label":"HTML","url":"https://arxiv.org/html/2605.11880v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2605.11880"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.24963/ijcai.2026/10"},{"label":"arXiv v1","url":"https://arxiv.org/abs/2605.11880v1"},{"label":"DOI 10.48550/arXiv.2605.11880","url":"https://doi.org/10.48550/arXiv.2605.11880"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.24963/ijcai.2026/10"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Yue Deng Affiliation: Zhongguancun Academy, Beijing, China Email: dengyue@zgci.ac.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2605.11880v1"}],"sources":["Crossref","arXiv","arXiv HTML"],"id":"doi-10-24963-ijcai-2026-10","updatedAt":"2026-09-17T05:05:28Z"},{"type":"preprint","title":"Reconstructing sequence-grammar trajectories enables interpretable and tunable cis-regulatory element design","authors":[{"name":"Mingqian Ma"},{"name":"Wanjuan Bu"},{"name":"Guoqing Liu"},{"name":"Yuxuan Liu"},{"name":"Sizhen Liu"},{"name":"Zhen Zhao"},{"name":"Shijie Yao"},{"name":"Qingru Hua"},{"name":"Yujie Zhang"},{"name":"Cuiting Zhong"},{"name":"Haitao Huang"},{"name":"Pan Deng","institutions":["zgca"]},{"name":"Peiran Jin","institutions":["zgca"]},{"name":"Qijin Yin","institutions":["zgca"]},{"name":"Chuan Cao","institutions":["zgca"]},{"name":"Haiguang Liu","institutions":["zgca"]},{"name":"Mo Xu"},{"name":"Yuan He","institutions":["zgca"]},{"name":"Tao Qin","institutions":["zgca"]},{"name":"Zeyu Chen"}],"institutions":["zgca"],"rawAffiliations":["Carnegie Mellon University","Peking University","Microsoft Research (United Kingdom)","National Institute of Biological Sciences, Beijing","University of Rochester","Hunan University","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Background Designing cis -regulatory elements (CREs) with cell-type-specific activity is critical for precise gene and cell therapies. Genomic language models (gLMs) have shown potential for CRE generation, but their optimization is often treated as a black box, limiting understanding of how regulatory grammar emerges and how unsuccessful generation trajectories can be corrected. Results We present Guided Optimization of C is -Regulatory Elements (GO-CRE), an interpretable deep learning framework for cell-type-specific CRE generation. GO-CRE is built on HybriDNA, a hybrid Transformer-Mamba2 gLM that supports iterative sequence optimization under practical computational constraints. GO-CRE then runs reinforcement learning (RL) and reconstructs sequence-grammar trajectories using deconvolved k-mer and motif features. The trajectories exhibit three phases - search, commitment, and optimization - each associated with the acquisition of distinct biological features. In HepG2, trajectory analysis identified a low-complexity polyG trap that motivated an updated reinforcement learning policy, and redirected optimization toward higher activity and specificity. In HepG2 and K562, GO-CRE generated diverse, cell-type-specific CREs with compact lineage-associated grammars, including HNF1B-FOXA1-HNF4A programs in HepG2 and GATA/RUNX-associated programs in K562. Lentiviral massively parallel reporter assays validated the cell-type-specific activity of the generated CREs in both cell types, and showed that HepG2 designs had higher average activity than endogenous CREs. Conclusions GO-CRE integrates efficient gLM-based generation, sequence grammar trajectory reconstruction, and biologically guided reward shaping. It links iterative sequence changes to regulatory grammar and feeds interpretable features back into reward design, and thus enables the design of diverse, experimentally validated, cell-type-specific CREs.","identifiers":{"doi":"10.21203/rs.3.rs-10913378/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.21203/rs.3.rs-10913378/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.21203/rs.3.rs-10913378/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.21203/rs.3.rs-10913378/v1"}],"sources":["Crossref"],"id":"doi-10-21203-rs-3-rs-10913378-v1","updatedAt":"2026-09-24T05:08:05Z"},{"type":"preprint","title":"Chemical Disorder Suppresses Thermal Transport through Local Stress Redistribution in SrTiO3-Based High-Entropy Perovskites","authors":[{"name":"Yongheng Li","institutions":["zgci"]},{"name":"Maomao Liu"},{"name":"Hongwei Du","institutions":["zgci"]},{"name":"Baole Wei","institutions":["zgci"]},{"name":"Bin Wei"},{"name":"Bonan Zhu"},{"name":"Ziheng Lu","institutions":["zgci"]},{"name":"Yuanhua Lin"}],"institutions":["zgci"],"rawAffiliations":["Zhongguancun Institute of Artificial Intelligence","Henan Polytechnic University","Beijing Institute of Technology","Tsinghua University"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract High-entropy disorder offers an effective route to suppress lattice heat transport in oxide thermoelectrics, yet the atomistic origin of phonon scattering in high-entropy SrTiO 3 remains unresolved. Here, we distill a fine-tuned MACE teacher into an accurate and efficient NEP for large-scale molecular dynamics. Using this potential, we find that A-site substitution rapidly lowers lattice thermal conductivity, while further reductions are modest in the four-and five-component compositions near 3 W m −1 K −1 . Phonon spectra reveal two distinct dopant effects: Pb introduces a low-frequency soft-mode-like response, whereas Ca produces stronger spectral broadening. Mass fluctuation and ionic-size mismatch alone do not describe the conductivity trend well. Instead, the local stress descriptor correlates more directly with thermal conductivity and captures the loss of phonon spectral coherence, linking local stress disorder to effective lifetime-like broadening. These results establish local stress disorder as a transport descriptor for low thermal conductivity and show that MACE-guided NEP distillation provides a practical route to large-scale atomistic studies of high-entropy oxides.","identifiers":{"doi":"10.21203/rs.3.rs-10754513/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.21203/rs.3.rs-10754513/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.21203/rs.3.rs-10754513/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.21203/rs.3.rs-10754513/v1"}],"sources":["Crossref"],"id":"doi-10-21203-rs-3-rs-10754513-v1","updatedAt":"2026-09-08T04:58:29Z"},{"type":"article","title":"SwipeWell: A Multi-Agent System for Short-Video-Based Mobile Psychological Self-Assessment among Older Adults","authors":[{"name":"Chenyu Gu","institutions":["zgca"]},{"name":"Yuxiao Sun","institutions":["zgca"]},{"name":"Zhilong Chen"},{"name":"Zhimin Wang"},{"name":"Yong Li"},{"name":"Kai Chen","institutions":["zgci"]},{"name":"Feng Lu"},{"name":"Yaojing Chen"},{"name":"Yuanyi Zhen","institutions":["zgca","zgci"]}],"institutions":["zgca","zgci"],"rawAffiliations":["State Key Lab. of VR System and Technology, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","Beijing Normal University, Beijing, China and Zhongguancun Academy, Beijing, China","Department of Electronic Engineering, Tsinghua University, Beijing, China","Key Laboratory of Social Computing and Cognitive Intelligence, Dalian University of Technology, Dalian, China","Zhongguancun Institute of Artificial Intelligence, Beijing, China","State Key Lab. of VR System and Technology, Beihang University, Beijing, China","Beijing Normal University, Beijing, China","Zhongguancun Academy, Beijing, China and Zhongguancun Institute of Artificial Intelligence, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies","status":"published","topics":[],"abstract":"Psychological self-assessment is fundamental to monitoring the mental well-being of the aging population. However, conventional paper-based methods and emerging LLM-driven conversational interfaces often suffer from low user engagement, high cognitive load, and usability barriers. This paper presents SwipeWell, a novel mobile psychological self-assessment system that leverages a human-in-the-loop multi-agent workflow to translate validated scales into psychometrically grounded and content-faithful animations, allowing users to intuitively log their status via sidebar interactions. We evaluated SwipeWell through a within-subjects study ( N = 27) with older adults, comparing it against traditional paper-and-pencil and LLM-based conversational assessments using three validated psychological scales covering emotion, cognition, and somatization. Psychometrically, SwipeWell maintained reliable and valid measurement performance. It established rank-order consistency for emotion and cognition, while facilitating somatic symptom interpretation through multimodal representations. Our empirical results demonstrate that SwipeWell is significantly more engaging, enjoyable, and time-efficient. Furthermore, participants experienced a significantly lower cognitive load and reported higher learnability compared to conversational systems. These findings highlight how embedding psychological self-assessment tasks into familiar, low-friction mobile interactions can foster accessible mobile health tools, providing design implications for future pervasive health technologies for older adults.","identifiers":{"doi":"10.1145/3832031"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3832031"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3832031"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"State Key Lab. of VR System and Technology, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3832031"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3832031"}],"sources":["Crossref"],"id":"doi-10-1145-3832031","updatedAt":"2026-10-01T06:16:42Z"},{"type":"article","title":"VectorWeaver: A Stage-Aware Automated Optimization Framework for LLM Inference on Edge Platform","authors":[{"name":"Pengfei Yang"},{"name":"Hui Zeng","institutions":["zgca"]},{"name":"Wenxuan Hou"},{"name":"Mingwei Wang"},{"name":"Weiye Ji"},{"name":"Tianyang Zheng"},{"name":"Hui Li"}],"institutions":["zgca"],"rawAffiliations":["Xidian University","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"ACM Transactions on Architecture and Code Optimization","status":"published","topics":[],"abstract":"The convergence of Large Language Models (LLMs) and the open-source RISC-V architecture creates significant opportunities for specialized AI hardware. However, achieving optimal LLM inference performance on RISC-V platforms is challenging due to the complex computational patterns of LLM operators and the architectural nuances of the RISC-V Vector (RVV) extension. Existing optimization strategies often rely on manual, expert-driven tuning, which is labor-intensive and not generalizable across diverse operators and hardware configurations. This article presents VectorWeaver, an automated framework to generate high-performance, stage-aware microkernels for LLM inference on RISC-V. VectorWeaver specifically addresses key performance bottlenecks in quantized models, such as memory discontinuities in vector dot-product operations, by employing a widen vectorization strategy. Guided by a stage-aware performance model, the framework systematically explores a rich optimization space encompassing loop unrolling, software pipelining, and mathematical approximations. The model generates distinct, highly-optimized kernels for the prefill and decode stages, selected at runtime via a dispatch mechanism. We evaluate VectorWeaver on the TH1520 and the Sophon SG2044 using Gemma3 and Qwen3 models with sizes ranging from 270M to 4B, where our generated microkernels achieve up to a 24x speedup over baseline implementations. For end-to-end evaluations, VectorWeaver delivers up to a 5.2x increase in inference throughput, demonstrating its effectiveness in automating the generation of high-performance code and advancing the state of LLM inference on the RISC-V ecosystem.","identifiers":{"doi":"10.1145/3799719"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3799719"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3799719"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3799719"}],"sources":["Crossref"],"id":"doi-10-1145-3799719","updatedAt":"2026-08-12T03:04:08Z"},{"type":"conference","title":"LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts","authors":[{"name":"Hao Liang","institutions":["zgca"]},{"name":"Qifeng Cai"},{"name":"Zhaoyang Han"},{"name":"Hejun Dong"},{"name":"Meiyi Qiang"},{"name":"Ruichuan An"},{"name":"Quanqing Xu"},{"name":"Bin Cui"},{"name":"Wentao Zhang","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Peking University, Beijing, China and Zhongguancun Academy, Beijing, China","East China Normal University, Shanghai, China","Huazhong University of Science and Technology, Wuhan, China","Beihang University, Beijing, China","Peking University, Beijing, China","OceanBase, Ant Group, Beijing, China","Peking University, Beijing, China and Zhongguancun Academy, Beijing, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the ACM Web Conference 2026","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3774904.3792099"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3774904.3792099"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3774904.3792099"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Peking University, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3774904.3792099"}],"sources":["Crossref"],"id":"doi-10-1145-3774904-3792099","updatedAt":"2026-09-27T05:30:19Z"},{"type":"conference","title":"PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers","authors":[{"name":"Linhao Zhang","institutions":["zgca"]},{"name":"Tong Xia"},{"name":"Jinghua Piao","institutions":["zgca"]},{"name":"Lizhen Cui"}],"institutions":["zgca"],"rawAffiliations":["School of Software, Shandong University, Jinan, China and Zhongguancun Academy, Beijing, China","Vanke School of Public Health, Tsinghua University, Beijing, China","Department of Electronic Engineering, Tsinghua University, Beijing, China and Zhongguancun Academy, Beijing, China","School of Software, Shandong University, Jinan, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770855.3818943"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3818943"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3818943"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"School of Software, Shandong University, Jinan, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3818943"}],"sources":["Crossref"],"id":"doi-10-1145-3770855-3818943","updatedAt":"2026-08-12T03:04:08Z"},{"type":"conference","title":"<scp>Caduceus:</scp> MoE Foundation Models for Unifying Biological and Natural Language","authors":[{"name":"Mingze Yin"},{"name":"Yiheng Zhu","institutions":["zgca"]},{"name":"Jialu Wu"},{"name":"Jian Ma","institutions":["zgca"]},{"name":"Hanjing Zhou"},{"name":"Mingyang Li"},{"name":"Yuhua Zhou"},{"name":"Jintai Chen"},{"name":"Tingjun Hou"},{"name":"Jieping Ye"},{"name":"Aimin Pan"}],"institutions":["zgca"],"rawAffiliations":["College of Computer Science and Technology, Zhejiang University, Hangzhou, China","Zhongguancun Academy, Beijing, China","College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, China","Cainiao Group, Alibaba Group, Hangzhou, China","Tongyi Lab, Beijing, China","AI Thrust, Information Hub, HKUST (GZ), Guangzhou, China","Zhejiang Lab, Hangzhou, China and VNET Group, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770855.3818885"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3818885"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3818885"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3818885"}],"sources":["Crossref"],"id":"doi-10-1145-3770855-3818885","updatedAt":"2026-08-13T13:57:27Z"},{"type":"conference","title":"Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction","authors":[{"name":"Yi Duan"},{"name":"Zhao Yang"},{"name":"Jiwei Zhu"},{"name":"Ying Ba"},{"name":"Chuan Cao","institutions":["zgca"]},{"name":"Bing Su"}],"institutions":["zgca"],"rawAffiliations":["Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China","Zhongguancun Academy, Beijing, China","Yi Duan Note: Yi Duan and Zhao Yang contributed equally to this work. This work was done while they were visiting Zhongguancun Academy. email: 2023200660@ruc.edu.cn Affiliation: Gaoling School of Artificial Intelligence , Renmin University of China , Beijing , China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"DNA cis-regulatory elements (CREs) such as enhancers control gene expression levels. Accurately predicting regulatory activity from DNA sequences is valuable but challenging, as it requires understanding complex biological regulatory processes. Existing methods typically regress activity scores from sequences in a black-box manner, limiting both interpretability and regression performance. Meanwhile, large language models (LLMs) benefit from explicit reasoning processes, yet directly applying LLMs to raw DNA sequences performs poorly. In this paper, we bridge this gap by introducing R3LM, a framework that teaches LLMs reasoning-informed regression on regulatory DNA through structured biological knowledge. Specifically, we design a biologically grounded data format that structures DNA's regulatory information for improved LLM understanding, and construct CRE-ReasonBench, the first dataset that associates DNA sequences and activity scores with mechanistic reasoning traces. Through two-stage training that first teaches LLMs reasoning over structured biological information then performs regression, R3LM achieves state-of-the-art performance on enhancer prediction across three cell types, outperforming both LLMs with raw sequence input and specialized DNA models while providing interpretable mechanistic explanations. We expect R3LM as an interpretable reward model that can effectively assist biologists in CRE design. Code is available at https://github.com/DuanYi516/R3LM.","identifiers":{"doi":"10.1145/3770855.3818836"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3818836"},{"label":"arXiv","url":"https://arxiv.org/abs/2606.08147"},{"label":"HTML","url":"https://arxiv.org/html/2606.08147v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2606.08147"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3818836"},{"label":"arXiv v1","url":"https://arxiv.org/abs/2606.08147v1"},{"label":"DOI 10.48550/arXiv.2606.08147","url":"https://doi.org/10.48550/arXiv.2606.08147"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3818836"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Yi Duan Note: Yi Duan and Zhao Yang contributed equally to this work. This work was done while they were visiting Zhongguancun Academy. email: 2023200660@ruc.edu.cn Affiliation: Gaoling School of Artificial Intelligence , Renmin University of China , Beijing , China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2606.08147v1"}],"sources":["Crossref","arXiv","arXiv HTML"],"id":"doi-10-1145-3770855-3818836","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"ResDIF: A Residual Disentanglement Framework for Interpretable Financial Time Series Forecasting via Spectrally-Enhanced Temporal Encoding","authors":[{"name":"Chengwei Fu","institutions":["zgca"]},{"name":"Gang Xiao"},{"name":"Yuchao Zhang","institutions":["zgca"]},{"name":"Yuhang Sun","institutions":["zgca"]},{"name":"Jiange Li"},{"name":"Yue Deng","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Beijing University of Aeronautics and Astronautics, Beijing, China and Zhongguancun Academy, Beijing, China","China Securities Co.,Ltd., Beijing, China","Beijing University of Aeronautics and Astronautics, Beijing, China and Zhongguancun Academy, Beijing, China, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770855.3818080"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3818080"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3818080"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beijing University of Aeronautics and Astronautics, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3818080"}],"sources":["Crossref"],"id":"doi-10-1145-3770855-3818080","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification","authors":[{"name":"Wenyao Cui","institutions":["zgca"]},{"name":"Huaping Zhang"},{"name":"Yongyi Huang"},{"name":"Qiuchi Li"},{"name":"Jian Xu","institutions":["zgca"]},{"name":"Cheng-Lin Liu","institutions":["zgca"]},{"name":"Chunxiao Gao"},{"name":"Juan Wang"},{"name":"Baohua Zhang"}],"institutions":["zgca"],"rawAffiliations":["Beijing Institute of Technology, Beijing, China and Zhongguancun Academy, Beijing, China","Xinjiang Future Enterprise Incubator Co., Ltd., Xinjiang, China and Beijing Institute of Technology, Beijing, China","Beijing Institute of Technology, Beijing, China","Zhongguancun Academy, Beijing, China and Institute of Automation, Chinese Academy of Sciencess, Beijing, China","Wenyao Cui OrcID: 0000-0002-2810-3824 Affiliation: Beijing Institute of Technology , Beijing , China Affiliation: Zhongguancun Academy , Beijing , China email: yao1970099540@gmail.com"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ''verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable critiques, and scalar rewards (e.g., PRMs/RMs) offer little insight into where a multi-step derivation fails.We propose SymDiag, a neuro-symbolic framework that reframes reasoning verification as structured failure diagnosis. SymDiag translates natural-language CoT into symbolic constraints and performs step-level satisfiability/entailment checks to (i) localize failing steps and (ii) produce verifiable diagnostic evidence, including counterexamples, inconsistency witnesses, and missing-premise indicators. A central challenge is that apparent ''logic violations'' can be caused either by genuine reasoning defects or by neural-to-symbolic translation noise. SymDiag therefore incorporates a Self-Auditor that disentangles TranslationError from ReasoningError via dual symbolic encodings consistency checks, enabling robust diagnosis under partial observability. Across diverse mathematical, logical, scientific, and general reasoning benchmarks, SymDiag improves detection of unfaithful reasoning and provides substantially more effective feedback for multi-round reasoning repair than outcome-only verification and LLM-based judging, offering a principled foundation for trustworthy and scalable reasoning diagnosis.","identifiers":{"doi":"10.1145/3770855.3818004"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3818004"},{"label":"arXiv","url":"https://arxiv.org/abs/2608.08786"},{"label":"HTML","url":"https://arxiv.org/html/2608.08786v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2608.08786"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3818004"},{"label":"arXiv v1","url":"https://arxiv.org/abs/2608.08786v1"},{"label":"DOI 10.48550/arXiv.2608.08786","url":"https://doi.org/10.48550/arXiv.2608.08786"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Beijing Institute of Technology, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3818004"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Wenyao Cui OrcID: 0000-0002-2810-3824 Affiliation: Beijing Institute of Technology , Beijing , China Affiliation: Zhongguancun Academy , Beijing , China email: yao1970099540@gmail.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2608.08786v1"}],"sources":["Crossref","arXiv","arXiv HTML"],"id":"doi-10-1145-3770855-3818004","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM","authors":[{"name":"Mengjie Liu"},{"name":"Jiahui Peng"},{"name":"Wenchang Ning"},{"name":"Pei Chu"},{"name":"Jiantao Qiu"},{"name":"Ren Ma"},{"name":"He Zhu"},{"name":"Rui Min"},{"name":"Lindong Lu"},{"name":"Linfeng Hou"},{"name":"Kaiwen Liu"},{"name":"Yuan Qu"},{"name":"Zhenxiang Li"},{"name":"Chao Xu"},{"name":"Zhongying Tu"},{"name":"Wentao Zhang","institutions":["zgca"]},{"name":"Conghui He"}],"institutions":["zgca"],"rawAffiliations":["Shanghai Artificial Intelligence Laboratory, Shanghai, China and Peking University, Beijing, China","Shanghai Artificial Intelligence Laboratory, Shanghai, China","Shanghai Artificial Intelligence Laboratory, Shanghai, Shanghai, China","Shanghai Artificial Intelligence Laboratory, Shanghai, Shanghai, China and Peking University, Beijing, Beijing, China","Shanghai Artificial Intelligence Laboratory, Shanghai, Shanghai, China, Peking University, Beijing, Beijing, China, Zhongguancun Academy, Beijing, Beijing, China, and Beijing Key Laboratory of Data Intelligence and Security (Peking University), Beijing, Beijing, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770855.3817915"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770855.3817915"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770855.3817915"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Shanghai Artificial Intelligence Laboratory, Shanghai, Shanghai, China, Peking University, Beijing, Beijing, China, Zhongguancun Academy, Beijing, Beijing, China, and Beijing Key Laboratory of Data Intelligence and Security (Peking University), Beijing, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770855.3817915"}],"sources":["Crossref"],"id":"doi-10-1145-3770855-3817915","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Generating Realistic Human Mobility Data with Hybrid Large Language Model Agent","authors":[{"name":"Chenyang Shao","institutions":["zgca"]},{"name":"Bingbing Fan"},{"name":"Jingtao Ding"},{"name":"Yuan Yuan"},{"name":"Meng Wang"},{"name":"Fengli Xu","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China and Zhongguancun Academy, Beijing, China","Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China","School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770854.3785685"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770854.3785685"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770854.3785685"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770854.3785685"}],"sources":["Crossref"],"id":"doi-10-1145-3770854-3785685","updatedAt":"2026-09-02T04:51:31Z"},{"type":"conference","title":"Tokenizing 3D Molecule Structure with Quantized Spherical Coordinates","authors":[{"name":"Kaiyuan Gao"},{"name":"Yusong Wang"},{"name":"Haoxiang Guan","institutions":["zgca"]},{"name":"Zun Wang"},{"name":"Qizhi Pei"},{"name":"John Hopcroft"},{"name":"Kun He"},{"name":"Lijun Wu"}],"institutions":["zgca"],"rawAffiliations":["Huazhong University of Science and Technology, Wuhan, China","Xi'an Jiaotong University, Xi'an, China","University of Science and Technology of China, Hefei, China and Zhongguancun Academy, Beijing, China","Shanghai Artificial Intelligence Laboratory, Beijing, China","Renmin University of China, Beijing, China and Shanghai Artificial Intelligence Laboratory, Shanghai, China","Cornell University, Ithaca, USA","Shanghai Artificial Intelligence Laboratory, Shanghai, China and Shanghai Innovation institute, Shanghai, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3770854.3780301"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3770854.3780301"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3770854.3780301"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"University of Science and Technology of China, Hefei, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3770854.3780301"}],"sources":["Crossref"],"id":"doi-10-1145-3770854-3780301","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"Multimodal Aspect-Based Sentiment Analysis With Plugin-Enhanced Large Language Models","authors":[{"name":"Yuanhe Tian","institutions":["zgci"]},{"name":"Yan Song"},{"name":"Yongdong Zhang"}],"institutions":["zgci"],"rawAffiliations":["Zhongguancun Institute of Artificial Intelligence, Beijing, China","University of Science and Technology of China, Hefei, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"IEEE Transactions on Neural Networks and Learning Systems","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/tnnls.2025.3622470"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/tnnls.2025.3622470"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/tnnls.2025.3622470"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1109/tnnls.2025.3622470"}],"sources":["Crossref"],"id":"doi-10-1109-tnnls-2025-3622470","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"Reinforced Context Augmentation for Multimodal Emotion Analysis","authors":[{"name":"Ruyi Gan"},{"name":"Yuanhe Tian","institutions":["zgci"]},{"name":"Kunhao Pan"},{"name":"Yan Song"},{"name":"Yongdong Zhang"}],"institutions":["zgci"],"rawAffiliations":["School of Information Science and Technology, University of Science and Technology of China","Zhongguancun Institute of Artificial Intelligence","International Digital Economics Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"IEEE Transactions on Multimedia","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/tmm.2026.3673543"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/tmm.2026.3673543"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/tmm.2026.3673543"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.1109/tmm.2026.3673543"}],"sources":["Crossref"],"id":"doi-10-1109-tmm-2026-3673543","updatedAt":"2026-09-27T05:30:19Z"},{"type":"article","title":"Feature Decomposition via Shared Low-Rank Matrix Recovery for CT Report Generation","authors":[{"name":"Yuanhe Tian","institutions":["zgci"]},{"name":"Yan Song"}],"institutions":["zgci"],"rawAffiliations":["Zhongguancun Institute of Artificial Intelligence, Beijing, China","University of Science and Technology of China, Hefei, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"IEEE Transactions on Medical Imaging","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/tmi.2025.3628159"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/tmi.2025.3628159"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/tmi.2025.3628159"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1109/tmi.2025.3628159"}],"sources":["Crossref"],"id":"doi-10-1109-tmi-2025-3628159","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"A Physical-Informed Super-Resolution Method for Electric Camera Inspired by Active Electric Sense","authors":[{"name":"Penghang Shuai"},{"name":"Zhanhua Xin"},{"name":"Xingyu Chen","institutions":["zgca"]},{"name":"Shihan Kong"},{"name":"Junzhi Yu"}],"institutions":["zgca"],"rawAffiliations":["Peking University","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"IEEE Transactions on Industrial Electronics","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/tie.2026.3720710"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/tie.2026.3720710"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/tie.2026.3720710"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.1109/tie.2026.3720710"}],"sources":["Crossref"],"id":"doi-10-1109-tie-2026-3720710","updatedAt":"2026-08-25T01:59:30Z"},{"type":"article","title":"Hybrid-CMLP: Hybrid CNN-MLP Networks for Low-to-standard-dose PET Synthesis","authors":[{"name":"Yuxin Xue"},{"name":"Mingyuan Meng","institutions":["zgca","zgci"]},{"name":"Lei Bi"},{"name":"Jinman Kim"}],"institutions":["zgca","zgci"],"rawAffiliations":["School of Computer Science, University of Sydney, NSW, Australia","Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China","Institute of Translational Medicine, National Center for Translational Medicine, Shanghai Jiao Tong University, Shanghai, China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"IEEE Journal of Biomedical and Health Informatics","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/jbhi.2026.3701924"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/jbhi.2026.3701924"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/jbhi.2026.3701924"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1109/jbhi.2026.3701924"},{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1109/jbhi.2026.3701924"}],"sources":["Crossref"],"id":"doi-10-1109-jbhi-2026-3701924","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models","authors":[{"name":"Dingkun Liu","institutions":["zgca"]},{"name":"Yuheng Chen"},{"name":"Zhu Chen"},{"name":"Zhenyao Cui"},{"name":"Yaozhi Wen","institutions":["zgca"]},{"name":"Jiayu An"},{"name":"Jingwei Luo"},{"name":"Dongrui Wu","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology , Wuhan, 430074","Zhongguancun Academy , Beijing, 100094 ,","Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology 1 , Wuhan, 430074","State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences , Beijing, 100190 ,"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"National Science Review","status":"published","topics":[],"abstract":"Abstract Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to learn transferable neural representations from large-scale heterogeneous recordings. Despite rapid progress, a fair and comprehensive comparison of existing EEG FMs is still lacking, owing to inconsistent pre-training objectives, preprocessing choices, and downstream evaluation protocols. To fill this gap, we present EEG-FM-Compass. We first review 55 representative models and organize their design choices into a unified taxonomic framework including data standardization, model architectures, and self-supervised pre-training strategies. We then evaluate 12 open source FMs and competitive specialist baselines across 13 EEG datasets spanning nine brain-computer interface paradigms. Emphasizing real-world deployments, we consider both cross-subject generalization under a leave-one-subject-out protocol and rapid calibration under a within-subject few-shot setting. We further compare full-parameter fine-tuning with linear probing to assess the transferability of pre-trained representations, and examine the relationship between model scale and downstream performance. Our results indicate that: 1) linear probing is frequently insufficient; 2) specialist models trained from scratch remain competitive across many tasks; and 3) larger FMs do not necessarily yield better generalization performance under current data regimes and training practices.","identifiers":{"doi":"10.1093/nsr/nwag466"},"links":[{"label":"DOI","url":"https://doi.org/10.1093/nsr/nwag466"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1093/nsr/nwag466"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy , Beijing, 100094 ,","source":"Crossref","sourceUrl":"https://doi.org/10.1093/nsr/nwag466"}],"sources":["Crossref"],"id":"doi-10-1093-nsr-nwag466","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"DECANT: Decoupling mechanism from context in single-cell drug perturbation representation","authors":[{"name":"Ren Qi"},{"name":"Wenjie Teng"},{"name":"Xin Yang"},{"name":"Yue Cheng","institutions":["zgca"]},{"name":"Alexey K Shaytan"},{"name":"Bin Liu","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Beijing Institute of Technology School of Computer Science and Technology, , Beijing 100081,","Zhongguancun Academy , Beijing 100094,","Jilin University School of Artificial Intelligence, , Jilin 130012,","Lomonosov Moscow State University Department of Biology, , Moscow, Russia","SMBU-MSU-BIT Joint Laboratory on Bioinformatics and Engineering Biology, Shenzhen MSU-BIT University , Shenzhen, Guangdong 518172,","School of Computer Science and Technology, Beijing Institute of Technology , Beijing 100081,","School of Artificial Intelligence, Jilin University , Jilin 130012,","Department of Biology, Lomonosov Moscow State University , Moscow 119991, Russia","SMBU-MSU-BIT Joint Laboratory on Bioinformatics and Engineering Biology, Shenzhen MSU-BIT University , Shenzhen 518172,"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Bioinformatics","status":"published","topics":[],"abstract":"Abstract Motivation Single-cell chemical perturbation profiling offers a powerful opportunity to organize drugs by shared mechanism-associated transcriptional responses, but observed transcriptional responses are entangled with contextual variation from cell identity, dose and treatment time. As a result, models that perform well in perturbation-response prediction may still learn latent spaces dominated by context-associated structure rather than transferable drug-associated signal. We developed DECANT to learn mechanism-aligned perturbation representations that remain stable across context shifts while preserving response fidelity. Results DECANT represents each perturbation as a matched treated–control cell set and separates a context-suppressed, mechanism-aligned perturbation representation from context-dependent response information. The resulting mechanism-aligned perturbation space is shaped to support drug-level retrieval and biological interpretation. Under a fixed drug-level unseen-compound benchmark, DECANT achieved the strongest overall response-difference profile among adapted published perturbation models and strong pseudo-bulk baselines across gene- and program-level metrics. Beyond prediction, DECANT produced embeddings that remained stable across changes in dose, cell line and treatment time, recovered drug neighborhoods enriched for shared mechanism-family annotations, and linked these neighborhoods to interpretable downstream consequence programs. Ablation analyses showed that mechanism–context decoupling provided the main signal-separation backbone, whereas retrieval-oriented shaping was critical for organizing local representation-space geometry. These results support DECANT as a framework for learning context-robust, mechanism-aligned perturbation representations from single-cell transcriptional responses, providing a basis for mechanism-aligned perturbation analysis and representation-based compound prioritization. Availability and Implementation The DECANT web server is publicly available at http://bliulab.net/DECANT. All source code and analysis scripts are available at https://github.com/bliulab/DECANT and archived on Zenodo at https://doi.org/10.5281/zenodo.21216567. Supplementary information Supplementary data are available at Bioinformatics online.","identifiers":{"doi":"10.1093/bioinformatics/btag662"},"links":[{"label":"DOI","url":"https://doi.org/10.1093/bioinformatics/btag662"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1093/bioinformatics/btag662"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy , Beijing 100094,","source":"Crossref","sourceUrl":"https://doi.org/10.1093/bioinformatics/btag662"}],"sources":["Crossref"],"id":"doi-10-1093-bioinformatics-btag662","updatedAt":"2026-09-08T04:58:29Z"},{"type":"article","title":"Competing Sn–O and Sn–C Bond Cleavage Pathways Control Cross-Linking in Tin-Oxo Clusters","authors":[{"name":"Taoli Guo","institutions":["zgca"]},{"name":"Chen Zhu"},{"name":"Lei Zhang"},{"name":"Feng Luo"},{"name":"Jin-Cheng Liu"}],"institutions":["zgca"],"rawAffiliations":["Nankai University , , ,","Zhongguancun Academy , ,"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"The Journal of Physical Chemistry Letters","status":"published","topics":[],"abstract":"Abstract Tin-oxo clusters, owing to their exceptionally high extreme ultraviolet (EUV) absorption cross sections, have emerged as promising photoresist materials for EUV lithography, yet their atomic-scale photochemical mechanisms remain poorly understood. Here, by combining density functional theory (DFT) and large-scale molecular dynamics enabled by machine-learning interatomic potentials (MLIPs), we reveal a previously underexplored atomistic pathway for cross-linking in tin-oxo clusters. Beyond conventional Sn–C bond cleavage, cross-linking is strongly influenced by ligand-controlled destabilization of the Sn–O cage framework. Cleavage of Sn–O bonds disrupts cage structural integrity and generates coordinatively unsaturated tin centers that actively facilitate intercluster linkage formation. Notably, the free-energy cost associated with Sn–O bond cleavage is comparable to that of Sn–C dissociation under the same simulation protocol, which identifies framework instability as a driving factor in the structural evolution and cross-linking of tin-oxo photoresists. We further demonstrate that ligand identity critically governs cross-linking behavior by modulating both Sn–C stability and cage resilience: vinyl-functionalized clusters form extensive cross-linked networks containing aggregates up to Sn160 during MLIP molecular dynamics simulation, whereas phenyl ligands largely suppress cross-linking due to stronger Sn–C bonding and steric stabilization. These findings expand the mechanistic picture of tin-oxo photoresist cross-linking and provide atomistic insight for the molecular design of tin-oxo photoresist materials.","identifiers":{"doi":"10.1021/acs.jpclett.6c01438"},"links":[{"label":"DOI","url":"https://doi.org/10.1021/acs.jpclett.6c01438"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1021/acs.jpclett.6c01438"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy , ,","source":"Crossref","sourceUrl":"https://doi.org/10.1021/acs.jpclett.6c01438"}],"sources":["Crossref"],"id":"doi-10-1021-acs-jpclett-6c01438","updatedAt":"2026-08-12T03:04:08Z"},{"type":"article","title":"GPU Accelerated Minimal Auxiliary Basis Approach TDDFT for Large Organic Molecules","authors":[{"name":"Zehao Zhou","institutions":["zgca"]},{"name":"Xiaojie Wu"},{"name":"Yanheng Li"},{"name":"Xinran Wei","institutions":["zgca"]},{"name":"Cheng Fan"},{"name":"Fusong Ju","institutions":["zgca"]},{"name":"Qiming Sun"},{"name":"Yi Qin Gao"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy","Bytedance Seed","The College of Chemistry and Molecular Engineering","Peking University","New Cornerstone Science Laboratory, The College of Chemistry and Molecular Engineering","West Lake University"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Journal of Chemical Theory and Computation","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1021/acs.jctc.6c00629"},"links":[{"label":"DOI","url":"https://doi.org/10.1021/acs.jctc.6c00629"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1021/acs.jctc.6c00629"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.1021/acs.jctc.6c00629"}],"sources":["Crossref"],"id":"doi-10-1021-acs-jctc-6c00629","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"A Synergistic Strategy for Data‐Constrained Deep Learning in Materials Science","authors":[{"name":"Chun Ting Shao"},{"name":"Yi Chen","institutions":["zgca"]},{"name":"Shan Man Song"},{"name":"Jian Xu","institutions":["zgca"]},{"name":"Peipei Yang"},{"name":"Qing Bo Yan"},{"name":"Gang Su"}],"institutions":["zgca"],"rawAffiliations":["Kavli Institute for Theoretical Sciences School of Physical Sciences University of Chinese Academy of Sciences Beijing China","State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation Chinese Academy of Sciences Beijing China","School of Artificial Intelligence University of Chinese Academy of Sciences Beijing China","Zhongguancun Academy Beijing China","Center of Materials Science and Optoelectronics Engineering College of Materials Science and Opto‐Electronic Technology University of Chinese Academy of Sciences Beijing China","Institute of Theoretical Physics Chinese Academy of Sciences Beijing China"],"relationType":"affiliation","publishedAt":"2026-01-01","year":2026,"venue":"Materials Genome Engineering Advances","status":"published","topics":[],"abstract":"ABSTRACT Materials science research increasingly benefits from the application of machine learning methods, yet encounters fundamental challenges from data scarcity, such as limited dataset sizes and severe distribution imbalance. In this paper, we propose a hybrid framework integrating attention pooling, multi‐task learning, auxiliary learning, and classification‐corrected regression. Using a 2D materials dataset as a case study, our approach demonstrates significantly enhanced prediction accuracy over the baseline crystal graph convolutional neural networks (CGCNN) method. Specifically, it reduces the mean absolute error for work function prediction from 0.312 to 0.240 eV, and for band gap from 0.301 to 0.230 eV. The framework also proves effective with other graph neural network methods such as atomistic line graph neural network (ALIGNN). These gains stem from the framework's ability to exploit underlying physical correlations between material properties and atomic structures. Through extensive experiments, we demonstrate that attention pooling serves as a generally effective component for diverse property prediction tasks, particularly with small datasets, which also offers the possibility of interpretability analysis through element attention weights. Our architecture enables seamless integration with various graph‐based or other end‐to‐end deep learning models, presenting a computationally efficient and easily implementable solution for constrained datasets in materials science.","identifiers":{"doi":"10.1002/mgea.70065"},"links":[{"label":"DOI","url":"https://doi.org/10.1002/mgea.70065"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1002/mgea.70065"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy Beijing China","source":"Crossref","sourceUrl":"https://doi.org/10.1002/mgea.70065"}],"sources":["Crossref"],"id":"doi-10-1002-mgea-70065","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents","authors":[{"name":"Zhaoxi Zhang"},{"name":"Yitong Duan"},{"name":"Yanzhi Zhang"},{"name":"Yiming Xu"},{"name":"Zhixiang Wang"},{"name":"Kun Liang"},{"name":"Weikang Li"},{"name":"Jiahui Liang"},{"name":"Deguo Xia"},{"name":"Jizhou Huang"},{"name":"Jiyan He"},{"name":"Yunfang Wu"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）"],"relationType":"official-output","publishedAt":"2025-12-24","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.SE","cs.AI"],"abstract":"Locating files and functions requiring modification in large software repositories is challenging due to their scale and structural complexity. Existing LLM-based methods typically treat this as a repository-level retrieval task and rely on multiple auxiliary tools, which often overlook code execution logic and complicate model control. We propose RepoNavigator, an LLM agent equipped with a single execution-aware tool: jumping to the definition of an invoked symbol. This unified design reflects the actual flow of code execution while simplifying tool manipulation. RepoNavigator is trained end-to-end via Reinforcement Learning (RL) directly from a base pretrained model, without relying on closed-source distillation. Experiments demonstrate that RL-trained RepoNavigator achieves state-of-the-art performance, with the 7B model outperforming 14B baselines, the 14B model surpassing 32B competitors, and the 32B model exceeding closed-source models such as GPT-5 on most metrics. These results confirm that integrating a single, structurally grounded tool with RL training provides an efficient and scalable solution for repository-level issue localization.","identifiers":{"arxiv":"2512.20957","doi":"10.48550/arXiv.2512.20957"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.20957"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.20957"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_mycb9w8jxc6gpcyeq7j8wc4cn6lga2l9"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.20957v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_mycb9w8jxc6gpcyeq7j8wc4cn6lga2l9"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_mycb9w8jxc6gpcyeq7j8wc4cn6lga2l9"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2512-20957","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale","authors":[{"name":"Linfeng Zhang"},{"name":"Siheng Chen"},{"name":"Yuzhu Cai"},{"name":"Jingyi Chai"},{"name":"Junhan Chang"},{"name":"Kun Chen"},{"name":"Zhi X. Chen"},{"name":"Zhaohan Ding"},{"name":"Yuwen Du"},{"name":"Yuanpeng Gao"},{"name":"Yuan Gao"},{"name":"Jing Gao"},{"name":"Zhifeng Gao"},{"name":"Qiangqiang Gu"},{"name":"Yanhui Hong"},{"name":"Yuan Huang"},{"name":"Xi Fang"},{"name":"Xiaohong Ji"},{"name":"Guolin Ke"},{"name":"Zixing Lei"},{"name":"Xinyu Li"},{"name":"Yongge Li"},{"name":"Ruoxue Liao"},{"name":"Hang Lin"},{"name":"Xiaolu Lin"},{"name":"Yuxiang Liu"},{"name":"Xinzijian Liu"},{"name":"Zexi Liu"},{"name":"Jintan Lu"},{"name":"Tingjia Miao"},{"name":"Haohui Que"},{"name":"Weijie Sun"},{"name":"Yanfeng Wang"},{"name":"Bingyang Wu"},{"name":"Tianju Xue"},{"name":"Rui Ye"},{"name":"Jinzhe Zeng"},{"name":"Duo Zhang"},{"name":"Jiahui Zhang"},{"name":"Linfeng Zhang"},{"name":"Tianhan Zhang"},{"name":"Wenchang Zhang"},{"name":"Yuzhi Zhang"},{"name":"Zezhong Zhang"},{"name":"Hang Zheng"},{"name":"Hui Zhou"},{"name":"Tong Zhu"},{"name":"Xinyu Zhu"},{"name":"Qingguo Zhou"},{"name":"Weinan E"}],"institutions":["zgca"],"rawAffiliations":["Jing Gao Affiliation: DP Technology, Beijing, China Affiliation: Shanghai Jiao Tong University, Shanghai, China Affiliation: Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2025-12-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.AI"],"abstract":"AI agents are emerging as a practical way to run multi-step scientific workflows that interleave reasoning with tool use and verification, pointing to a shift from isolated AI-assisted steps toward \\emph{agentic science at scale}. This shift is increasingly feasible, as scientific tools and models can be invoked through stable interfaces and verified with recorded execution traces, and increasingly necessary, as AI accelerates scientific output and stresses the peer-review and publication pipeline, raising the bar for traceability and credible evaluation. However, scaling agentic science remains difficult: workflows are hard to observe and reproduce; many tools and laboratory systems are not agent-ready; execution is hard to trace and govern; and prototype AI Scientist systems are often bespoke, limiting reuse and systematic improvement from real workflow signals. We argue that scaling agentic science requires an infrastructure-and-ecosystem approach, instantiated in Bohrium+SciMaster. Bohrium acts as a managed, traceable hub for AI4S assets -- akin to a HuggingFace of AI for Science -- that turns diverse scientific data, software, compute, and laboratory systems into agent-ready capabilities. SciMaster orchestrates these capabilities into long-horizon scientific workflows, on which scientific agents can be composed and executed. Between infrastructure and orchestration, a \\emph{scientific intelligence substrate} organizes reusable models, knowledge, and components into executable building blocks for workflow reasoning and action, enabling composition, auditability, and improvement through use. We demonstrate this stack with eleven representative master agents in real workflows, achieving orders-of-magnitude reductions in end-to-end scientific cycle time and generating execution-grounded signals from real workloads at multi-million scale.","identifiers":{"arxiv":"2512.20469","doi":"10.48550/arXiv.2512.20469"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.20469"},{"label":"HTML","url":"https://arxiv.org/html/2512.20469v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.20469"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.20469v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Jing Gao Affiliation: DP Technology, Beijing, China Affiliation: Shanghai Jiao Tong University, Shanghai, China Affiliation: Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2512.20469v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2512-20469","updatedAt":"2026-08-25T02:44:34Z"},{"type":"preprint","title":"Degradation-Aware Metric Prompting for Hyperspectral Image Restoration","authors":[{"name":"Binfeng Wang"},{"name":"Di Wang"},{"name":"Haonan Guo"},{"name":"Ying Fu"},{"name":"Jing Zhang"}],"institutions":["zgca"],"rawAffiliations":["Binfeng Wang Affiliation: Beijing Institute of Technology Affiliation: Zhongguancun Academy wbf_bit@163.com; {d_wang,haonan.guo}@whu.edu.cn; fuying@bit.edu.cn; jingzhang.cv@gmail.com"],"relationType":"affiliation","publishedAt":"2025-12-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.CV","eess.IV"],"abstract":"Unified hyperspectral image (HSI) restoration aims to recover diverse degradations within a single model. However, current methods often rely on impractical explicit priors or opaque black-box representations that overfit to training distributions, hampering generalization to unseen scenarios. To bridge this gap, we propose Degradation-Aware Metric Prompting (DAMP), a novel framework that characterizes multi-dimensional degradations through interpretable spatial-spectral metrics. These metrics serve as Degradation Prompts (DP), enabling the model to capture shared characteristics across tasks and adapt to unknown corruptions. Central to our framework is the Degradation-Adaptive Mixture-of-Experts (DAMoE), where Spatial-Spectral Adaptive Modules (SSAMs) serve as experts that utilize learnable fusion coefficients to specialize in distinct degradation degrees. By using DP as a gating router, DAMoE dynamically activates specialized experts tailored to the specific degradation profile. Extensive experiments on natural and remote sensing HSI datasets demonstrate that DAMP achieves state-of-the-art performance and exhibits exceptional zero-shot generalization on unseen restoration tasks. Code is publicly available at \\href{DAMP}{https://github.com/MiliLab/DAMP}.","identifiers":{"arxiv":"2512.20251","doi":"10.48550/arXiv.2512.20251"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.20251"},{"label":"HTML","url":"https://arxiv.org/html/2512.20251v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.20251"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.20251v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Binfeng Wang Affiliation: Beijing Institute of Technology Affiliation: Zhongguancun Academy wbf_bit@163.com; {d_wang,haonan.guo}@whu.edu.cn; fuying@bit.edu.cn; jingzhang.cv@gmail.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2512.20251v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2512-20251","updatedAt":"2026-08-25T02:44:34Z"},{"type":"preprint","title":"Understanding Generalization in Role-Playing Models via Information Theory","authors":[{"name":"Yongqi Li"},{"name":"Hao Lang"},{"name":"Fei Huang"},{"name":"Tieyun Qian"},{"name":"Yongbin Li"}],"institutions":["zgca"],"rawAffiliations":["Tieyun Qian Affiliation: Zhongguancun Academy {liyongqi,qty}@whu.edu.cn , {hao.lang,f.huang,shuide.lyb}@alibaba-inc.com"],"relationType":"affiliation","publishedAt":"2025-12-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI","cs.CL"],"abstract":"Role-playing models (RPMs) are widely used in real-world applications but underperform when deployed in the wild. This degradation can be attributed to distribution shifts, including user, character, and dialogue compositional shifts. Existing methods like LLM-as-a-judge fall short in providing a fine-grained diagnosis of how these shifts affect RPM generalization, and thus there lack formal frameworks to characterize RPM generalization behaviors. To bridge these gaps, we introduce an information-theoretic metric, named reasoning-based effective mutual information difference (R-EMID), to measure RPM performance degradation in an interpretable way. We also derive an upper bound on R-EMID to predict the worst-case generalization performance of RPMs and theoretically reveal how various shifts contribute to the RPM performance degradation. Moreover, we propose a co-evolving reinforcement learning framework to adaptively model the connection among user, character, and dialogue context and thus enhance the estimation of dialogue response generation probability, which is critical for calculating R-EMID. Finally, we evaluate the generalization performance of various RPMs using R-EMID, finding that user shift poses the highest risk among all shifts and reinforcement learning is the most effective approach for enhancing RPM generalization.","identifiers":{"arxiv":"2512.17270","doi":"10.48550/arXiv.2512.17270"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.17270"},{"label":"HTML","url":"https://arxiv.org/html/2512.17270v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.17270"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.17270v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Tieyun Qian Affiliation: Zhongguancun Academy {liyongqi,qty}@whu.edu.cn , {hao.lang,f.huang,shuide.lyb}@alibaba-inc.com","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2512.17270v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2512-17270","updatedAt":"2026-08-24T02:51:25Z"},{"type":"preprint","title":"Red Teaming Large Reasoning Models","authors":[{"name":"Jiawei Chen"},{"name":"Yang Yang"},{"name":"Chao Yu"},{"name":"Yu Tian"},{"name":"Zhi Cao"},{"name":"Xue Yang"},{"name":"Linghao Li"},{"name":"Hang Su"},{"name":"Zhaoxia Yin"}],"institutions":["zgca"],"rawAffiliations":["Zhi Cao Affiliation: East China Normal University, Tsinghua University, Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2025-12-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.CR","cs.AI"],"abstract":"Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical consistency through explicit chains of thought (CoT). However, these models introduce novel safety and reliability risks, such as CoT-hijacking and prompt-induced inefficiencies, which are not fully captured by existing evaluation methods. To address this gap, we propose RT-LRM, a unified benchmark designed to assess the trustworthiness of LRMs. RT-LRM evaluates three core dimensions: truthfulness, safety and efficiency. Beyond metric-based evaluation, we further introduce the training paradigm as a key analytical perspective to investigate the systematic impact of different training strategies on model trustworthiness. We achieve this by designing a curated suite of 30 reasoning tasks from an observational standpoint. We conduct extensive experiments on 26 models and identify several valuable insights into the trustworthiness of LRMs. For example, LRMs generally face trustworthiness challenges and tend to be more fragile than Large Language Models (LLMs) when encountering reasoning-induced risks. These findings uncover previously underexplored vulnerabilities and highlight the need for more targeted evaluations. In addition, we release a scalable toolbox for standardized trustworthiness research to support future advancements in this important field. Our code and datasets will be open-sourced.","identifiers":{"arxiv":"2512.00412","doi":"10.48550/arXiv.2512.00412"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.00412"},{"label":"HTML","url":"https://arxiv.org/html/2512.00412v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.00412"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.00412v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhi Cao Affiliation: East China Normal University, Tsinghua University, Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2512.00412v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2512-00412","updatedAt":"2026-08-24T02:51:25Z"},{"type":"preprint","title":"Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model","authors":[{"name":"Jing He"},{"name":"Haodong Li"},{"name":"Mingzhi Sheng"},{"name":"Ying-Cong Chen"}],"institutions":["zgca"],"rawAffiliations":["1 College of Artificial Intelligence, Zhejiang University · 2 Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2025-11-30","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.CV"],"abstract":"Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings between 2D observations and 3D structures. While discriminative regression models achieve strong performance through large-scale supervision, their success is bounded by the scale, quality, and diversity of available data, as well as by limited physical reasoning. Recent diffusion models exhibit powerful world priors that encode geometry and semantics learned from massive image-text data, yet directly reusing their stochastic generative formulation is suboptimal for deterministic geometric inference: the former is optimized for diverse and high-fidelity image generation, whereas the latter requires stable and accurate predictions. In this work, we propose Lotus-2, a two-stage deterministic framework for stable, accurate and fine-grained geometric dense prediction, aiming to provide an optimal adaptation protocol to fully exploit the pre-trained generative priors. Specifically, in the first stage, the core predictor employs a single-step deterministic formulation with a clean-data objective and a lightweight local continuity module (LCM) to generate globally coherent structures without grid artifacts. In the second stage, the detail sharpener performs a constrained multi-step rectified-flow refinement within the manifold defined by the core predictor, enhancing fine-grained geometry through noise-free deterministic flow matching. Using only 59K training samples, less than 1% of existing large-scale datasets, Lotus-2 establishes new state-of-the-art results in monocular depth estimation and highly competitive surface normal prediction. These results demonstrate that diffusion models can serve as deterministic world priors, enabling high-quality geometric reasoning beyond traditional discriminative and generative paradigms.","identifiers":{"arxiv":"2512.01030","doi":"10.48550/arXiv.2512.01030"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2512.01030"},{"label":"PDF","url":"https://arxiv.org/pdf/2512.01030"},{"label":"Code","url":"https://github.com/longxiang-ai/TransNormal-2"},{"label":"Model","url":"https://huggingface.co/black-forest-labs/FLUX.2-klein-base-9B"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2512.01030v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 College of Artificial Intelligence, Zhejiang University · 2 Zhongguancun Academy","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TransNormal-2"}],"sources":["arXiv","Official GitHub project README"],"id":"doi-10-48550-arxiv-2512-01030","updatedAt":"2026-09-07T05:00:23Z"},{"type":"preprint","title":"DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models","authors":[{"name":"Cheng Yin"},{"name":"Yankai Lin"},{"name":"Wang Xu"},{"name":"Sikyuen Tam"},{"name":"Xiangrui Zeng"},{"name":"Zhiyuan Liu"},{"name":"Zhouping Yin"}],"institutions":["zgca"],"rawAffiliations":["Cheng Yin Affiliation: The School of Mechanical Science and Engineering, Huazhong University of Science and Techn-ology, China Affiliation: Beijing Zhongguancun Academy, China yinchenghust@hust.edu.cnyankailin@ruc.edu.cnzeng@hust.edu.cn"],"relationType":"affiliation","publishedAt":"2025-11-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI","cs.RO"],"abstract":"Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsistent gains, yet no prior work has rigorously diagnosed when and why CoT helps robots act. Through systematic experiments, we identify two necessary conditions that must be jointly satisfied for CoT to be effective in VLA: (1) Decoding Alignment: CoT and actions must be generated with modality-appropriate mechanisms; forcing both through a single autoregressive decoder is not merely suboptimal but actively harmful, degrading performance by 4.2 percentage points; (2) Causal Alignment: CoT must be causally linked to task success via outcome-based optimization; without it, supervised CoT is indistinguishable from no reasoning at all under action-execution-sensitive dynamics shift, exhibiting a 32.0 pp performance drop nearly identical to the 31.6 pp drop of a reasoning-free baseline. Guided by these findings, we build DeepThinkVLA: a hybrid-attention decoder satisfies Condition 1 by pairing causal attention for language with bidirectional attention for parallel action decoding, while a two-stage SFT-then-RL pipeline satisfies Condition 2 by aligning the full reasoning: action chain with sparse task-success rewards. DeepThinkVLA achieves 97.0\\% success on LIBERO, 79.0\\% robustness on LIBERO-Plus (vs. 61.6\\% for $π_0$-FAST), and 59.3\\% success on RoboTwin 2.0, exceeding the strongest baseline by 21.7 points. Furthermore, real-robot experiments provide preliminary evidence for the physical applicability of our CoT data construction and hybrid architecture. Our codes are available at https://github.com/OpenBMB/DeepThinkVLA.","identifiers":{"arxiv":"2511.15669","doi":"10.48550/arXiv.2511.15669"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2511.15669"},{"label":"HTML","url":"https://arxiv.org/html/2511.15669v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2511.15669"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2511.15669v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Cheng Yin Affiliation: The School of Mechanical Science and Engineering, Huazhong University of Science and Techn-ology, China Affiliation: Beijing Zhongguancun Academy, China yinchenghust@hust.edu.cnyankailin@ruc.edu.cnzeng@hust.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2511.15669v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2511-15669","updatedAt":"2026-08-24T02:51:25Z"},{"type":"preprint","title":"Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network","authors":[{"name":"Keyu Zhao"},{"name":"Weiquan Lin"},{"name":"Qirui Zheng"},{"name":"Fengli Xu"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2025-11-01","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.AI"],"abstract":"Novel research ideas play a critical role in advancing scientific inquiries. Recent advancements in Large Language Models (LLMs) have demonstrated their potential to generate novel research ideas by leveraging large-scale scientific literature. However, previous work in research ideation has primarily relied on simplistic methods, such as keyword co-occurrence or semantic similarity. These approaches focus on identifying statistical associations in the literature but overlook the complex, contextual relationships between scientific concepts, which are essential to effectively leverage knowledge embedded in human literature. For instance, papers that simultaneously mention \"keyword A\" and \"keyword B\" often present research ideas that integrate both concepts. Additionally, some LLM-driven methods propose and refine research ideas using the model's internal knowledge, but they fail to effectively utilize the scientific concept network, limiting the grounding of ideas in established research. To address these challenges, we propose the Deep Ideation framework to address these challenges, integrating a scientific network that captures keyword co-occurrence and contextual relationships, enriching LLM-driven ideation. The framework introduces an explore-expand-evolve workflow to iteratively refine research ideas, using an Idea Stack to track progress. A critic engine, trained on real-world reviewer feedback, guides the process by providing continuous feedback on the novelty and feasibility of ideas. Our experiments show that our approach improves the quality of generated ideas by 10.67% compared to other methods, with ideas surpassing top conference acceptance levels. Human evaluation highlights their practical value in scientific research, and ablation studies confirm the effectiveness of each component in the workflow. Code repo is available at https://github.com/kyZhao-1/Deep-Ideation.","identifiers":{"arxiv":"2511.02238","doi":"10.48550/arXiv.2511.02238"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2511.02238"},{"label":"HTML","url":"https://arxiv.org/html/2511.02238v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2511.02238"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2511.02238v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2511.02238v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2511-02238","updatedAt":"2026-08-24T02:51:25Z"},{"type":"preprint","title":"Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction","authors":[{"name":"Shaohua Liu"},{"name":"Ning Gao"},{"name":"Zuoya Gu"},{"name":"Hongkun Dou"},{"name":"Yue Deng"},{"name":"Hongjue Li"}],"institutions":["zgca"],"rawAffiliations":["Yue Deng OrcID: 0000-0003-2871-8922 Affiliation: School of Artificial Intelligence , Beihang University , Beijing , China Affiliation: Zhongguancun Academy , Beijing , China email: ydeng@buaa.edu.cn"],"relationType":"affiliation","publishedAt":"2025-10-25","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatiotemporal degradations in underwater imaging, including caustics, flickering, attenuation, and backscattering, frequently result in inaccurate geometry and appearance in existing 3D reconstruction methods. While a few recent works have explored underwater degradation-aware reconstruction, they often address either spatial or temporal degradation alone, falling short in more real-world underwater scenarios where both types of degradation occur. We propose MarineSTD-GS, a novel 3D Gaussian Splatting-based framework that explicitly models both temporal and spatial degradations for realistic underwater scene reconstruction. Specifically, we introduce two paired Gaussian primitives: Intrinsic Gaussians represent the true scene, while Degraded Gaussians render the degraded observations. The color of each Degraded Gaussian is physically derived from its paired Intrinsic Gaussian via a Spatiotemporal Degradation Modeling (SDM) module, enabling self-supervised disentanglement of realistic appearance from degraded images. To ensure stable training and accurate geometry, we further propose a Depth-Guided Geometry Loss and a Multi-Stage Optimization strategy. We also construct a simulated benchmark with diverse spatial and temporal degradations and ground-truth appearances for comprehensive evaluation. Experiments on both simulated and real-world datasets show that MarineSTD-GS robustly handles spatiotemporal degradations and outperforms existing methods in novel view synthesis with realistic, water-free scene appearances.","identifiers":{"arxiv":"2604.23551","doi":"10.48550/arXiv.2604.23551","publishedDoi":"10.1145/3746027.3754888"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2604.23551"},{"label":"HTML","url":"https://arxiv.org/html/2604.23551v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2604.23551"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2604.23551v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yue Deng OrcID: 0000-0003-2871-8922 Affiliation: School of Artificial Intelligence , Beihang University , Beijing , China Affiliation: Zhongguancun Academy , Beijing , China email: ydeng@buaa.edu.cn","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2604.23551v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2604-23551","updatedAt":"2026-08-13T11:57:17Z"},{"type":"preprint","title":"Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement","authors":[{"name":"Ruofan Wu"},{"name":"Yifan Zhao"},{"name":"Jia Li"}],"institutions":["zgca"],"rawAffiliations":["2 Zhongguancun Academy, Beijing, China"],"relationType":"affiliation","publishedAt":"2025-10-19","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Class-Incremental Semantic Segmentation (CISS) requires continuous learning of newly introduced classes while retaining knowledge of past classes. By abstracting mainstream methods into two stages (visual feature extraction and prototype-feature matching), we identify a more fundamental challenge termed catastrophic semantic entanglement. This phenomenon involves Prototype-Feature Entanglement caused by semantic misalignment during the incremental process, and Background-Increment Entanglement due to dynamic data evolution. Existing techniques, which rely on visual feature learning without sufficient cues to distinguish targets, introduce significant noise and errors. To address these issues, we introduce a Language-inspired Bootstrapped Disentanglement framework (LBD). We leverage the prior class semantics of pre-trained visual-language models (e.g., CLIP) to guide the model in autonomously disentangling features through Language-guided Prototypical Disentanglement and Manifold Mutual Background Disentanglement. The former guides the disentangling of new prototypes by treating hand-crafted text features as topological templates, while the latter employs multiple learnable prototypes and mask-pooling-based supervision for background-incremental class disentanglement. By incorporating soft prompt tuning and encoder adaptation modifications, we further bridge the capability gap of CLIP between dense and sparse tasks, achieving state-of-the-art performance on both Pascal VOC and ADE20k, particularly in multi-step scenarios.","identifiers":{"arxiv":"2509.00527","doi":"10.48550/arXiv.2509.00527","publishedDoi":"10.1109/iccv51701.2025.02008"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2509.00527"},{"label":"HTML","url":"https://arxiv.org/html/2509.00527v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2509.00527"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2509.00527v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"2 Zhongguancun Academy, Beijing, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2509.00527v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2509-00527","updatedAt":"2026-08-12T04:55:57Z"},{"type":"preprint","title":"RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation","authors":[{"name":"Chao Yu"},{"name":"Yuanqing Wang"},{"name":"Zhen Guo"},{"name":"Hao Lin"},{"name":"Si Xu"},{"name":"Hongzhi Zang"},{"name":"Quanlu Zhang"},{"name":"Yongji Wu"},{"name":"Chunyang Zhu"},{"name":"Junhao Hu"},{"name":"Zixiao Huang"},{"name":"Mingjie Wei"},{"name":"Yuqing Xie"},{"name":"Ke Yang"},{"name":"Bo Dai"},{"name":"Zhexuan Xu"},{"name":"Jiakun Du"},{"name":"Xiangyuan Wang"},{"name":"Xu Fu"},{"name":"Letong Shi"},{"name":"Zhihao Liu"},{"name":"Kang Chen"},{"name":"Weilin Liu"},{"name":"Gang Liu"},{"name":"Boxun Li"},{"name":"Jianlei Yang"},{"name":"Zhi Yang"},{"name":"Guohao Dai"},{"name":"Yu Wang"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2025-09-19","year":2025,"venue":"arXiv","status":"preprint","topics":["cs.LG","cs.AI","cs.DC"],"abstract":"Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\\times-2.43\\times$ speedup in end-to-end training throughput.","identifiers":{"arxiv":"2509.15965","doi":"10.48550/arXiv.2509.15965"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2509.15965"},{"label":"PDF","url":"https://arxiv.org/pdf/2509.15965"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_yps1zc0z9d3bf6owk6ighwbvst93ajw1"},{"label":"Code","url":"https://github.com/RLinf/RLinf"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2509.15965v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_yps1zc0z9d3bf6owk6ighwbvst93ajw1"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2509-15965","updatedAt":"2026-08-08T11:59:38Z"},{"id":"doi-10-1093-bioinformatics-btaf489","type":"article","title":"Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning framework","authors":[{"name":"Ke Yan","institutions":["zgca"]},{"name":"Shutao Chen"},{"name":"Bin Liu","institutions":["zgca"]},{"name":"Hao Wu"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy, Beijing 100094, China","School of Computer Science and Technology, Beijing Institute of Technology , Beijing 100081,","Zhongguancun Academy , Beijing 100094,","SMBU-MSU-BIT Joint Laboratory on Bioinformatics and Engineering Biology, Shenzhen MSU-BIT University , Shenzhen, Guangdong 518172,"],"relationType":"affiliation","publishedAt":"2025-09-04","year":2025,"venue":"Bioinformatics","status":"published","topics":["Bioinformatics","Peptides","Machine Learning","Toxicity Prediction"],"abstract":"Abstract Motivation Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. Results This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. Availability and implementation The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L.","identifiers":{"doi":"10.1093/bioinformatics/btaf489"},"links":[{"label":"DOI","url":"https://doi.org/10.1093/bioinformatics/btaf489"},{"label":"PMC","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12457739/"},{"label":"Code","url":"http://bliulab.net/ToxPre-2L"}],"versions":[{"label":"Open access version","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12457739/"},{"label":"Publisher version","url":"https://doi.org/10.1093/bioinformatics/btaf489"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing 100094, China","source":"PubMed Central","sourceUrl":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12457739/"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing 100094, China","source":"Publisher PDF","sourceUrl":"https://doi.org/10.1093/bioinformatics/btaf489"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy , Beijing 100094,","source":"Crossref","sourceUrl":"https://doi.org/10.1093/bioinformatics/btaf489"}],"sources":["Europe PMC","Crossref","Publisher"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"conference","title":"CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing","authors":[{"name":"Tianhui Liu"},{"name":"Hetian Pang"},{"name":"Xin Zhang"},{"name":"Tianjian Ouyang"},{"name":"Zhiyuan Zhang"},{"name":"Jie Feng"},{"name":"Yong Li"},{"name":"Pan Hui"}],"institutions":["zgca"],"rawAffiliations":["北京中关村学院官网科研创新页面"],"relationType":"official-output","publishedAt":"2025-05-31","year":2025,"venue":"ICLR 2026","status":"published","topics":["cs.AI","cs.CL"],"abstract":"Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introduce \\textit{CityLens}, a comprehensive benchmark designed to evaluate the capabilities of Large Vision-Language Models (LVLMs) in predicting socioeconomic indicators from satellite and street view imagery. We construct a multi-modal dataset covering a total of 17 globally distributed cities, spanning 6 key domains: economy, education, crime, transport, health, and environment, reflecting the multifaceted nature of urban life. Based on this dataset, we define 11 prediction tasks and utilize 3 evaluation paradigms: Direct Metric Prediction, Normalized Metric Estimation, and Feature-Based Regression. We benchmark 17 state-of-the-art LVLMs across these tasks. These make CityLens the most extensive socioeconomic benchmark to date in terms of geographic coverage, indicator diversity, and model scale. Our results reveal that while LVLMs demonstrate promising perceptual and reasoning capabilities, they still exhibit limitations in predicting urban socioeconomic indicators. CityLens provides a unified framework for diagnosing these limitations and guiding future efforts in using LVLMs to understand and predict urban socioeconomic patterns. The code and data are available at https://github.com/tsinghua-fib-lab/CityLens.","identifiers":{"arxiv":"2506.00530","doi":"10.48550/arXiv.2506.00530"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2506.00530"},{"label":"PDF","url":"https://arxiv.org/pdf/2506.00530"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_xb8ky5d1m0w0i3zjeh4jpnsks22oazlc"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2506.00530v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院官网科研创新页面","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_xb8ky5d1m0w0i3zjeh4jpnsks22oazlc"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2506-00530","updatedAt":"2026-08-08T11:59:38Z"},{"type":"dataset","title":"TransLab","authors":[{"name":"longxiang-ai"}],"institutions":["zgca"],"rawAffiliations":["1 Zhejiang University, 2 Zhongguancun Academy, Beijing, 3 Xi'an Jiaotong University, 4 Beijing Normal University"],"relationType":"official-output","publishedAt":"2025-04-17","year":2025,"venue":"Hugging Face Datasets","status":"released","topics":["Open Dataset"],"abstract":"Dataset released alongside TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors","identifiers":{},"links":[{"label":"Dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransLab"},{"label":"Project","url":"https://github.com/longxiang-ai/TSGS"}],"versions":[{"label":"Hugging Face dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransLab"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, 2 Zhongguancun Academy, Beijing, 3 Xi'an Jiaotong University, 4 Beijing Normal University","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TSGS"}],"sources":["Official GitHub project README","Hugging Face"],"id":"work-b19a6ad16790","updatedAt":"2026-08-08T12:09:53Z"},{"type":"conference","title":"TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors","authors":[{"name":"Mingwei Li"},{"name":"Pu Pang"},{"name":"Hehe Fan"},{"name":"Hua Huang"},{"name":"Yi Yang"}],"institutions":["zgca"],"rawAffiliations":["1 Zhejiang University, 2 Zhongguancun Academy, Beijing, 3 Xi'an Jiaotong University, 4 Beijing Normal University","Zhejiang University, Hangzhou, Zhejiang, China and Zhongguancun Academy, Beijing, China","Xi'an Jiaotong University, Xi'an, Shaanxi, China and Zhongguancun Academy, Beijing, China","Zhejiang University, Hangzhou, Zhejiang, China","Beijing Normal University, Beijing, China"],"relationType":"affiliation","publishedAt":"2025-04-17","year":2025,"venue":"Proceedings of the 33rd ACM International Conference on Multimedia","status":"published","topics":["cs.CV"],"abstract":"Reconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gaussian Splatting (3DGS). These methods often encounter a transparency-depth dilemma, where the pursuit of photorealistic rendering through standard $α$-blending undermines geometric precision, resulting in considerable depth estimation errors for transparent materials. To address this issue, we introduce Transparent Surface Gaussian Splatting (TSGS), a new framework that separates geometry learning from appearance refinement. In the geometry learning stage, TSGS focuses on geometry by using specular-suppressed inputs to accurately represent surfaces. In the second stage, TSGS improves visual fidelity through anisotropic specular modeling, crucially maintaining the established opacity to ensure geometric accuracy. To enhance depth inference, TSGS employs a first-surface depth extraction method. This technique uses a sliding window over $α$-blending weights to pinpoint the most likely surface location and calculates a robust weighted average depth. To evaluate the transparent surface reconstruction task under realistic conditions, we collect a TransLab dataset that includes complex transparent laboratory glassware. Extensive experiments on TransLab show that TSGS achieves accurate geometric reconstruction and realistic rendering of transparent objects simultaneously within the efficient 3DGS framework. Specifically, TSGS significantly surpasses current leading methods, achieving a 37.3% reduction in chamfer distance and an 8.0% improvement in F1 score compared to the top baseline. The code and dataset are available at https://longxiang-ai.github.io/TSGS/.","identifiers":{"arxiv":"2504.12799","doi":"10.48550/arXiv.2504.12799"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2504.12799"},{"label":"PDF","url":"https://arxiv.org/pdf/2504.12799"},{"label":"Code","url":"https://github.com/longxiang-ai/TSGS"},{"label":"Dataset","url":"https://huggingface.co/datasets/Longxiang-ai/TransLab"},{"label":"DOI","url":"https://doi.org/10.1145/3746027.3754548"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2504.12799v1"},{"label":"Publisher version","url":"https://doi.org/10.1145/3746027.3754548"},{"label":"DOI 10.1145/3746027.3754548","url":"https://doi.org/10.1145/3746027.3754548"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"1 Zhejiang University, 2 Zhongguancun Academy, Beijing, 3 Xi'an Jiaotong University, 4 Beijing Normal University","source":"Official GitHub project README","sourceUrl":"https://github.com/longxiang-ai/TSGS"},{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhejiang University, Hangzhou, Zhejiang, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3746027.3754548"}],"sources":["arXiv","Official GitHub project README","Crossref"],"id":"doi-10-48550-arxiv-2504-12799","updatedAt":"2026-08-08T12:09:53Z"},{"id":"doi-10-1109-tkde-2025-3545948","type":"article","title":"A Universal Pre-Training and Prompting Framework for General Urban Spatio-Temporal Prediction","authors":[{"name":"Yuan Yuan"},{"name":"Jingtao Ding"},{"name":"Jie Feng","institutions":["zgca"]},{"name":"Depeng Jin"},{"name":"Yong Li"}],"institutions":["zgca"],"rawAffiliations":["Beijing Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2025-02-28","year":2025,"venue":"IEEE Transactions on Knowledge and Data Engineering","status":"published","topics":["Urban Computing","Spatio-Temporal Prediction","Foundation Models","Prompting"],"abstract":"UniST is a universal model for urban spatio-temporal prediction across grid- and graph-based scenarios. It combines diverse pre-training data with knowledge-guided prompts and is evaluated on more than twenty scenarios, including few-shot and zero-shot settings.","identifiers":{"doi":"10.1109/TKDE.2025.3545948"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/TKDE.2025.3545948"},{"label":"DBLP","url":"https://dblp.org/rec/journals/tkde/YuanDFJL25"},{"label":"Code","url":"https://github.com/tsinghua-fib-lab/UniST"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/TKDE.2025.3545948"}],"evidence":[{"level":"structured","institution":"zgca","matchedText":"Beijing Zhongguancun Academy","source":"Scholarly metadata record","sourceUrl":"https://www.researchgate.net/publication/389369039_A_Universal_Pre-training_and_Prompting_Framework_for_General_Urban_Spatio-Temporal_Prediction"}],"sources":["OpenAlex","Crossref","DBLP"],"updatedAt":"2026-08-08T00:00:00Z"},{"type":"conference","title":"DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration","authors":[{"name":"Hebaixu Wang","institutions":["zgca"]},{"name":"Jing Zhang"},{"name":"Haonan Guo"},{"name":"Di Wang"},{"name":"Jiayi Ma"},{"name":"Bo Du"}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy","The University of Sydney","Wuhan University","Nanyang Technological University"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Advances in Neural Information Processing Systems 38","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.52202/085713-5624"},"links":[{"label":"DOI","url":"https://doi.org/10.52202/085713-5624"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.52202/085713-5624"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.52202/085713-5624"}],"sources":["Crossref"],"id":"doi-10-52202-085713-5624","updatedAt":"2026-08-12T03:04:08Z"},{"type":"conference","title":"Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition","authors":[{"name":"Yu Li"},{"name":"Jin Jiang"},{"name":"Jianhua Zhu"},{"name":"Shuai Peng"},{"name":"Baole Wei","institutions":["zgci"]},{"name":"Yuxuan Zhou"},{"name":"Liangcai Gao"}],"institutions":["zgci"],"rawAffiliations":["Peking University","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Advances in Neural Information Processing Systems 38","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.52202/085713-4297"},"links":[{"label":"DOI","url":"https://doi.org/10.52202/085713-4297"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.52202/085713-4297"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.52202/085713-4297"}],"sources":["Crossref"],"id":"doi-10-52202-085713-4297","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products","authors":[{"name":"Yunyang Li"},{"name":"Lin Huang"},{"name":"Zhihao Ding"},{"name":"Xinran Wei"},{"name":"Chu Wang"},{"name":"Han Yang"},{"name":"Zun Wang"},{"name":"Chang Liu"},{"name":"Yu Shi"},{"name":"Peiran Jin"},{"name":"Tao Qin","institutions":["zgca"]},{"name":"Mark Gerstein"},{"name":"Jia Zhang"}],"institutions":["zgca"],"rawAffiliations":["Yale","BUPT","The Hong Kong Polytechnic University, Hong Kong Polytechnic University","Microsoft","ByteDance Inc.","Shanghai Artificial Intelligence Laboratory","Shanghai Jiaotong University","Microsoft Research","Zhongguancun Academy","Yale University"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Advances in Neural Information Processing Systems 38","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.52202/085713-0797"},"links":[{"label":"DOI","url":"https://doi.org/10.52202/085713-0797"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.52202/085713-0797"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.52202/085713-0797"}],"sources":["Crossref"],"id":"doi-10-52202-085713-0797","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization","authors":[{"name":"Yuantian Shao"},{"name":"Yuanteng Chen"},{"name":"Peisong Wang"},{"name":"Jinpu Yu"},{"name":"Jing Lin"},{"name":"Yiwu Yao"},{"name":"Zhihui Wei"},{"name":"Jianwei Cheng"}],"institutions":["zgca"],"rawAffiliations":["Yuanteng Chen 1 1 footnotemark: 1 Affiliation: C DL, Institute of Automation, Chinese Academy of Sciences, Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences, Affiliation: Zhongguancun Academy,"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Quantization plays a crucial role in accelerating the inference of large-scale models, and rotational matrices have been shown to effectively improve quantization performance by smoothing outliers. However, end-to-end fine-tuning of rotational optimization algorithms incurs high computational costs and is prone to overfitting. To address this challenge, we propose an efficient distribution-aware rotational calibration method, DartQuant, which reduces the complexity of rotational optimization by constraining the distribution of the activations after rotation. This approach also effectively reduces reliance on task-specific losses, thereby mitigating the risk of overfitting. Additionally, we introduce the QR-Orth optimization scheme, which replaces expensive alternating optimization with a more efficient solution. In a variety of model quantization experiments, DartQuant demonstrates superior performance. Compared to existing methods, it achieves 47$\\times$ acceleration and 10$\\times$ memory savings for rotational optimization on a 70B model. Furthermore, it is the first to successfully complete rotational calibration for a 70B model on a single 3090 GPU, making quantization of large language models feasible in resource-constrained environments. Code is available at https://github.com/CAS-CLab/DartQuant.git.","identifiers":{"arxiv":"2511.04063","doi":"10.48550/arXiv.2511.04063","publishedDoi":"10.52202/085713-4819"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2511.04063"},{"label":"HTML","url":"https://arxiv.org/html/2511.04063v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2511.04063"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2511.04063v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Yuanteng Chen 1 1 footnotemark: 1 Affiliation: C DL, Institute of Automation, Chinese Academy of Sciences, Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences, Affiliation: Zhongguancun Academy,","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2511.04063v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2511-04063","updatedAt":"2026-08-13T08:02:46Z"},{"type":"preprint","title":"Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL","authors":[{"name":"Ronghui Wu"},{"name":"Yifan Zhao"},{"name":"Guangyao Chen"},{"name":"Jia Li"}],"institutions":["zgca"],"rawAffiliations":["2 Zhongguancun Academy 3 Peking University Corresponding authors."],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Few-Shot Class-Incremental Learning (FSCIL) challenges models to sequentially learn new classes from minimal examples without forgetting prior knowledge, a task complicated by the stability-plasticity dilemma and data scarcity. Current FSCIL methods often struggle with generalization due to their reliance on limited datasets. While diffusion models offer a path for data augmentation, their direct application can lead to semantic misalignment or ineffective guidance. This paper introduces Diffusion-Classifier Synergy (DCS), a novel framework that establishes a mutual boosting loop between diffusion model and FSCIL classifier. DCS utilizes a reward-aligned learning strategy, where a dynamic, multi-faceted reward function derived from the classifier's state directs the diffusion model. This reward system operates at two levels: the feature level ensures semantic coherence and diversity using prototype-anchored maximum mean discrepancy and dimension-wise variance matching, while the logits level promotes exploratory image generation and enhances inter-class discriminability through confidence recalibration and cross-session confusion-aware mechanisms. This co-evolutionary process, where generated images refine the classifier and an improved classifier state yields better reward signals, demonstrably achieves state-of-the-art performance on FSCIL benchmarks, significantly enhancing both knowledge retention and new class learning.","identifiers":{"arxiv":"2510.03608","doi":"10.48550/arXiv.2510.03608","publishedDoi":"10.52202/085713-2153"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2510.03608"},{"label":"HTML","url":"https://arxiv.org/html/2510.03608v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2510.03608"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2510.03608v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"2 Zhongguancun Academy 3 Peking University Corresponding authors.","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2510.03608v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2510-03608","updatedAt":"2026-08-13T06:56:52Z"},{"type":"preprint","title":"GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution","authors":[{"name":"Fengxiang Wang"},{"name":"Mingshuo Chen"},{"name":"Yueying Li"},{"name":"Di Wang"},{"name":"Haotian Wang"},{"name":"Zonghao Guo"},{"name":"Zefan Wang"},{"name":"Shan Boqi"},{"name":"Long Lan"},{"name":"Yulin Wang"},{"name":"Hongzhen Wang"},{"name":"Wenjing Yang"},{"name":"Boxue Du"},{"name":"Jing Zhang"}],"institutions":["zgca"],"rawAffiliations":["4 School of Computer Science, Wuhan University, China 5 Zhongguancun Academy, China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce SuperRS-VQA (avg. 8,376$\\times$8,376) and HighRS-VQA (avg. 2,000$\\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: Background Token Pruning and Anchored Token Selection, to reduce the memory footprint while preserving key semantics.Integrating these techniques, we introduce GeoLLaVA-8K, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench.","identifiers":{"arxiv":"2505.21375","doi":"10.48550/arXiv.2505.21375","publishedDoi":"10.52202/085713-5317"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2505.21375"},{"label":"HTML","url":"https://arxiv.org/html/2505.21375v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2505.21375"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2505.21375v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"4 School of Computer Science, Wuhan University, China 5 Zhongguancun Academy, China","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2505.21375v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2505-21375","updatedAt":"2026-08-08T18:30:46Z"},{"type":"preprint","title":"STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible Benchmarking","authors":[{"name":"Sicheng Shen"},{"name":"Dongcheng Zhao"},{"name":"Linghao Feng"},{"name":"Zeyang Yue"},{"name":"Jindong Li"},{"name":"T. Li"},{"name":"Guobin Shen"},{"name":"Yi Zeng"}],"institutions":["zgca"],"rawAffiliations":["4 Zhongguancun Academy 5 Beihang University"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"arXiv","status":"preprint","topics":[],"abstract":"Spiking Transformers have recently emerged as promising architectures for combining the efficiency of spiking neural networks with the representational power of self-attention. However, the lack of standardized implementations, evaluation pipelines, and consistent design choices has hindered fair comparison and principled analysis. In this paper, we introduce STEP a unified benchmark framework for Spiking Transformers that supports a wide range of tasks, including classification, segmentation, and detection across static, event-based, and sequential datasets. STEP provides modular support for diverse components such as spiking neurons, input encodings, surrogate gradients, and multiple backends (e.g., SpikingJelly, BrainCog). Using STEP, we reproduce and evaluate several representative models, and conduct systematic ablation studies on attention design, neuron types, encoding schemes, and temporal modeling capabilities. We also propose a unified analytical model for energy estimation, accounting for spike sparsity, bitwidth, and memory access, and show that quantized ANNs may offer comparable or better energy efficiency. Our results suggest that current Spiking Transformers rely heavily on convolutional frontends and lack strong temporal modeling, underscoring the need for spike-native architectural innovations. The full code is available at: https://github.com/Fancyssc/STEP","identifiers":{"arxiv":"2505.11151","doi":"10.48550/arXiv.2505.11151","publishedDoi":"10.52202/085713-5334"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2505.11151"},{"label":"HTML","url":"https://arxiv.org/html/2505.11151v1"},{"label":"PDF","url":"https://arxiv.org/pdf/2505.11151"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2505.11151v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"4 Zhongguancun Academy 5 Beihang University","source":"arXiv HTML v1","sourceUrl":"https://arxiv.org/html/2505.11151v1"}],"sources":["arXiv","arXiv HTML"],"id":"doi-10-48550-arxiv-2505-11151","updatedAt":"2026-08-08T17:52:29Z"},{"type":"preprint","title":"Integrating World Models into Vision Language Action and Navigation: A Comprehensive Survey","authors":[{"name":"Jingwen Sun","institutions":["zgca"]},{"name":"Hongjin Chen","institutions":["zgca"]},{"name":"Zezhi Liu","institutions":["zgca"]},{"name":"Xin Jin","institutions":["zgca"]},{"name":"Wenyao Zhang"},{"name":"Tonghua Su"},{"name":"Chen Gao","institutions":["zgca"]},{"name":"Zhibo Chen","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["University of Science and Technology of China","Zhongguancun Academy","Harbin Institute of Technology","Nankai University","Eastern Institute of Technology","Shanghai Jiao Tong University","Tsinghua University"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Crossref","status":"published","topics":[],"abstract":"World models are a transformative paradigm in embodied AI, enabling agents to learn efficiently and plan by simulating environmental dynamics. As vision-language action and navigation systems grow more sophisticated, integrating world models is crucial for bridging multimodal perception, language understanding, and sequential decision-making. However, existing surveys, which focus on application domains or learning paradigms, have overlooked a critical architectural question: how world models should be structurally integrated into these systems. To address this gap, we introduce an integration-centric taxonomy that classifies research into three fundamental architectural paradigms: (1) Modular Architectures, where world models and policies are distinct modules; (2) Sequential Architectures, implementing hierarchical plan-then-execute workflows; and (3) Unified Architectures, which fuse world prediction and action generation into an end-to-end network. Through this framework, we systematically analyze the inherent trade-offs across these paradigms: modular designs excel in interpretability, sequential approaches enable hierarchical reasoning, and unified models achieve tight prediction-control coordination. Based on this analysis, we outline promising research directions and provide architectural principles for effectively integrating world models, ultimately to advance more efficient, interpretable, and generalizable agents.","identifiers":{"doi":"10.36227/techrxiv.176531987.77979037/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.36227/techrxiv.176531987.77979037/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.36227/techrxiv.176531987.77979037/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.36227/techrxiv.176531987.77979037/v1"}],"sources":["Crossref"],"id":"doi-10-36227-techrxiv-176531987-77979037-v1","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Model-Based Closed-Loop Control Algorithm for Stochastic Partial Differential Equation Control","authors":[{"name":"Peiyan Hu"},{"name":"Haodong Feng"},{"name":"Yue Wang","institutions":["zgca"]},{"name":"Zhiming Ma"}],"institutions":["zgca"],"rawAffiliations":["Academy of Mathematics and Systems Science","Westlake University","Zhongguancun Academy"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence","status":"published","topics":[],"abstract":"Neural operators have demonstrated promise in modeling and controlling systems governed by Partial Differential Equations (PDEs). Beyond PDEs, Stochastic Partial Differential Equations (SPDEs) play a critical role in modeling systems influenced by randomness, with applications in finance, physics, and beyond. However, controlling SPDE-governed systems remains a significant challenge. On the one hand, the regularity of the system's state (which can be intuitively understood as smoothness) deteriorates, making modeling and generalization more challenging. On the other hand, this stochasticity also renders control more unstable and thus less accurate. To address this gap, we propose the Model-Based Closed-Loop Control Algorithm (MB-CC), the first model-based closed-loop control method for SPDEs. MB-CC introduces two key innovations to enhance control robustness and efficiency: a Regularity Feature (RF) block and a closed-loop strategy with an operator-encoded policy network. The RF block, inspired by the regularity structure theory of SPDEs, addresses noise-induced irregularities by transforming the network's input—including the system state and noise-perturbed external forces—into a refined feature space for improved forward prediction. Compared to previous works using regularity features, we introduce a new parameterization, data augmentation, and extend the RF block as a plug-and-play component. Additionally, to achieve closed-loop control, we introduce an operator-encoded policy network to map the current state to optimal control, which integrates physical priors and swiftly makes decisions based on states returned by the environment. We conduct a systematic evaluation of MB-CC on two notable SPDEs, showcasing its effectiveness and efficiency. The ablation studies show its ability to handle stochasticity more effectively.","identifiers":{"doi":"10.24963/ijcai.2025/599"},"links":[{"label":"DOI","url":"https://doi.org/10.24963/ijcai.2025/599"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.24963/ijcai.2025/599"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy","source":"Crossref","sourceUrl":"https://doi.org/10.24963/ijcai.2025/599"}],"sources":["Crossref"],"id":"doi-10-24963-ijcai-2025-599","updatedAt":"2026-08-12T03:04:08Z"},{"type":"preprint","title":"AI-Enhanced Subseasonal Forecasting of Extreme Temperature Risks","authors":[{"name":"Jia Xing"},{"name":"Siwei Li"},{"name":"Shuxin Zheng","institutions":["zgci"]},{"name":"Ge Song"},{"name":"Jiaxin Dong"},{"name":"Gonzalo Ferrada"},{"name":"Tie-Yan Liu","institutions":["zgci"]},{"name":"Joshua Fu"}],"institutions":["zgci"],"rawAffiliations":["University of Tennessee at Knoxville","Wuhan University","Zhongguancun Institute of Artificial Intelligence"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Crossref","status":"published","topics":[],"abstract":"Abstract Sub-seasonal weather prediction remains a significant scientific challenge due to the chaotic nature of the atmosphere, with current numerical and AI-driven models exhibiting limited skill, particularly at the fine spatial scales for human exposure, agriculture, and infrastructure. Here, we introduce DeepMet, a high-resolution, AI-driven sub-seasonal forecasting system designed to improve the prediction of temperature extremes and their associated health risks, demonstrated successfully over the continental United States. Specifically, DeepMet substantially outperforms the benchmark of European Centre for Medium-Range Weather Forecasts, reducing the root mean square error by 20–60% for key surface variables, including daily maximum and minimum 2-meter temperature, specific humidity, and 10-meter wind speed. The model also improves the detection of extreme heat and cold events by over 40% across all evaluation metrics. By enhancing early warning capabilities, DeepMet enables more accurate identification of extreme weather conditions, potentially improving risk communication to prevent additional extreme-weather related deaths in the United States. Remarkably, such performance is achieved using only a single GPU for training, making the method highly accessible for local agencies to enhance early warning systems and protect public health. This underscores its strong potential to transform long-range forecasting and significantly enhance public health preparedness in a changing climate.","identifiers":{"doi":"10.21203/rs.3.rs-7314380/v1"},"links":[{"label":"DOI","url":"https://doi.org/10.21203/rs.3.rs-7314380/v1"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.21203/rs.3.rs-7314380/v1"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.21203/rs.3.rs-7314380/v1"}],"sources":["Crossref"],"id":"doi-10-21203-rs-3-rs-7314380-v1","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation","authors":[{"name":"Xin Lu","institutions":["zgca"]},{"name":"Chuanqing Zhuang"},{"name":"Chenxi Jin"},{"name":"Zhengda Lu"},{"name":"Yiqun Wang"},{"name":"Wu Liu"},{"name":"Jun Xiao","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["University of Chinese Academy of Sciences, Beijing, China and Zhongguancun Academy, Beijing, China","University of Chinese Academy of Sciences, Beijing, China","National university of Singapore, Singapore, Singapore","Chongqing University, Chongqing, China","University of Science and Technology of China, Hefei, China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Proceedings of the SIGGRAPH Asia 2025 Conference Papers","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3757377.3763887"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3757377.3763887"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3757377.3763887"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"University of Chinese Academy of Sciences, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3757377.3763887"}],"sources":["Crossref"],"id":"doi-10-1145-3757377-3763887","updatedAt":"2026-09-08T04:58:29Z"},{"type":"conference","title":"RATopo: Improving Lane Topology Reasoning via Redundancy Assignment","authors":[{"name":"Han Li","institutions":["zgca"]},{"name":"Shaofei Huang"},{"name":"Longfei Xu"},{"name":"Yulu Gao"},{"name":"Beipeng Mu"},{"name":"Si Liu"}],"institutions":["zgca"],"rawAffiliations":["School of Artificial Intelligence, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","Faculty of Science and Technology, University of Macau, Macau, China","School of Computer Science and Engineering, Beihang University, Beijing, China","School of Artificial Intelligence, Beihang University, Beijing, China and Hangzhou International Innovation Institute, Beihang University, Hangzhou, China","Meituan, Beijing, China","School of Artificial Intelligence, Beihang University, Beijing, China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Proceedings of the 33rd ACM International Conference on Multimedia","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3746027.3755845"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3746027.3755845"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3746027.3755845"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"School of Artificial Intelligence, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3746027.3755845"}],"sources":["Crossref"],"id":"doi-10-1145-3746027-3755845","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Try Harder: Hard Sample Generation and Learning for Cloth-Changing Person Re-ID","authors":[{"name":"Hankun Liu"},{"name":"Yujian Zhao","institutions":["zgca"]},{"name":"Guanglin Niu"}],"institutions":["zgca"],"rawAffiliations":["School of Computer Science and Engineering, Beihang University, Beijing, China","School of Artifical Intelligence, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","School of Artificial Intelligence, Beihang University, Beijing, China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Proceedings of the 33rd ACM International Conference on Multimedia","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1145/3746027.3755394"},"links":[{"label":"DOI","url":"https://doi.org/10.1145/3746027.3755394"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1145/3746027.3755394"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"School of Artifical Intelligence, Beihang University, Beijing, China and Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1145/3746027.3755394"}],"sources":["Crossref"],"id":"doi-10-1145-3746027-3755394","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation","authors":[{"name":"Yuanhe Tian","institutions":["zgci"]},{"name":"Lei Mao"},{"name":"Yan Song"}],"institutions":["zgci"],"rawAffiliations":["Zhongguancun Institute of Artificial Intelligence","Origin Omics","University of Science and Technology of China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/bibm66473.2025.11356692"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/bibm66473.2025.11356692"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/bibm66473.2025.11356692"}],"evidence":[{"level":"exact-affiliation","institution":"zgci","matchedText":"Zhongguancun Institute of Artificial Intelligence","source":"Crossref","sourceUrl":"https://doi.org/10.1109/bibm66473.2025.11356692"}],"sources":["Crossref"],"id":"doi-10-1109-bibm66473-2025-11356692","updatedAt":"2026-08-08T11:35:31Z"},{"type":"conference","title":"ManiNet: Manifold Network for Few-Shot Learning","authors":[{"name":"Rui-Qi Wang"},{"name":"Hengcan Shi"},{"name":"Yi Chen","institutions":["zgca"]},{"name":"Yaonan Wang"}],"institutions":["zgca"],"rawAffiliations":["University of Science and Technology Beijing,School of Intelligence Science and Technology Institute of Artificial Intelligence,Beijing,China","Hunan University,School of Artificial Intelligence and Robotics,Changsha,China","University of Chinese Academy of Sciences,State Key Laboratory of MAIS, CASIA Zhongguancun Academy,Beijing,China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR)","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1109/aihcir67580.2025.11405062"},"links":[{"label":"DOI","url":"https://doi.org/10.1109/aihcir67580.2025.11405062"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1109/aihcir67580.2025.11405062"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"University of Chinese Academy of Sciences,State Key Laboratory of MAIS, CASIA Zhongguancun Academy,Beijing,China","source":"Crossref","sourceUrl":"https://doi.org/10.1109/aihcir67580.2025.11405062"}],"sources":["Crossref"],"id":"doi-10-1109-aihcir67580-2025-11405062","updatedAt":"2026-08-08T11:35:31Z"},{"type":"preprint","title":"FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling","authors":[{"name":"Jianwei Zhu","institutions":["zgca"]},{"name":"Yu Shi","institutions":["zgca"]},{"name":"Ran Bi","institutions":["zgca"]},{"name":"Peiran Jin","institutions":["zgca"]},{"name":"Chang Liu","institutions":["zgca"]},{"name":"Zhe Zhang","institutions":["zgca"]},{"name":"Haitao Huang","institutions":["zgca"]},{"name":"Zekun Guo","institutions":["zgca"]},{"name":"Pipi Hu"},{"name":"Fusong Ju","institutions":["zgca"]},{"name":"Lin Huang"},{"name":"Xinwei Tai","institutions":["zgca"]},{"name":"Chenao Li"},{"name":"Kaiyuan Gao"},{"name":"Xinran Wei","institutions":["zgca"]},{"name":"Huanhuan Xia","institutions":["zgca"]},{"name":"Jia Zhang"},{"name":"Yaosen Min","institutions":["zgca"]},{"name":"Zun Wang"},{"name":"Yusong Wang","institutions":["zgca"]},{"name":"Liang He","institutions":["zgca"]},{"name":"Haiguang Liu","institutions":["zgca"]},{"name":"Tao Qin","institutions":["zgca"]}],"institutions":["zgca"],"rawAffiliations":["Zhongguancun Academy, Beijing, China","Hunan University, Changsha, China","Beijing Institute of Mathematical Sciences and Applications, Beijing, China","Ubiquant, Beijing, China","Huazhong University of Science and Technology, Wuhan, China","Institute of Biophysics, Chinese Academy of Sciences, Beijing, China","Shanghai Artificial Intelligence Laboratory, Shanghai, China"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"Crossref","status":"published","topics":[],"abstract":"A bstract Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-based language models (e.g., ESM-2) capture sequence semantics at scale, and a number of recent works incorporate structural signals into sequence encoders. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting evolutionary couplings, but their reliance on homologous sequences makes them less reliable in highly mutated or alignment-sparse regimes. We present FlexRibbon ‡ , a pretrained protein model that jointly learns from amino acid sequences and three-dimensional structures. Our pretraining strategy combines masked language modeling with diffusion-based denoising, enabling bidirectional sequence-structure learning without requiring MSAs. Trained on both experimentally resolved structures and AlphaFold 2 predictions, FlexRibbon captures global folds as well as flexible conformations critical for biological function. Evaluated across diverse tasks spanning interface design, intermolecular interaction prediction, and protein function prediction, FlexRibbon establishes new state-of-the-art performance on 12 different tasks, with particularly strong gains in mutation-rich settings where MSA-based methods often struggle.","identifiers":{"doi":"10.1101/2025.10.08.681293"},"links":[{"label":"DOI","url":"https://doi.org/10.1101/2025.10.08.681293"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1101/2025.10.08.681293"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy, Beijing, China","source":"Crossref","sourceUrl":"https://doi.org/10.1101/2025.10.08.681293"}],"sources":["Crossref"],"id":"doi-10-1101-2025-10-08-681293","updatedAt":"2026-08-08T11:35:31Z"},{"type":"article","title":"Neuromorphic spike-based large language model","authors":[{"name":"Han Xu"},{"name":"Xuerui Qiu","institutions":["zgca"]},{"name":"Yunhui Xu"},{"name":"Mohammed E Elbtity"},{"name":"Peng Zhou"},{"name":"Yang Tian"},{"name":"Rui-Jie Zhu"},{"name":"Jiahong Zhang"},{"name":"Shaowei Gu","institutions":["zgca"]},{"name":"Yuqi Pan"},{"name":"Yuhong Chou"},{"name":"Qinghao Wen"},{"name":"Man Yao"},{"name":"Jiangbo Qian"},{"name":"Yonghong Tian"},{"name":"Lei Ma"},{"name":"Tiejun Huang"},{"name":"Jason K Eshraghian"},{"name":"Bo Xu"},{"name":"Guoqi Li"}],"institutions":["zgca"],"rawAffiliations":["Institute of Automation, Chinese Academy of Sciences , Beijing 100190 ,","School of Artificial Intelligence, University of Chinese Academy of Sciences , Beijing 100049 ,","Brain-Inspired Models Group, Beijing Academy of Artificial Intelligence , Beijing 100862 ,","School of Future Technology, University of Chinese Academy of Sciences , Beijing 101408 ,","Zhongguancun Academy , Beijing 100094 ,","Department of Psychology, Tsinghua University , Beijing 100084 ,","Advanced Micro Devices Inc , Santa Clara 95054 ,","Department of Research, LuxiTech , Shenzhen 518055 ,","Faculty of Data Science, City University of Macau , Macau 999078 ,","Department of Electrical and Computer Engineering, University of California , Santa Cruz 95064 ,","Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University , Hong Kong 999077 ,","School of Aerospace, Mechanical and Mechatronic Engineering, The University of Sydney , Sydney 2006 ,","Faculty of Electrical Engineering and Computer Science, Ningbo University , Ningbo 315211 ,","School of Computer Science, Peking University , Beijing 100871 ,","Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology , Beijing 100190 ,","Spiking Intelligence Lab, Tianqiao & Chirssy Chen Institute , Shanghai 201203 ,"],"relationType":"affiliation","publishedAt":"2025-01-01","year":2025,"venue":"National Science Review","status":"published","topics":[],"abstract":"ABSTRACT This work proposes a unified neuromorphic spike-based large-language-model (NSLLM) framework to simultaneously address the challenges of high energy consumption and low interpretability in LLMs. Our framework transforms LLMs into efficient NSLLMs by converting their behaviors into neural dynamics—such as spike trains—through rigorous mathematical modeling and complemented by advanced techniques including quantization and sparsification. This transformation also enables the analysis of information encoding processes using computational neuroscience tools, thereby offering a novel neuroscientific perspective that conceptualizes LLMs as neural populations to enhance their interpretability. Leveraging a hardware-algorithm co-design paradigm, an NSLLM can completely eliminate matrix multiplication (MatMul) while maintaining high performance. We designed a custom MatMul-free hardware core on the VCK190 field-programmable gate array to validate the 1.5-billion-parameter NSLLM, achieving a dynamic power consumption of only 13.849 W and an inference throughput of 161.8 tokens per second. Compared with the A800 GPU, this implementation improves energy efficiency, memory usage and inference throughput by 19.8$\\times$, 21.3$\\times$ and 2.2$\\times$, respectively. This work provides a novel perspective within a unified framework to enhance both the energy efficiency and interpretability of LLMs, offering valuable insights for future neuromorphic chip designs tailored for large models.","identifiers":{"doi":"10.1093/nsr/nwaf551"},"links":[{"label":"DOI","url":"https://doi.org/10.1093/nsr/nwaf551"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1093/nsr/nwaf551"}],"evidence":[{"level":"exact-affiliation","institution":"zgca","matchedText":"Zhongguancun Academy , Beijing 100094 ,","source":"Crossref","sourceUrl":"https://doi.org/10.1093/nsr/nwaf551"}],"sources":["Crossref"],"id":"doi-10-1093-nsr-nwaf551","updatedAt":"2026-09-18T04:57:17Z"},{"type":"article","title":"AI-Driven Accelerated Discovery of High-Performance Perovskite Quantum Dots Via Predictive LightGBM Modeling","authors":[{"name":"Zhenze Zhao"},{"name":"Yanan Zhu"},{"name":"Gaoxu Li"},{"name":"Haojie Dong"},{"name":"Kai Chen"},{"name":"Yibo Zhang"},{"name":"Roman B. Vasiliev"},{"name":"Jiyan He"},{"name":"Na Liu"},{"name":"Shuai Chang"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）"],"relationType":"official-output","publishedAt":"2025-01-01","year":2025,"venue":"The Journal of Physical Chemistry Letters","status":"published","topics":[],"abstract":"","identifiers":{"doi":"10.1021/acs.jpclett.5c02902"},"links":[{"label":"DOI","url":"https://doi.org/10.1021/acs.jpclett.5c02902"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_f2t7awv1560v78jfkwtbbg69qgzppkzs"}],"versions":[{"label":"Publisher version","url":"https://doi.org/10.1021/acs.jpclett.5c02902"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_f2t7awv1560v78jfkwtbbg69qgzppkzs"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_f2t7awv1560v78jfkwtbbg69qgzppkzs"}],"sources":["Crossref","北京中关村学院官网"],"id":"doi-10-1021-acs-jpclett-5c02902","updatedAt":"2026-08-08T11:59:38Z"},{"type":"preprint","title":"The Value of AI-Generated Metadata for UGC Platforms: Evidence from a Large-scale Field Experiment","authors":[{"name":"Xinyi Zhang"},{"name":"Chenshuo Sun"},{"name":"Renyu Zhang"},{"name":"Khim-Yong Goh"}],"institutions":["zgca","zgci"],"rawAffiliations":["北京中关村学院与中关村人工智能研究院（中关村两院）"],"relationType":"official-output","publishedAt":"2024-12-24","year":2024,"venue":"arXiv","status":"preprint","topics":["econ.GN","cs.AI","cs.HC"],"abstract":"AI-generated content (AIGC), such as advertisement copy, product descriptions, and social media posts, is becoming ubiquitous in business practices. However, the value of AI-generated metadata, such as titles, remains unclear on user-generated content (UGC) platforms. To address this gap, we conducted a large-scale field experiment on a leading short-video platform in Asia to provide about 1 million users access to AI-generated titles for their uploaded videos. Our findings show that the provision of AI-generated titles significantly boosted content consumption, increasing valid watches by 1.6% and watch duration by 0.9%. When producers adopted these titles, these increases jumped to 7.1% and 4.1%, respectively. This viewership-boost effect was largely attributed to the use of this generative AI (GAI) tool increasing the likelihood of videos having a title by 41.4%. The effect was more pronounced for groups more affected by metadata sparsity. Mechanism analysis revealed that AI-generated metadata improved user-video matching accuracy in the platform's recommender system. Interestingly, for a video for which the producer would have posted a title anyway, adopting the AI-generated title decreased its viewership on average, implying that AI-generated titles may be of lower quality than human-generated ones. However, when producers chose to co-create with GAI and significantly revised the AI-generated titles, the videos outperformed their counterparts with either fully AI-generated or human-generated titles, showcasing the benefits of human-AI co-creation. This study highlights the value of AI-generated metadata and human-AI metadata co-creation in enhancing user-content matching and content consumption for UGC platforms.","identifiers":{"arxiv":"2412.18337","doi":"10.48550/arXiv.2412.18337"},"links":[{"label":"arXiv","url":"https://arxiv.org/abs/2412.18337"},{"label":"PDF","url":"https://arxiv.org/pdf/2412.18337"},{"label":"Official","url":"https://www.bza.edu.cn/detail/inews_8jt1d001mt4c3kzt090aknzlnitbtnz3"}],"versions":[{"label":"arXiv v1","url":"https://arxiv.org/abs/2412.18337v1"}],"evidence":[{"level":"official-listing","institution":"zgca","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_8jt1d001mt4c3kzt090aknzlnitbtnz3"},{"level":"official-listing","institution":"zgci","matchedText":"北京中关村学院与中关村人工智能研究院（中关村两院）","source":"北京中关村学院官网","sourceUrl":"https://www.bza.edu.cn/detail/inews_8jt1d001mt4c3kzt090aknzlnitbtnz3"}],"sources":["arXiv","北京中关村学院官网"],"id":"doi-10-48550-arxiv-2412-18337","updatedAt":"2026-08-12T03:04:08Z"}]
