Table of Contents
- cs.CL [Total: 44]
- cs.CV [Total: 161]
- cs.HC [Total: 3]
- cs.AI [Total: 15]
- cs.NE [Total: 1]
- cs.RO [Total: 11]
- eess.SP [Total: 1]
- cs.IR [Total: 1]
- cs.LG [Total: 6]
- cs.MM [Total: 1]
- cs.CR [Total: 1]
cs.CL [Back]
[1] Unified Hallucination Fuzzing for Multimodal Large Language Models cs.CL | cs.AIPDF
Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng
TL;DR: 本文针对多模态大语言模型(MLLMs)中的幻觉问题,提出了一个系统性的评估框架,该框架整合了细粒度基准数据集UniHall和自适应的压力测试方法SAMF。研究发现,现有最先进的MLLMs在模糊测试下性能显著下降,揭示了其推理能力与事实基础之间的脱节,并指出了强化学习对齐可能加剧指令跟随任务中的迎合行为。
Details
Motivation: 解决MLLMs中持续存在的幻觉问题,该问题严重限制了模型在高风险应用中的可靠性。现有基于静态基准的评估方法存在分类覆盖范围窄、性能快速饱和的缺陷,无法反映模型在动态现实场景中的鲁棒性。
Result: 在提出的评估框架下进行广泛实验,结果表明,与常规设置相比,最先进的MLLMs在模糊测试下表现出显著的性能退化。研究还识别出在指令跟随任务中存在有用性与幻觉之间的权衡。
Insight: 创新点包括:1)提出了一个统一的、细粒度的幻觉分类法(涵盖对象、指令和知识维度)及相应基准数据集UniHall;2)提出了自适应的多模态模糊测试框架SAMF,采用进化突变策略探索模型幻觉边界;3)引入了一个由多模态预言机集合驱动的结构化度量套件,用于可靠评估动态输入。从客观角度看,该工作将软件测试中的模糊测试思想系统性地引入MLLM评估,并强调了动态、自适应测试的重要性。
Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing. First, we introduce UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi-modal oracles. Our extensive experiments reveal that state-of-the-art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness-hallucination trade-off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction-following tasks. The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.
[2] DocAtlas: Long-Document Understanding as Mutable-State Interaction cs.CL | cs.AIPDF
Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai
TL;DR: DocAtlas将长文档理解建模为一个可变状态的信息寻求过程,通过一个可变的文档管理环境来动态控制信息的搜索、读取、存储和呈现。该系统结合了自改进检索、选择性证据访问和主动工作记忆,支持在固定上下文预算下使用大型视觉语言模型进行推理,或通过端到端强化学习训练紧凑型VLM代理。
Details
Motivation: 现有检索增强系统通常在生成前从静态索引中选择证据,而近期基于代理的系统增加了多轮工具使用,但往往依赖冻结的专有骨干模型且行为由提示设定。本文旨在解决长文档理解中需要跨多页、布局、表格、图表寻找和整合证据的挑战。
Result: 在MMLongBench-Doc基准测试中,使用GPT-5.4的DocAtlas达到71.4%,超过了65.8%的人类专家参考分数。通过在DocAtlas环境中进行端到端RL训练的Qwen3.5-4B VLM达到63.7%,显著优于54.4%的直接输入基线。
Insight: 核心创新在于将长文档理解视为一个可变状态交互过程,并设计了可变的文档管理环境作为外部环境来动态管理信息流。该方法将自改进检索、选择性访问和主动工作记忆统一在固定上下文预算下,为紧凑型文档代理提供了有效的训练和推理框架,显著提升了性能。
Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen proprietary backbones whose behavior is set by prompts. We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process. We instantiate DocAtlas as a mutable document harness: an external environment that determines what document information is searched, read, stored, reviewed, and shown to the model at each step. Given a document and question, the harness exposes search, reading, note-taking, and review tools, maintains a hierarchical tree and note store, and updates both as the agent records evidence. DocAtlas combines self-improving retrieval, selective evidence access, and active working memory under a fixed context budget. The same harness supports inference-time use with large VLMs and end-to-end reinforcement learning for compact VLM agents. With GPT-5.4, DocAtlas reaches 71.4% on MMLongBench-Doc, exceeding the human-expert reference of 65.8%. A Qwen3.5-4B VLM trained with end-to-end RL in the DocAtlas environment reaches 63.7%, compared with a 54.4% direct-input baseline, showing that mutable document-harness design can improve compact document agents by a large margin.
[3] WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management cs.CL | cs.AIPDF
Yi Zhang, Hongyang Wang, Zheng Hao Leong, Zihao Wu, Kaijun Lin
TL;DR: 本文提出了WuYuEval,一个用于评估大语言模型在固体废物管理领域能力的多层次基准。该基准包含基础模块和专家模块,涵盖从基础知识到复杂决策的任务。通过对33个LLM的评估,发现模型在计算、实验设计等复杂任务上表现显著下降,并指出结构化推理仅在锚定工程约束时有效。
Details
Motivation: 现有基准侧重于通用知识,难以评估LLM在固体废物管理这一专业领域中,面对工程、环境和政策约束时的专业决策能力。
Result: 在33个LLM的评估中,领先模型在基础模块准确率达到94.64%,但平均准确率从简单问题的84.14%下降到困难问题的42.50%。在专家模块的开放式任务上表现普遍较低。
Insight: 创新点在于构建了首个针对固体废物管理的多层次评估基准,并引入了结合锚点校准的LLM-as-a-Judge评分与基于Elo的成对比较方法。研究揭示了显式推理链的有效性依赖于其对单位、假设和工程约束的锚定,否则可能导致答案偏离边界,这为开发具有专业推理链的领域大模型提供了实证基础。
Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64% accuracy on the Foundation Module, but average accuracy still fell from 84.14% on easy questions to 42.50% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.
[4] Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards cs.CL | cs.AIPDF
Cheng Ruoxi, Ma Haoxuan, Zhang Hongyi, Zhang Junming, Duan Ranjie
TL;DR: 本文提出了Search-G1,一种基于表征的内在奖励框架,旨在提升搜索增强语言代理的接地性(grounding)与搜索效率的权衡。该框架通过两个经过干预校准的读出器来评估代理答案的操作性接地程度,从而在无需过程标注或LLM评判推理的情况下,为强化学习策略优化提供密集的、与接地性相关的奖励信号。
Details
Motivation: 现有外部奖励要么是稀疏的结果监督,要么是依赖昂贵标注或LLM推理的丰富反馈,难以有效区分必要的、接地的检索与冗余搜索。基于策略侧信号的内在奖励主要反映模型置信度而非证据接地性。
Result: 在多个基于搜索的问答基准测试和两种模型规模上的实验表明,Search-G1在保持任务准确率竞争力的同时,改善了接地性与搜索成本的权衡,产生了更短的响应侧轨迹。
Insight: 创新点在于提出了一种基于表征的、通过干预校准来测量操作性接地性的内在奖励框架,其两个读出器分别评估闭卷知识充分性(从而定义检索必要性)和答案对证据的依赖程度。该框架允许奖励信号与策略在优化过程中协同进化,无需外部过程监督。
Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent’s answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding–search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search-G1.
[5] Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation cs.CL | cs.CVPDF
Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn, Katarzyna Kowol, Karolina Piosek
TL;DR: 本文介绍了PoVisLE,一个针对波兰语设计的单文化视觉语言评测基准,旨在评估文化背景下的多模态理解能力。该数据集包含1,117张图像和2,366个人工标注的视觉问答对,通过基于情境的评估范式来测试模型对区域特定含义、符号内容和上下文视觉线索的深层理解。
Details
Motivation: 当前视觉语言模型主要基于英语数据训练,在处理文化相关的视觉理解方面存在局限,无法有效解释区域特定含义和上下文视觉线索,而现有文化能力评测基准多基于模板且关注表面识别,不足以评估文化情境下的深层语言和语用理解。
Result: 论文提出了PoVisLE数据集,作为一个受控且具有挑战性的资源,用于评估超越表面识别的文化背景视觉语言理解,但摘要中未提及具体定量结果或与现有模型的比较。
Insight: 创新点在于构建了一个针对波兰语文化的视觉语言评测基准,强调基于情境的评估范式,以促进对文化背景多模态理解的深层评测,这为评估非英语中心模型的跨文化能力提供了新思路。
Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.
[6] Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions cs.CL | cs.AIPDF
Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou
TL;DR: 本文提出了一种考试风格的评估框架,用于研究推理语言模型在共享推理计算(token预算)约束下如何跨多个问题分配计算资源。研究发现,现有模型无法根据问题难度和分值进行战略性预算分配,而是表现出贪婪的顺序求解行为,优先处理呈现顺序靠前的问题,且对分值不敏感。
Details
Motivation: 现有推理模型评估通常逐个问题地研究测试时计算,但在实际应用中,多个问题往往共享端到端的成本或延迟约束,模型需要决定如何在有限的计算资源下分配推理计算以最大化总分。
Result: 在多个开源和前沿推理模型上的实验表明,模型无法在难度和分值不同的问题间进行战略性预算分配,其行为类似于贪婪的顺序求解器,且随着问题数量增加,这种倾向更加明显。显式的规划提示虽能更均匀地分配计算,但无法实现基于价值或难度的优先级排序。该行为模式从数学推理扩展到代码推理任务。
Insight: 论文的创新点在于提出了一个全局预算分配的评估框架,揭示了推理模型在跨问题资源分配能力上的不足,这是一个未被传统逐问题评估捕获的、对当前模型构成挑战的独特能力。
Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must decide how to divide limited inference compute among them. We introduce an exam-style evaluation framework for studying this setting, in which a model must distribute one shared token budget across questions with different difficulty and point values to maximize its total score. Across several open and frontier reasoning models, we find that models fail to allocate a shared budget strategically across questions of varying difficulties and values. Models behave largely as greedy sequential solvers: they prioritize questions by presentation order, front-load effort on early questions, and remain insensitive to value, with these tendencies becoming more pronounced as the number of questions grows. Explicit planning prompts spread compute more evenly but do not produce value- or difficulty-aware prioritization. The same behavioral pattern extends from mathematical to code reasoning. These findings establish global budget allocation as a distinct capability that is not captured by conventional per-question evaluation and remains a challenge for current reasoning models.
[7] DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects cs.CL | cs.AIPDF
Yi Shu, Tianyu Peng, Yingzhuo Deng, Wen Yang, Jun Lin
TL;DR: 论文提出了DialectS2S,一个面向低资源中文方言的端到端语音对话模型。它通过一个可扩展的方言语音对话合成流水线来高效构建数据,并引入了一种带有自对齐语音监督的两阶段后训练策略,以提升方言语音生成质量。
Details
Motivation: 当前端到端语音对话模型主要针对主流语言优化,在低资源方言场景下因数据稀缺而受限;且在方言适应过程中,模型的语义表示空间不断演变,而传统的语音监督保持不变,导致隐藏表示与语音目标之间的语义不一致,降低了语音的稳定性和自然度。
Result: 实验结果表明,DialectS2S在多种中文方言的语音对话任务中持续优于现有基线,在方言一致性、响应质量和语音清晰度方面取得了显著提升。
Insight: 创新点在于提出了一个可扩展的方言语音数据合成流水线,以及一种自对齐的语音监督后训练策略,该策略将语音监督的语义内容与模型演变的语义表示对齐,从而改善了低资源方言场景下的语音生成质量。论文还开源了框架、模型、数据和代码,促进了相关研究和应用。
Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.
[8] Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs cs.CLPDF
Rama Alomair, Remas Alsubaie, Walaa Saifalislam, Rima Alsonbul, Mona Alnajjar
TL;DR: 本文研究了多语言训练在谚语比喻语言识别中的作用,通过引入一个多维注释框架(包括隐喻、道德/建议、因果和文化特定四种比喻形式),评估了五种模型在不同多语言监督水平下的表现。研究发现,约50%的翻译多语言训练数据即可达到接近最优的性能,且结合多种比喻形式能带来最强整体表现,其中文化特定形式在多语言监督下提升最大,而道德/建议和文化特定形式对指令调优大语言模型的性能贡献显著。
Details
Motivation: 研究动机是超越语言同质训练数据,明确翻译多语言监督在比喻语言识别中的贡献,以推动该领域从以隐喻为中心的分类转向概念级多维框架。
Result: 在包含7种语言、6,787个翻译实例的谚语数据集上,实验表明约50%的多语言训练数据可达到接近最优性能;结合多种比喻形式(特别是道德/建议和文化特定形式)能提升指令调优大语言模型的识别效果,文化特定形式在多语言监督下增益最大。
Insight: 创新点在于提出了一个多维注释框架来表征谚语的互补比喻形式,并实证发现多语言训练中数据效率和形式多样性的重要性,为比喻语言识别提供了更全面的概念级建模思路。
Abstract: Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We examine this question using 742 proverb concepts across 6,787 translated instances in seven languages. We evaluate five models, including multilingual encoders and instruction tuned LLMs, under progressively increasing levels of multilingual supervision. Moreover, we introduce a multidimensional annotation framework for proverbs that characterizes them through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture Specific. Our findings show that approximately 50% of the translated multilingual training data is sufficient to achieve near-optimal figurative language identification performance. We further show that combining diverse figurative forms yields the strongest overall performance. A notable finding is that the least frequent figurative form, Culture Specific, exhibits the largest performance gains under multilingual supervision. Furthermore, the Moral/Advisory and Culture Specific forms contribute most to the performance of instruction-tuned LLMs on figurative language identification. These findings motivate multilingual figurative language identification to move beyond metaphor-centric taxonomies toward concept level multidimensional frameworks that explicitly model complementary forms of figurative meaning.
[9] Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models cs.CLPDF
Xuning He, Zinan Sheng, Yongding Tao, Huanyu Liu, Ge Li
TL;DR: 本文提出了Archer方法,一种无需训练、用于支持回滚的扩散语言模型(DLMs)的键值缓存技术。该方法通过非对称地保持可变响应与当前假设同步,同时有界地重用提示的键值状态,从而在不缓存可变响应状态的情况下分摊重复的提示计算,并延迟来自试探性令牌的反馈,减少对瞬时高置信度错误的过早强化,为回滚提供更多纠正机会。
Details
Motivation: 扩散语言模型(DLMs)具有迭代细化和回滚能力,这使其区别于不可逆的自回归生成,但也导致推理成本高昂。每次去噪更新都会改变全局上下文,迫使提示和响应状态都需要重新计算,尽管只有响应令牌是可修改的。传统的键值缓存假设历史状态不可变,难以与回滚机制协调,因此需要一种新的缓存方法来降低DLMs的推理成本。
Result: 在主要测试套件上,Archer方法取得了33.63%的最佳平均性能,并实现了2.57倍的平均加速。在评估的各种设置中,它将Pass@1指标提升了最高3.05个百分点,并达到了最高2.95倍的加速。
Insight: 论文的核心创新点在于提出了“有界状态邻域内重用提示键值”的非对称缓存策略,将提示重用概念化为与可逆性对齐的缓存边界,并分析了其状态相关的近似误差界限。该方法通过延迟提示反馈来获得质量提升,并通过状态感知的刷新来保持完全刷新决策,在加速的同时提升了生成质量,而非以质量换取速度。
Abstract: Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. Every denoising update alters the global context, forcing both prompt and response states to be recomputed even though only response tokens are revisable. Key-value (KV) caching could reduce this cost, yet conventional caching assumes immutable historical states and is therefore difficult to reconcile with rollback.In this paper, we introduce Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer), a training-free KV caching method for rollback-capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Although prompt representations also change under bidirectional attention, their token identities remain fixed; bounded reuse therefore amortizes repeated prompt computation without caching mutable response states. It also delays feedback from tentative tokens, reducing premature reinforcement of transient high-confidence errors and giving rollback more opportunity to correct them. Our analysis characterizes prompt reuse as a reversibility-aligned cache boundary, bounds its state-dependent approximation error, and gives a decoder-margin condition for preserving full-refresh decisions.Existing DLM acceleration often trades quality for speed. Archer shifts this frontier, attaining the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite. Across evaluated settings, it improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup. Controlled analyses connect the quality gain to delayed prompt feedback and validate state-aware refresh. Our code is available at https://github.com/Hxnng/Archer.
[10] NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs cs.CL | cs.AIPDF
Jiayue Jin, Jingwei Zhang, Chen Wang, Jing Liu, Longteng Guo
TL;DR: 本文提出NeuPAT框架,通过分析预训练LLM神经元在多模态学习中的异质性可塑性,在指令微调中为不同神经元分配差异化更新约束,以保护语言敏感神经元并促进多模态适应,从而在扩展多模态能力的同时最大程度保留语言智能。
Details
Motivation: 解决多模态大语言模型在扩展感知能力时,通常会损害预训练阶段获得的语言能力的问题。
Result: 在多种LLM架构上的实验表明,NeuPAT在11个语言基准测试上恢复了94.5%由标准微调造成的语言能力退化,同时保持了可比的多模态性能。
Insight: 核心创新点在于揭示了预训练LLM神经元在多模态学习中的异质性可塑性,并据此提出了轻量级、架构无关的神经元感知可塑性分配调优框架,实现了能力保留式的多模态扩展。
Abstract: Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective of internal adaptation dynamics and discover that neurons in pretrained LLMs exhibit heterogeneous plasticity during multimodal learning: some neurons are critical for preserving language capabilities, while others are more adaptive to multimodal knowledge. Based on this insight, we propose NeuPAT (Neuron-aware Plasticity Allocation Tuning), a lightweight and architecture-agnostic framework that allocates neuron-wise update constraints during multimodal instruction tuning. NeuPAT uses a small-scale probing stage to estimate neuron adaptation patterns and selectively protects language-sensitive neurons while promoting multimodal adaptation through more plastic neurons. Experiments across diverse LLM families demonstrate that NeuPAT recovers 94.5% of the language capability degradation caused by vanilla tuning on 11 language benchmarks while maintaining comparable multimodal performance, providing an effective approach for capability-preserving multimodal expansion.
[11] STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs cs.CL | cs.AIPDF
Nuthakki Siva Gopala Krishna, Kanishka Jain
TL;DR: 本文提出STEMMA框架,这是一个多模态多代理系统,用于评估大型语言模型在自我身份表示上的一致性。研究发现,知识蒸馏过程中学生模型不仅学习功能知识,还继承了教师模型的行为模式,可能导致输出同质化、偏见和问责问题。
Details
Motivation: 随着知识蒸馏在LLM训练中规模与复杂度提升,需探究学生模型从教师模型真正继承的知识类型,特别是行为模式与自我身份表示,这关系到模型偏见与问责。
Result: 研究通过手动设计的对抗性提示评估多个模型,结果表明大多数模型在自我表示上存在一定程度的不一致性。
Insight: 创新点在于提出多代理协作框架STEMMA来系统性评估LLM的身份一致性,并引入对抗性提示集,为模型行为分析提供了新工具。
Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.
[12] Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders cs.CLPDF
Bo Cheng, Qiaolin Lu, Yi Chang, Yuan Wu
TL;DR: 本文通过Top-K稀疏自编码器(SAEs)分析DeepSeek-R1-Distill-Qwen-7B模型在数学推理任务中的中间表示,对比了其使用思维链(Thinking模式)与直接生成答案(NoThinking模式)的神经机制差异。研究发现,Thinking模式依赖稀疏且高强度的特征激活进行语言演绎,而NoThinking模式则表现出自适应、弥散的模式,侧重于符号操作。通过抑制最活跃的稀疏特征,揭示了推理与句法结构紧密耦合、Thinking模式受干扰后会产生补偿性过度生成、以及连贯思维链行为依赖于脆弱特征协调等原则。
Details
Motivation: 尽管采用思维链(CoT)的大语言模型展现出卓越的推理能力,但区分这种显式Thinking模式与直接生成答案(NoThinking模式)的神经机制仍不明确。本文旨在解构这一认知过程,理解模型在不同推理模式下的行为差异。
Result: 研究在三个不同难度级别的数学解题任务上进行了观察和因果干预实验。通过抑制三个最活跃的稀疏特征(总激活量干预),发现干预会一致性地破坏LaTeX和框定解决方案的格式,并导致Thinking模式产生补偿性过度生成(表现为元认知线索增加和重复性低信息延续),而输出结构则持续受损。
Insight: 论文的创新点在于应用稀疏自编码器来解耦和解释大语言模型的推理机制,并进行了因果干预分析。从客观角度看,其将稀疏特征激活模式与推理行为关联,并揭示了推理与句法格式的紧密耦合、以及思维链对特征间脆弱协调的依赖,为理解模型内部工作提供了新的可解释性视角。
Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poorly understood. To deconstruct this cognitive process, we apply Top-K Sparse Autoencoders (SAEs) to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examine the model’s divergent behaviors across math-solving tasks of three distinct difficulty levels. Observationally, we identify a clear distinction in how the model functions under two reasoning modes: Thinking mode relies on sparse and high-intensity feature activations driving verbal deduction independent of problem complexity, whereas NoThinking mode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: (i) reasoning and syntactic structure are tightly coupled, as interventions consistently degrade \LaTeX{} and boxed-solution formatting; (ii) Thinking responds to disruption with compensatory over-generation marked by increased metacognitive cues and repetitive, low-information continuations; and (iii) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure.
[13] Hidden Language Consistency Phenomena in Reasoning LLMs cs.CL | cs.AIPDF
Muhammad Ali Shafique, Kelly Marchisio
TL;DR: 本文研究了多语言推理模型在任务难度增加时的语言一致性现象,发现模型在推理和回答过程中可能偏离输入语言,尤其是在非拉丁文字和代表性较弱的语言中。通过PolyMath基准测试,揭示了语言一致性随难度变化的四种行为模式,并提出了语言一致性崩溃效应。
Details
Motivation: 当前多语言推理模型评估主要关注答案正确性,忽略了模型在推理过程中是否保持输入语言的一致性,这掩盖了任务难度增加时的重要多语言行为。
Result: 在八个语言和四个难度级别的PolyMath基准测试中,发现语言一致性随难度变化呈现四种行为:保持对齐、保持错位、逐渐退化或突然崩溃。量化方法如GPTQ和AWQ在语言一致性方面优于AutoRound(基于ε=1.0的容忍度投票)。
Insight: 创新点在于揭示了语言一致性崩溃效应,即难度增加可能导致输出语言一致性突然下降,而模型可能转向其内部主导语言以保持或提高准确性。这强调了多语言评估需综合考虑任务准确性、语言一致性和任务难度。
Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors that emerge as tasks become harder. In this paper, we study task difficulty, task accuracy, thinking-language consistency (TC), and answer-language consistency (AC) across reasoning models using PolyMath benchmark in eight languages and four difficulty levels. We uncover four findings: (1) language consistency exhibits four difficulty-dependent behaviors: output-language consistency remains aligned with input, remains misaligned, degrades gradually, or collapses abruptly. (2) We identify the language consistency breakdown effect, where increasing difficulty can cause a sudden drop in output-language consistency, especially in less strongly represented and non-Latin-script languages. (3) Due to this breakdown effect, accuracy can be preserved or even improved at a harder difficulty level as the model shifts to its internal dominant language. (4) Quantization can improve or degrade output-language consistency independently of its effect on accuracy, with GPTQ and AWQ often outperforming AutoRound under tolerance-based voting with ε = 1.0. These results show that multilingual capability cannot be characterized by accuracy alone; reliable evaluation should jointly consider task accuracy, language consistency, and task difficulty for multilingual benchmarks.
[14] Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction cs.CL | cs.AI | cs.DBPDF
Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang, Jiazhuo Chen, Nan Tang
TL;DR: 该论文提出了Doc2DB-Bench,一个用于评估将长文档转换为关系型数据库的基准测试集。它旨在解决现有文档转表格基准的不足,通过包含多个领域、复杂模式和跨表关系,来测试系统能否构建出关系一致、可查询的数据库。
Details
Motivation: 现有文档转表格基准将信息扁平化为单表,导致实体重复、多对多关系模糊、记录稀疏,且无法验证提取的事实是否能构成有效的数据库实例。因此,需要一个新的基准来评估将文档理解作为数据库构建任务的能力,以支持下游的分析、合规和基于SQL的决策等工作流。
Result: 论文构建了Doc2DB-Bench基准,包含7个领域组、42个模式下的203个长文档实例,总计117个实体表、132个关系表、7,341行和41,935个单元格。该基准通过可控的数据库到文档合成流程构建,并经过真实性验证,与真实世界文档难以区分。
Insight: 核心创新在于将文档理解任务从字段提取提升到关系型数据库构建,并为此创建了一个系统性的基准。其可控的合成流程和基于模式内提取与模式间推理的分类法,为评估LLM在构建可靠、可审计且关系一致的数据系统方面提供了新的测试平台。
Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream workflows rely on normalized schemas, entity identities, keys, cross-table relationships, and integrity constraints for analytics, compliance, auditing, and SQL-backed decision making. Existing Document-to-Table benchmarks are insufficient for this setting: flattening evidence into single tables can duplicate entities, obscure many-to-many relationships, create sparse records, and avoid testing whether extracted facts form a valid database instance. This creates an urgent need to evaluate document understanding as database construction rather than field extraction. We introduce Doc2DB-Bench, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells. Built through a controllable DB-to-Doc synthesis pipeline and organized by a taxonomy of intra-table extraction and inter-table reasoning, the generated documents undergo authenticity verification, proving indistinguishable from real-world references. Doc2DB-Bench thus provides a testbed for reliable, auditable, and relationally faithful LLM-based data systems. The benchmark is publicly available at https://github.com/SetonLiang/Doc2DB-Bench.
[15] Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization cs.CL | cs.AIPDF
Haojie Yu, Ziyou Jiang, Junjie Wang, Mingyang Li, Yuekai Huang
TL;DR: 该论文提出了一种名为BRACE的方法,用于检测不断演变的恶意聊天对话。该方法通过识别恶意对话中不变的有序推理链(ORC),包括重复主题、恶意语言指标、严重性层级和类型特征,将ORC编码为四个可微分的阶段,并结合直接预测头、基于原型的特征增强和特征路径解耦进行正则化,以捕捉频繁变化的词汇表达中的关键信息。
Details
Motivation: 恶意聊天对话通过类型转换和词汇规避不断演变,难以检测。作者发现这些对话共享不变的原则,即一个有序推理链,这有助于捕捉变化词汇中的关键信息,从而解决恶意对话检测的挑战。
Result: 在4个领域和5种恶意类别的评估中,BRACE使用RoBERTa-wwm-ext编码器实现了0.934的恶意类型宏平均F1分数(3次种子平均),使用Qwen3-1.7B LoRA解码器骨干网络达到0.949。消融研究表明所有组件都有效,且ORC的结构分解使BRACE能区分具有语义模糊性的恶意类型。
Insight: 创新点在于将恶意对话建模为一个有序推理链(ORC),并将其编码为可微分的结构化正则化阶段,结合了中间监督和特征增强。这提供了一种结构化方法来处理动态变化的恶意内容,通过分解语义元素提高类型区分能力,可借鉴于其他需要处理演变或模糊模式的分类任务。
Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic -> Indicator -> Severity -> Type) with intermediate supervision, serving as a structured regularizer blended with direct heads, and supported by prototype-based feature augmentation and feature path disentanglement. The evaluation results show that, across 4 domains and 5 harm categories, BRACE achieves harm-type macro F1 of 0.934 (RoBERTa-wwm-ext, 3-seed mean), with decoder backbones (Qwen3-1.7B LoRA) reaching 0.949. Ablation studies show that all components contribute to BRACE, and the structural decomposition of ORC enables BRACE to distinguish harmful types with semantic ambiguity. Disclaimer: This paper may contain content that is disturbing to some readers.
[16] From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios cs.CLPDF
Thai-Binh Nguyen, Zhaolin Li, Jan Niehues, Alexander Waibel
TL;DR: 本文分析了在鸡尾酒会场景下处理多说话人语音识别的多种系统策略。研究基于CHiME-9 MCoRec任务,该任务要求系统从视听输入中识别说话人组并转录各自的对话。最佳系统实现了高达57%的相对错误率降低。
Details
Motivation: 解决鸡尾酒会场景下,语音识别系统难以在多人同时说话的环境中分离并转录特定说话人对话的挑战。
Result: 在CHiME-9 MCoRec基准测试中,最佳系统实现了高达57%的相对错误率降低。分析表明,高语音重叠本身并不能完全解释性能差异。
Insight: 论文总结了三种互补的核心策略:显式或隐式的视听目标语音分离、针对每个目标说话人的改进视听语音识别,以及利用大语言模型对说话人进行分组并增强对话一致性。研究挑战了重叠是鸡尾酒会识别主要困难来源的常见假设。
Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This “cocktail party” scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a testbed where systems must recognize groups of speakers and transcribe each of their conversations from audio-visual input. In this work, we analyze a diverse set of systems, representing different design directions for addressing the cocktail-party scenario, where the best system achieves up to 57% relative error reduction. We identify three main strategies: (1) explicit or implicit audio-visual target speech separation, (2) improved audio-visual speech recognition for each target speaker, and (3) the use of large language models to group speakers into conversations and enhance conversational consistency. Our analysis shows that these directions address complementary failure modes of the cocktail-party problem, and that high speech overlap alone does not explain performance differences, challenging the common assumption that overlap is the primary source of difficulty in cocktail-party recognition.
[17] VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use cs.CLPDF
Juan S. Santillana
TL;DR: 本文提出了VectraYX-Vision-1B,一个参数量小于20亿的西班牙语/拉丁美洲网络安全视觉语言模型。它通过MLP将冻结的SigLIP-so400m视觉编码器与一个10.4亿参数的西班牙语安全解码器连接,专门用于分析网络安全工具界面图像,并支持西班牙语问答、结构化推理和原生工具调用。论文报告了初步的视觉接地结果不佳,并提出了修复方案,同时通过三种变体的消融实验研究了无位置编码层对视觉块注意力的影响。
Details
Motivation: 解决西班牙语/拉丁美洲网络安全领域缺乏专用、轻量级视觉语言模型的问题,该模型需要能够理解网络安全工具界面,并用西班牙语进行结构化推理和工具调用。
Result: 初步视觉接地结果不佳,B6分数接近零(工具识别为0.08),表明当前模型忽略了图像内容。论文提供了文本骨干的B1-B5分数、文本控制、初步B6/B7分数、运行时间、CPU上的GGUF效率,以及包含10个领域14596个问答对的数据集。
Insight: 主要创新点包括:1) 首个专用于西班牙语网络安全UI的轻量级VLM;2) 通过原生<|think|>令牌实现结构化推理;3) 通过模型上下文协议(<|tool_call|>)实现原生工具调用;4) 引入三种位置编码变体进行消融实验,研究无位置编码层对长视觉序列注意力的影响;5) 支持导出到llama.cpp的LLaVA mmproj格式以实现物理隔离部署。
Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1.04B Spanish/LATAM security decoder via an MLP. To our knowledge, it is the first sub-2B VLM specialized for cyber UI (IDA, Ghidra, Wireshark, Nmap, Metasploit, Volatility) that answers in Spanish, emits structured reasoning via native <|think|> tokens, invokes tools via Model Context Protocol (<|tool_call|>), and exports to llama.cpp’s LLaVA mmproj format for air-gapped deployment. We report a negative preliminary visual-grounding result: despite fully functional pipelines, the current vision SFT (400-1900 steps, ~16M tokens) yields near-zero B6 scores (0.08 tool-identification), ignoring image content. We specify remediation (longer SFT, >=60% replay, lower LR) and expose a checkpoint-loader bug (unstripped llm. prefix) masquerading as training collapse. Crucially, we introduce a 3-variant ablation matrix (V0: NoPE-every-4, V1: all-RoPE, V2: NoPE+learned 2D) to study if periodic no-positional-encoding (NoPE) layers help or hurt attention over the 729-token visual block. Code, configs, and weights are released to establish priority on this architectural question. We provide B1-B5 for the text backbone, text controls, preliminary B6/B7 scores, wall times, GGUF efficiency on CPU, and a corpus of 14,596 QA pairs across 10 domains. We open-source all models and trajectories: jsantillana/vectrayx-1b, jsantillana/vectrayx-vision-1b, and jsantillana/vectrayx-vision-1b-checks.
[18] OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories cs.CL | cs.CV | cs.LGPDF
Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma
TL;DR: 本文提出了OpenVisTool框架,用于构建具有教学意义的视觉工具使用轨迹数据集,以解决现有方法中仅基于答案正确性筛选轨迹导致监督信号不足的问题。该框架通过难度筛选、轨迹合成和监督验证三个阶段,确保轨迹既答案正确又工具观察对答案有因果贡献,从而生成更有效的监督数据。
Details
Motivation: 现有视觉工具使用学习方法通常从教师生成的轨迹中学习,仅过滤答案正确的轨迹,但忽略了工具调用可能并非答案正确的因果因素,导致学生模仿工具调用模式而非学习何时及如何获取视觉证据。
Result: 基于OpenVisTool框架构建了OpenVisTool-42K数据集(涵盖五个视觉推理领域)和OpenVisTool-Bench基准测试。在四个骨干模型(4B-27B)上微调后,视觉工具使用性能持续提升,并在两个分布外基准上取得增益;较大模型接近领先的闭源系统水平。
Insight: 创新点在于提出轨迹应同时满足答案正确性(结果有效性)和工具观察对答案的因果贡献(因果效用)的双重标准,强调从因果基础的监督中学习有效视觉工具使用,而非单纯模仿工具调用模式。
Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.
[19] Mitigating Gender Bias in English to Romanian Machine Translation cs.CL | cs.AIPDF
Ioana Grigore, Sergiu Nisioi
TL;DR: 本文提出了一种混合流水线方法,通过结合基于大语言模型(LLM)的性别分类与神经机器翻译(NMT),来缓解从性别中性的英语到有性别形态的罗马尼亚语机器翻译中的性别偏见问题。系统首先使用微调的LLM检测英语句子中目标词的预期性别并插入内联性别提示标签,然后将带标签的句子输入经过微调的Transformer模型以生成形态正确的罗马尼亚语翻译。
Details
Motivation: 机器翻译系统在翻译性别时经常出错,尤其是在从英语这类性别中性语言翻译到罗马尼亚语这类有性别形态的目标语言时,这种偏见导致翻译默认使用阳性形式或强化性别刻板印象。
Result: 在WinoMT和WinoGender基准测试上,该方法将性别准确率相比基线MT系统提高了超过40个百分点。
Insight: 主要创新点在于提出了一个结合LLM性别消歧与标签感知翻译的混合流水线,并为此引入了三个新的性别消歧和翻译数据集。这是首个在英罗翻译中同时利用LLM推理和标签感知翻译来明确处理和评估性别偏见的方法。
Abstract: Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to masculine forms or reinforce gender stereotypes. We propose a hybrid pipeline to mitigate this issue by combining large language model (LLM)-based gender classification with neural machine translation (NMT). Our system uses a fine-tuned LLM to detect the intended gender of target words in English sentences and insert inline gender hint tags. These tagged sentences are then passed to a Transformer model fine-tuned to generate morphologically correct Romanian translations. To support this, we introduce three novel datasets for gender disambiguation and translation. Our approach improves gender accuracy on the WinoMT and WinoGender benchmarks by over 40 percentage points compared to a baseline MT system. This is the first method to explicitly address and evaluate gender bias in English-Romanian MT using both LLM inference and tag-aware translation.
[20] OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents cs.CLPDF
Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng, Albert Ventayol-Boada, Gabriel Mejia Gonzalez
TL;DR: 本文介绍了OmnilingualGAIA2,这是一个通过机器翻译(并经过部分人工专家验证)扩展的GAIA2智能体基准测试,涵盖了五种书写系统的十种目标语言,并配备了本地化且经过人工校准的多语言验证器。通过评估七个前沿和开源模型,研究发现了一个普遍存在的跨语言性能差距,该差距在智能体间不对称,主要集中在工具编排而非定量推理上,且不随模型规模扩大而消失。
Details
Motivation: 现有的智能体基准测试几乎全是英文的,而AI智能体正被全球部署给语言多样的用户。因此,衡量在英文中评估的智能体能力是否能迁移到其他语言,是一个亟待解决的开放性问题。
Result: 评估发现了一个普遍的跨语言差距,pass@3分数下降8.8-18.4点。该差距在智能体间不对称,主要集中在工具编排任务上,且不随模型规模扩大而闭合。误差归因分析表明,差距主要由模型本身驱动(55%),而翻译污染造成的误差下限仅为6.4%。
Insight: 创新点在于构建了一个覆盖多语言、多书写系统的智能体评估基准,并系统性地量化了跨语言性能差距及其成因。客观来看,该研究强调了将多语言智能体评估作为全球部署智能体标准报告协议的必要性,并揭示了非拉丁文字语言中形态线索丢失和歧义放大是主要的失败机制。
Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.
[21] Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue cs.CLPDF
Esam Ghaleb, Hugh Mee Wong, Kristina Kobrock
TL;DR: 本文研究视频中介对话中不同伙伴可见性条件下多模态信息(尤其是手势)的指称信息量。通过构建基于语音转录、手势骨架表示或两者融合的模型来识别指称对象,发现手势本身具有预测性,且当基于转录的模型不确定时多模态融合最有效。与人类交互数据对比进一步揭示了对话者可见性对手势产生和信息度的语用影响,以及跨轮次交互中语音和多模态(而非手势)表现的趋同效应。
Details
Motivation: 解决对话模型通常仅依赖语音转录而忽略手势等多模态信息的问题,探究在不同伙伴可见性条件下手势及其与语音结合所携带的指称信息量。
Result: 在视频中介指称通信游戏中,手势单独模型能预测目标指称对象;多模态融合在基于转录的模型不确定时最有益;通过训练时与指称图像对齐学习表示进一步提升了融合模型性能。
Insight: 创新点在于系统量化了不同可见性条件下手势的信息贡献,并展示了多模态融合在模型不确定性高时的优势;技术贡献包括利用学习表示分析人类交互数据中的语用效应和趋同现象。
Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.
[22] ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making cs.CLPDF
Haohao Zhu, Xiaolin Shi, Jiayu Zhou
TL;DR: 本文提出了EHR2Dial-Triage,这是一个基于MIMIC-IV-ED电子健康记录(EHR)的、面向急诊分诊的对话生成框架与基准。该框架模拟了急诊分诊中基于角色和时间信息边界的交互式对话过程,将患者陈述与EHR中的时序事件相关联,用于评估模型在信息获取、证据使用、紧急程度预测和医患沟通等方面的能力。
Details
Motivation: 现有急诊分诊基准大多基于固定的临床快照来评估病情严重程度预测,未能捕捉分诊中通过交互对话动态获取和解释证据的过程;同时,现有医疗对话数据集中的对话陈述并不总是与EHR中的时序事件相关联。因此,需要一个新的基准来研究分诊作为一个动态的临床信息获取、推理和沟通过程。
Result: 论文提出了EHR2Dial-Triage基准,它支持对信息获取、证据使用、五级急诊严重指数(Emergency Severity Index)预测以及面向患者的沟通进行受控评估。该基准为在不同模型和患者角色下研究对话分诊提供了一个结构化环境。
Insight: 创新点在于提出了一个基于EHR的、具有明确角色和时间信息边界的对话生成框架,将对话回合与EHR中的时序事件精确关联,从而能够更真实地模拟和评估急诊分诊中的动态交互过程。这为研究临床对话中的信息获取和推理提供了新的结构化基准。
Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be limited to a chief complaint and initial vital signs. Clinically important details, including symptom onset and progression, associated symptoms, medical history, and medication use, are often obtained through focused conversation. Effective triage therefore requires clinicians to identify information gaps, ask appropriate follow-up questions, and update their assessment as new evidence becomes available. Most existing ED benchmarks evaluate acuity prediction from a fixed clinical snapshot. Although this formulation measures predictive performance after patient information has been assembled, it does not capture the interactive process through which triage-relevant evidence is elicited and interpreted. Existing medical dialogue datasets support the study of clinical communication, but dialogue statements are not always linked to temporally ordered events in the electronic health record (EHR). We introduce EHR2Dial-Triage, an agentic conversation-generation framework and benchmark grounded in MIMIC-IV-ED. The framework constructs triage conversations under explicit role-based and temporal information boundaries. Each accepted patient disclosure is linked to its supporting EHR event and the first dialogue turn at which it becomes available. EHR2Dial-Triage enables controlled evaluation of information elicitation, evidence use, five-level Emergency Severity Index prediction, and patient-facing communication across models and patient personas. It provides a structured setting for studying conversational triage as a dynamic process of clinical information acquisition, reasoning, and communication.
[23] Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents cs.CLPDF
Zihao Deng, Yining Zhu, Leiming Wang, Jingfei Lu, Junbo Wang
TL;DR: 本文提出了Tree-of-Experience (ToE)框架,旨在解决LLM智能体在持续自我进化中经验管理效率低下的问题。该框架将经验组织成与智能体分层推理过程对齐的共享树结构,从而支持系统化的更新、迁移和高效检索。在Game of 24和FinEvolveBench基准测试中,ToE显著提升了智能体的问题解决性能和效率。
Details
Motivation: 现有方法将经验表示为孤立的轨迹或抽象知识,与底层的推理过程脱节,这限制了反馈归因、跨任务迁移以及更新和检索的效率,尤其是在仅提供结果级反馈的复杂推理任务中。
Result: 在Game of 24任务上,ToE相比无经验的ToT基线,准确率相对提升了31.4%。在FinEvolveBench的12个评估设置中,ToE相比无经验流水线平均提升了41.24%的tsIC指标,而传统经验管理方法通常表现不如无经验基线。
Insight: 核心创新在于将经验组织与LLM智能体的分层推理过程对齐,构建一个共享的分析视角和推理路径树,并通过环境结果校准其可靠性。这为经验管理提供了一种结构化、可系统更新的新范式,有效提升了复杂任务中的经验复用和迁移效率。
Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related trajectories, but their experience representations are often disconnected from the underlying reasoning process. This limits feedback attribution, cross-task transfer, and update and retrieval efficiency, particularly in complex reasoning tasks with outcome-level feedback. To overcome this limitation, we propose \textbf{T}ree-\textbf{o}f-\textbf{E}xperience (ToE), a structured experience-management framework that aligns experience organization with the hierarchical reasoning process of LLM agents. Specifically, ToE organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient retrieval. The experimental results on \textsc{Game of 24} and \textsc{FinEvolveBench} show that ToE substantially improves both problem-solving performance and efficiency. On \textsc{Game of 24}, ToE achieves a 31.4% relative improvement in accuracy over the experience-free ToT baseline. On \textsc{FinEvolveBench}, ToE improves tsIC by an average of 41.24% over the experience-free pipeline across 12 evaluation settings, whereas conventional experience-management methods often underperform experience-free baselines.
[24] When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information cs.CL | cs.AI | cs.HC | cs.LGPDF
Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
TL;DR: 本文系统分析了大型语言模型在临床信息不确定性下的行为,发现模型在不确定性增加时准确率下降,但置信度并未相应降低,导致不安全置信错误显著增加。研究基于MedMCQA数据集构建了两种不确定性评估框架:通过提示修改引入语言不确定性线索,以及构造答案移除场景强制模型识别信息不足并弃答。
Details
Motivation: 解决LLMs在临床不确定性环境下的可靠性问题,特别是在高风险临床部署中,模型过度自信的错误预测可能误导临床决策。
Result: 在500个医学问题上使用校准间隙、预期校准误差和不安全置信错误率等指标评估,结果显示模型置信度与准确度错位,且不同模型在正确答案不可用时的弃答能力存在显著差异,部分模型持续产生高置信度的幻觉答案。
Insight: 创新点在于提出了针对临床不确定性的系统评估框架,揭示了LLMs在信息缺失时置信度校准的普遍失败模式;客观而言,该研究强调了在临床工作流部署前开发不确定性感知评估方法的必要性。
Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.
[25] The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora cs.CL | cs.AIPDF
E. M. Freeburg
TL;DR: 本文研究了预训练语料库中文本提取和标记符号(notation)对模型行为的影响,提出了“干净窗口存活率”(clean-window survival)这一度量指标,并通过对13个公共语料库的普查、基础模型的预测实验以及写作风格分析,发现结构声明(announcement)而非具体标记符号是影响模型预测的关键线索。
Details
Motivation: 当前领域已认识到文本提取选择会影响模型行为,但从未测量过这些选择所引入语料库的标记符号;本文旨在量化并分析文档编排的标记符号如何作为训练变量影响模型能力。
Result: 在13个公共语料库的普查中,视觉转换的PDF切片存活率低至0.153,而C4语料库高达0.889;通过预注册的预测实验发现,删除结构声明会使后续文本预测显著变难,但交换标记符号则无影响;基础模型不会在作者基线之上强加标记风格。
Insight: 关键创新点在于提出了“干净窗口存活率”来量化边界推断需求,并实证揭示了影响模型预测的核心是结构声明本身,而非其具体标记符号;这促使数据格式设计应基于其训练的能力而非保真度,并建议在数据卡片中记录提取器身份和存活率。
Abstract: How a document’s arrangement is written down, its notation, is a training variable that no dataset card records. The field has established that text-extraction choices change model behaviour, and has never once measured the notation of what those choices put into the corpus. We define clean-window survival, a deterministic count of how much of a stream still demands the boundary inference, and measure notation on three fronts. What corpora carry: a census of thirteen public corpora, where survival falls to 0.153 in a vision-converted PDF slice against 0.889 in C4; the scarce resource is not unmarked text but long unmarked text; a pre-registered supply test finds what remains institutional, not consumer. Our own pre-registered prediction failed: converters do not fabricate structure on prose, and that null forced the reliability mechanism that survives it. What readers use: across five base models spanning 0.6B to 8.2B and two pipelines, deleting a structural announcement makes the following prose measurably harder to predict, while swapping its notation moves nothing. That zero does not make notation unimportant; it relocates the variable: the operative cue is the announcement, not the sigil. What writers impose: a bounded null. Base models do not impose the marked register above the authored baseline, and handed prose with every announcement deleted they do not put one back, at a rate indistinguishable from zero against an authored reference of zero. We ship the format those measurements imply: the pure frame, paragraphs in authored order, every announcement deleted into a reversible sidecar, mixed against the marked copy over announcement presence rather than notation. Choose format operators by the capability they train, not by the fidelity they preserve, and record extractor identity and survival on data cards.
[26] Evo-Bench: Can Language Models Improve Agent Harness? cs.CLPDF
Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen
TL;DR: 本文提出了Evo-Bench,这是首个用于评估语言模型在自主优化其操作框架(即harness evolution)方面内在能力的基准测试。该基准覆盖搜索、办公和通用智能体领域,通过一种新颖的框架引导构建方法,隔离了框架改进能力与基础模型性能,并防止了任务过拟合。
Details
Motivation: 当前LLM驱动的自主智能体评估主要集中于静态任务解决,而缺乏对智能体自主优化其操作框架(harness evolution)这一新兴前沿能力的系统性评测。现有评估方法难以将框架改进能力与基础模型强度分离,也无法防止任务特定过拟合或捕捉长视野的迭代研究过程。
Result: 在九个前沿和开源模型上的广泛评估显示,顶级模型在Evo-Bench上取得了高达16.6分的绝对增益,接近最先进的人工设计基线水平。自主演化在通用任务上优于人工框架,在搜索任务上表现出色,但在需要高度特定处理流程的办公任务上表现不佳。
Insight: 论文的创新点在于提出了首个专门评估智能体框架演化能力的基准Evo-Bench,并设计了框架引导的构建框架(包括辅助任务演化和敏感度感知分层分割)来严格隔离和衡量这种能力。客观来看,其提出的评估范式和对不同领域(搜索、办公、通用)演化能力差异的发现,为理解和推进智能体的自我改进能力提供了新的视角和工具。
Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution—the agent’s capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models’ intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
[27] LexKairos: Benchmarking Legal Temporal Capabilities in LLMs cs.CLPDF
Chenyang Li, Zejia Feng, Yuqin Huang, Yuxiao Ye, Huiyuan Xie
TL;DR: 本文提出了LexKairos,一个用于评估大型语言模型在中国法律语境下时间能力的综合基准。该基准涵盖法规时间知识、案件时间建模和法规-案件时间推理三个维度,包含九个来自真实中国司法案例和法规的子任务。作者对八个LLM在多种推理设置下进行了系统评估,发现Gemini-3-Flash表现最佳,但即使是最好模型在需要精确时间敏感法规元数据回忆或复杂时限推理的任务上仍存在明显局限。
Details
Motivation: 现有法律AI基准对法律时间能力探索不足,而时间在法律实践中是决定法规有效性、案件进展和程序截止日期的关键概念,因此需要专门的基准来评估LLM在这方面的能力。
Result: 在LexKairos基准上,Gemini-3-Flash取得了最强的综合性能,但所有模型在需要精确时间敏感法规元数据回忆或复杂时限推理的任务上均表现出显著局限性。
Insight: 创新点在于构建了一个专门针对中国法律语境、多维度(法规、案件、法规-案件推理)评估LLM时间能力的综合基准。客观来看,该工作揭示了当前LLM在法律时间知识与推理方面仍是一个开放挑战,为未来研究指明了方向。
Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the enforcement of procedural deadlines. However, legal temporal capabilities remain underexplored in existing legal AI benchmarks. To address this gap, we propose LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute-case temporal reasoning. LexKairos comprises nine sub-tasks drawn from real-world Chinese judicial cases and statutes. We conduct systematic evaluations of eight LLMs under multiple inference settings, including vanilla, Chain-of-Thought (CoT), and thinking modes. Our results show that Gemini-3-Flash achieves the strongest overall performance, yet even the best-performing model exhibits notable limitations on tasks demanding precise time-sensitive statutory metadata recall or complex reasoning in time limits, indicating that legal temporal knowledge and reasoning remain open challenges for current LLMs. Data and code are available at https://github.com/thunlp/LexKairos.
[28] An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer cs.CLPDF
Mengxian Lyu, Cheng Peng, Tim Jang, Ang Li, Mengyuan Zhang
TL;DR: 本文提出了GatorOnco,一个用于结直肠癌治疗规划的智能体化大语言模型。该模型通过整合大规模生物医学文本预训练、模型融合、两阶段后训练以及基于智能体的强化学习进行领域适应,并采用智能体化检索增强生成技术动态集成时效性临床指南。在由五位肿瘤学家进行的盲法随机临床评估中,其性能显著优于开源大语言模型,并在多个维度上达到与专家相当的水平。
Details
Motivation: 精准肿瘤学的治疗规划需要综合异质性患者信息与快速演进的临床指南,以确保符合指南的护理。现有大语言模型在高风险治疗规划中的应用受到复杂推理、对时效性指南的遵循以及安全性问题的阻碍。
Result: 在佛罗里达大学健康中心五位肿瘤学家进行的盲法随机临床评估中,GatorOnco显著优于开源大语言模型,并达到了与专家相当的性能。具体而言,在可读性和完整性上显著优于专家,在正确性、时效性和安全性上与专家表现无统计学差异。
Insight: 论文的创新点在于将智能体化推理与大规模领域适应相结合,特别是通过智能体化RAG动态集成时效性临床指南,以解决高风险医疗场景中LLM的可靠性问题。这为生成式AI在关键医疗决策中的应用提供了一种可行的技术路径。
Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnostic tasks, their adoption for high-stakes treatment planning is hindered by complex reasoning, adherence to timely clinical guidelines, and safety concerns. In this study, we present GatorOnco, an agentic LLM for colorectal cancer (CRC) treatment planning. GatorOnco is developed using a total of 282 billion tokens of biomedical text, including healthcare system-scale clinical text comprising 166 billion tokens from UF Health. We implemented a domain-adaptation method that integrates pre-training, model merging, a two-stage post-training approach, and agent-based reinforcement learning. An agentic retrieval-augmented generation (RAG) approach dynamically integrates time-sensitive clinical guidelines into the reasoning process. In a blind, randomized clinical evaluation conducted by five UF Health oncologists, GatorOnco significantly outperformed open-source LLMs (P < 0.01) and achieved expert-level performance comparable to UF Health oncologists. Compared with expert oncologists, GatorOnco received significantly higher ratings for readability (4.46 vs. 4.19, P < 0.01) and completeness (3.91 vs. 3.52, P < 0.01), while showing statistically comparable performance in correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999). These findings demonstrate that integrating agentic reasoning with large-scale domain adaptation can help bridge the gap for generative AI in high-stakes cancer treatment planning.
[29] Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments cs.CL | cs.AI | cs.MAPDF
Keyu He, Xuhui Zhou, Maarten Sap
TL;DR: 该论文提出了Social Gym和SPaRTan两个框架,用于评测和改进大语言模型(LLM)在多智能体社交环境中的推理能力。Social Gym是一个包含21个多智能体社交游戏(如狼人杀、抵抗组织)的基准测试环境,通过基于规则的胜负判定提供客观、可验证的性能评估,并采用Elo等级分锦标赛生成跨游戏排行榜。SPaRTan则是一种无需训练的自改进循环,通过自我对弈、反思轨迹并生成可迁移的策略手册来提升模型在特定游戏角色中的表现。
Details
Motivation: 当前评估LLM在需要合作、协商和适应的多智能体社交环境中的能力非常困难,因为社交互动缺乏客观的ground truth,现有方法依赖成本高、主观且噪声大的LLM法官,且模型无法获得可靠的学习信号。
Result: 基准测试表明,GPT-5-mini在Social Gym的排行榜上领先,但没有模型能在所有游戏或所有游戏角色中表现一致优异,揭示了社交推理的局限性。SPaRTan方法帮助GPT-5-mini智能体在其较弱的角色上提升了表现,但对Qwen3-32B模型的改进效果有限。
Insight: 主要创新点在于构建了一个基于规则、结果可验证的客观基准测试环境(Social Gym)来量化评估LLM的社交推理能力,并提出了一种无需权重更新的、基于自我对弈和反思生成可迁移策略的自改进方法(SPaRTan),为LLM社交能力的测量和提升提供了可复现的基础。
Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B’s performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.
[30] Verifiably grounded machine interpretation of lunar geology cs.CL | cs.LGPDF
Tom Sander, Kay Wohlfarth, Christian Wöhler
TL;DR: 本文提出了一种可验证的、基于多模态视觉语言架构的自动化‘机器智能地质学家’方法,用于月球地质解释。该方法通过整合配准的地形、光谱和地质图,生成可验证的地质解释,并引入开放书籍检索机制来准确引用已发表的地质年代数据。
Details
Motivation: 传统行星地质学依赖历史性和解释性推理,本文旨在将这种地质知识发现和推理的方法论嵌入到多模态架构中,以实现自动化地质解释,解决从多样化观测中重建过去事件的挑战。
Result: 模型在描述月球玄武岩月海火山作用的层序和地形时,能有效平衡先验地质知识与局部视觉证据,但仅基于视觉的数值定年会依赖记忆先验;通过集成开放书籍检索机制,模型能准确引用已发表的地质年代,实现了可验证的自动化地质推断。
Insight: 创新点在于将地质推理方法论嵌入多模态视觉语言模型,并引入检索机制分离视觉解释与定量历史上下文,为自动化地质推断提供了必要的架构设计,即局部数据视觉解释与科学记录检索相结合。
Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an automated “machine intelligence geologist” by embedding this distinct methodology of geologic knowledge discovery and inference into a multimodal vision-language architecture. Focusing on the stratigraphy of lunar basaltic mare volcanism, we train a model to generate verifiably grounded geologic interpretations directly from co-registered topographic, spectral, and geologic maps. We demonstrate that while the system successfully balances established geological priors with local visual evidence to accurately describe stratigraphy and terrain, numeric age dating derived solely from vision defaults to memorized priors. Integrating an open-book retrieval mechanism resolves this, enabling the model to faithfully cite published chronologies. Our findings delineate the necessary architecture for automated geologic inference: site evidence must be visually interpreted from local data, while quantitative historical context must be retrieved from the scientific record.
[31] EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models cs.CLPDF
Junyu Wang, Siyuan Zhang, Peiyuan Jiang, Jian Zong, Jingyu Zhang
TL;DR: 该论文提出了EmoS框架,这是一个基于四分支情绪智力理论构建的评估与对齐口语语言模型情绪智力的系统。研究首先创建了EmoSBench基准,用于全面评估模型在感知、理解、使用和管理情绪方面的能力,发现现有顶尖模型表现远低于人类水平。为解决此差距,作者开发了EmoS评估器模型,通过监督微调、GRPO优化以及结合SEAR和RFR的奖励机制进行训练,使其准确率接近人类水平,并在真实口语对话中展现出强大的泛化能力。
Details
Motivation: 当前口语语言模型在情绪智力评估方面缺乏系统、理论驱动的认知框架,评估仅局限于初级的副语言感知,无法全面衡量模型的情感智能。
Result: 在EmoSBench基准上的初步评估显示,即使是GPT-4o-Audio这样的领先专有模型准确率也仅为52.6%,远低于人类基线。经过训练的EmoS模型达到了83.8%的准确率,接近人类水平,并在真实无约束的口语交互评估中验证了其强大的现实世界泛化能力。
Insight: 论文的创新点在于将心理学中的四分支情绪智力理论系统地引入SLM评估,构建了首个全面的理论驱动基准EmoSBench。同时,提出了结合SFT、GRPO以及SEAR和RFR奖励机制的专门评估器训练框架,并创建了带有精细情绪分级标注的双语对话数据集EmoDialogue,为推进情感智能对话系统奠定了方法论基础。
Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
[32] Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing cs.CLPDF
Steve Woollaston, Brendan Flanagan, Hiroaki Ogata
TL;DR: 本研究针对日本初中生英语写作,提出了一种分层LLM校正流程,以区分语法错误与不自然地道的表达。通过分析3830份写作样本,量化了准确度差距和地道性差距,揭示了不同语法结构在准确性和地道性使用上的不同模式。
Details
Motivation: 动机在于解决自动写作评估中常将语法准确性与地道性混为一谈的问题,旨在分离结构性错误与不自然表达,以提供更精细的诊断。
Result: 研究结果在3830份日本初中生英语写作样本上量化了准确度差距和地道性差距。例如,定冠词、第三人称单数-s和情态动词(would, could)存在显著的准确性问题,而-ing形式和假设性情态动词(would)则显示出最大的地道性使用不足。
Insight: 创新点在于提出了一个二维教学分类法,将错误率与地道性差距进行映射,从而区分出准确但过度使用的语法项目与易错或回避的复杂形式。该框架能帮助教师诊断学习者的困难是源于执行不准确、结构回避还是母语映射的过度依赖,从而支持基于证据的针对性干预。
Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that isolates structural errors from unnaturalness by generating literal error corrections and idiomatic revisions for 3,830 English writing samples from 120 Japanese junior high school students. Applying the regex-based CEFR-J grammar extractor, we quantify two diagnostic measures: accuracy gaps (structures attempted but incorrectly produced) and idiomatic gaps (grammatically correct structures underused or overused relative to native norms). Results reveal distinct patterns: definite articles, third-person singular -s, and modals (would, could) exhibit significant accuracy difficulties, while -ing forms and hypothetical modals (would) show the largest idiomatic underuse, with simple present verbs, subject-verb-object patterns, and modal can conversely exhibiting the most pronounced overuse. A two-dimensional instructional typology maps error rates against idiomatic gaps, distinguishing accurate but overused grammar items from error-prone or avoided complex forms requiring targeted production practice. This framework advances pedagogical feedback by enabling teachers to diagnose whether learner difficulties arise from inaccurate execution, structural avoidance, or L1-mapped overreliance, supporting evidence-based interventions tailored to the specific needs of each learner.
[33] Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages cs.CLPDF
Zijie Zhang
TL;DR: 本文研究了针对突厥语族语言的跨语言迁移中,通用罗马化与语族特定文字统一方案的效果对比。通过比较通用罗马化工具uroman和突厥语族专用文字CTS,在11种突厥语上训练fastText模型,并在WikiANN命名实体识别和Universal Dependencies词性标注任务上评估。结果表明,两种方案在NER任务上无显著差异且均大幅超越单语基线,而POS任务表现则取决于跨语言字符n-gram覆盖度和目标语言监督情况。
Details
Motivation: 解决密切相关的语言因使用不同文字导致表面重叠度低,从而限制多语言模型跨语言迁移能力的问题。
Result: 在WikiANN NER任务上,CTS与uroman无显著差异,两者均显著优于官方单语fastText基线;在UD POS任务中,无普遍最优方案,CANINE-c整体平均分更高,但fastText系统在多个树库上仍具竞争力。
Insight: 创新点在于系统比较了通用与语族特定的文字统一策略,并揭示其效果受语言特性、诱导的子词重叠度和可用监督共同影响;客观来看,研究强调了跨语言任务中文字表征与语言结构适配性的重要性。
Abstract: Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare two approaches to script unification: the general-purpose uroman romanizer and the family-specific Common Turkic Script (CTS). We train matched fastText models on transliterated Wikipedia corpora from 11 Turkic languages and evaluate them on WikiANN named entity recognition and Universal Dependencies part-of-speech tagging. CTS and uroman show no significant difference on NER, while both substantially outperform the official monolingual fastText baselines. POS results reveal no universal winner: language-specific differences are associated with the cross-lingual character n-gram coverage induced by each representation, while within-language coverage becomes more important when target-language supervision is available. Although CANINE-c achieves higher overall POS averages, the substantially simpler fastText-based systems remain competitive on several treebanks. Overall, the effectiveness of script unification depends on the language, the induced subword overlap, and the available supervision.
[34] Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law cs.CL | cs.AI | cs.IRPDF
Rose Cymbler, Daniel Guez, Laurent Fabre
TL;DR: 本文针对法律RAG系统中的时间错位问题,提出了一个版本化语料库基准FiscalQA Pro,用于评估模型在法国税法问答中检索适用版本法律条文的能力。研究发现现有模型在封闭式测试中无法准确恢复日期适用的答案,静态RAG检索适用版本的成功率为0%,而作者提出的多版本索引端到端检索器达到了98.3%的严格准确率。
Details
Motivation: 解决法律RAG系统中普遍存在的时间错位问题,即系统错误地检索并引用当前生效版本的法律条文,而实际适用的可能是过去或未来的版本。作者认为法律问答是一个时间索引的检索问题,而非静态语料库处理。
Result: 在FiscalQA Pro基准上,11个模型(包括前沿闭源API系统和开源模型)的参数知识平均严格准确率仅为3.0%,基于静态当前版本语料库的RAG为2.7%。静态RAG检索日期适用版本的成功率为0%。作者提出的多版本索引端到端检索器(无需先知信息)达到了98.3%的平均严格准确率,而先知文章消融实验达到99.1%。
Insight: 创新点在于将法律RAG重新定义为时间索引检索问题,并构建了包含32,436个条文版本、跨越93年的版本化语料库基准。方法上采用原子事实“金块”(正则表达式和容错数值)进行确定性评分,避免了使用LLM作为评判者可能引入的时间偏差。多版本索引检索器显著提升了版本选择准确率,揭示了现有系统在时间推理上的根本缺陷。
Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem. We introduce FiscalQA Pro, pairing a versioned corpus of 32,436 article-versions of the French tax code (93 years, 1938-2031) with an all-model-hard temporal-reasoning track: 209 scored, expert-reviewed questions across 33 CGI articles (221 released; twelve flagged out of the answerable scope). At selection time, no evaluated model recovered its date-applicable answer closed-book in any of four sampling draws, and the currently in-force text lacks the gold value for all but one of the scored questions. Answers are scored deterministically via atomic ground-truth “nuggets” (regex and numeric-with-tolerance), never LLM-as-judge: an LLM judge would inherit the temporal bias it is meant to score. Across eleven models (five frontier closed-API systems plus Gemini 2.5 Pro as a substitute entry, and five open-weight), parametric knowledge yields 3.0% mean strict accuracy and RAG over a static current-version corpus 2.7%. Static RAG retrieves the date-applicable version 0% of the time, confidently citing a real but inapplicable version. Our end-to-end retriever over a multi-version index, with no oracle, reaches 98.3% mean strict; an oracle-article ablation reaches 99.1%, locating the residual gap in first-stage recall, not version selection. We additionally release a version-aware jurisprudence dataset of 69,208 citation links, together with the corpus, benchmark, model responses, and pipeline code.
[35] ZetaGPT: A Reference Implementation of Positional–Encoding–Free State–Space–Attention Language Models cs.CL | cs.AIPDF
Róisín Luo
TL;DR: 本文提出了ZetaGPT,一种无需显式位置编码的语言模型参考实现。该模型将因果状态空间方程集成到每个模型块的自注意力层之前,通过循环状态动态隐式编码序列信息,从而使注意力层能够处理具有位置感知的表示。ZetaGPT是一个完全开源、端到端的紧凑型混合语言模型,旨在用于研究、快速原型设计、算法验证和教育应用。
Details
Motivation: 动机在于探索无需显式位置编码的架构。现有Transformer模型通过位置嵌入或编码(如RoPE)显式引入位置信息,这被视为一种架构获得的能力而非模型固有属性。本文旨在设计一种能隐式编码位置信息的模型。
Result: 摘要中未提及具体的定量实验结果或基准测试性能。论文主要宣称ZetaGPT是首个开源的、无显式位置编码的小型语言模型,并提供了一个可复现的参考实现。
Insight: 主要创新点在于提出了一种将因果状态空间方程与自注意力结合的混合架构,以隐式、递归的方式在注意力计算前编码位置信息,从而无需显式位置编码。此外,提供了一个涵盖从数据构建到RLHF和CoT推理的完整开源训练流程,为无位置编码语言模型的研究和开发建立了紧凑的参考基准。
Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating positional information through learned positional embeddings or hand-crafted positional encodings, such as rotary positional encoding (RoPE), treating positional information as an architecturally acquired capability rather than an inherent property of the model. Motivated by the pursuit of positional-encoding-free architectures, this work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation. Specifically, each model block applies a causal state-space equation before self-attention, allowing recurrent state dynamics to encode sequential information into token representations. Consequently, subsequent attention layers operate on position-aware representations without requiring explicit positional encodings while retaining the expressive modeling capacity of self-attention. We present \textsc{ZetaGPT}, a compact hybrid language model designed for research, rapid prototyping, algorithm verification, and educational applications. In addition to the proposed architecture, \textsc{ZetaGPT} provides a fully open-source, end-to-end training pipeline encompassing dataset construction, tokenizer training, pretraining, supervised fine-tuning, reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning via pure reinforcement learning. To the best of our knowledge, \textsc{ZetaGPT} is the first open-source small language model without explicit positional encoding and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.
[36] Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation cs.CLPDF
Bo Wang, Ruixing Zhang, Yunqi Liu, Yang Zhang, Liangzhe Han
TL;DR: 本文提出了一种名为UserIDA的可控用户模拟方法,其核心是将用户意图与语言表达分离,通过显式的每轮意图指令来指导用户模拟的生成。该方法通过监督微调学习指令条件生成,并在基于群体的强化学习中使用意图校准策略优化,以同时保证响应质量和意图准确性。
Details
Motivation: 现有用户模拟器在生成下一轮用户对话时存在一对多问题,即相同上下文可能产生多种合理但意图不同的回应,导致模拟对话可能通过不恰当的意图(如接受而非修正)推进。论文旨在解决如何精确控制用户模拟的局部交互意图,而不仅仅是模仿响应风格。
Result: 在LMSYS-USP基准测试上,UserIDA实现了86.6%的意图准确率,比最强的专用用户模拟基线高出24.3个百分点,同时提升了语义和风格相似性。在上下文干预评估中,它在91.7%的对话状态下实现了至少四种目标意图,而最强外部基线仅为22.9%。
Insight: 论文的创新点在于将每轮交互意图作为显式指令(六类意图接口)引入用户模拟,实现了意图与表达的解耦控制。其提出的意图校准策略优化和混合组排序奖励机制,为构建高保真且意图可控的对话模拟环境提供了新范式。
Abstract: User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7% of evaluated dialogue states, compared with 22.9% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.
[37] Reducing Pretraining-Generation Mismatch in Diffusion Language Models cs.CLPDF
Xiaocheng Lu, Huabin Liu, Song Guo, Jianguo Li
TL;DR: 本文提出了一种名为PCD(前缀条件扩散)的预训练目标,旨在解决扩散语言模型(dLLM)中预训练与生成之间的不匹配问题。该方法通过结合自回归前缀监督和无偏移后缀去噪,使训练时的上下文分布与评估时基于提示的生成条件对齐,从而在不改变推理过程的情况下提升模型性能。
Details
Motivation: 扩散语言模型在预训练时可能随机破坏提示词和后续词,削弱了基于干净前缀进行提示条件生成所需的接口,导致预训练与生成之间存在不匹配。本文旨在解决这一不匹配问题,以改善提示续写的效果。
Result: 在LLaDA2-Mini和Qwen-1.7B骨干网络上,PCD相比同系列原生dLLM基线模型取得了一致性改进:在LLaDA2-Mini的六项基准测试平均上获得4.2%的相对提升(+2.56分),在Qwen的主要机制比较中获得14.2%的相对提升(+4.86分)。
Insight: 创新点在于提出了PCD预训练目标,它在训练目标层面通过调整注意力掩码、破坏掩码和标签构建,将自回归监督仅用于干净前缀,而将扩散过程仅应用于未知后续部分,从而在局部训练接口上模拟了评估时的块扩散查询方式。这种方法无需自回归解码器、验证器或新的推理模式,即可有效对齐预训练与生成分布。
Abstract: Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native dLLM pretraining can randomly corrupt prompt and continuation tokens together, weakening the clean-prefix interface needed for prompt-conditioned generation. We identify this mismatch for prompt continuation and propose PCD (Prefix-Conditioned Diffusion), a pretraining objective that combines AR prefix supervision with no-shift suffix denoising. At the training-objective level, PCD changes the attention mask, corruption mask, and label construction in continued pretraining; it does not require an autoregressive decoder, verifier, or new inference mode. By supervising the clean-prefix side autoregressively and applying diffusion only to the unknown continuation, PCD makes the local training interface resemble how block-diffusion models are queried at evaluation time. We further separate intra-sample prefix conditioning from inter-sample objective mixing, allowing us to identify the local alignment signal separately from the optional batch-level mixing knob. Across LLaDA2-Mini and Qwen-1.7B backbones, PCD consistently improves over same-family native dLLM stable baselines, reaching a 4.2% relative gain on the main LLaDA2-Mini six-benchmark average (+2.56 points) and a 14.2% relative gain in the primary Qwen mechanism comparison (+4.86 points). These results suggest that aligning the pretraining context distribution with prompt-conditioned generation can recover a measurable part of the dLLM continuation gap without changing inference.
[38] Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning cs.CL | cs.AIPDF
Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo
TL;DR: 本文提出AlignXada框架,通过基于语言的强化学习(verbal reinforcement learning)实现任务特定的偏好适应,将通用的用户偏好摘要转化为针对下游任务的精简表示,从而提升大型语言模型个性化性能并节省上下文容量。
Details
Motivation: 通用用户偏好摘要包含与特定下游任务无关的信息,直接使用会浪费上下文容量并引入跨任务干扰,而手动设计任务特定偏好视图难以扩展,因此需要自动化的任务特定偏好适应方法。
Result: 在13个任务和三个下游模型(共39个任务-模型组合)上,AlignXada平均提升3.82分,在33个组合中取得改进,仅保留22.8%的原始偏好令牌,并在36个组合中优于检索增强生成(RAG)方法。
Insight: 创新点在于提出无需训练的无元学习框架,通过语言强化学习迭代优化文本精炼策略,实现可重用的偏好适应;该方法在保持源偏好忠实度的同时有效保留任务相关的个性化信号,为终身个性化智能体提供了一种实用的补充方案。
Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task–model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
[39] REFRAMED: Towards Realistic Audio Description Generation for Movies cs.CL | cs.CVPDF
Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
TL;DR: 本文提出了一种新的音频描述(AD)生成任务,要求模型同时决定描述内容和时机,并引入了高质量数据集REFRAMED,包含2023个电影视频片段,用于支持该任务的研究。
Details
Motivation: 现有AD生成方法在人工设定下工作,预先指定描述内容和时机,且依赖噪声转录和对齐流程,缺乏建模叙事上下文所需的丰富并行数据,无法满足真实电影AD生成的需求。
Result: 实验表明,最先进的AD系统和多模态大语言模型在REFRAMED数据集上优于简单基线,但仍远低于专家人类表现。
Insight: 创新点在于将AD生成重新定义为联合决策任务,并提供了包含专业AD转录、字幕和对齐剧本的高质量数据集,以及利用对话间隙和多参考比较的评估协议,为视频理解研究奠定了基础。
Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,023 videos that span 3,302 scenes from 206 movies, with professional AD transcripts (both American and British versions), professional subtitles and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that they outperform trivial baselines but fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on video understanding.
[40] Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification cs.CL | cs.SD | eess.ASPDF
Abner Hernandez, Tomás Arias Vergara, Daiqi Liu, Andreas Maier, Paula Andrea Pérez-Toro
TL;DR: 本文研究了基于音频训练的语音学模型PhonoQ在音频-发音建模中的应用,通过提取其Conformer模块的表示来提升发音轮廓的分类性能。实验表明,结合PhonoQ特征能提高未见语音和未见说话者场景下语音学目标(如发音方式、位置、浊音、元音高度和元音后位)的宏F1分数,并改善细粒度音素分类。
Details
Motivation: 实时MRI能观察语音产生时的声道发音,但将发音模式映射到语音学和音系学类别仍具挑战;研究旨在探索基于音频的语音学模型PhonoQ是否能为音频-发音建模提供有用信息。
Result: 在未见语音和未见说话者设置中,结合PhonoQ特征的模型在语音学目标分类上优于WavLM-large和HuBERT-large基线,宏F1分数提升,并改善了39音素分类;在仅使用发音轮廓的推理设置中,音频监督带来小幅但一致的增益。
Insight: 创新点在于利用结构化语音学特征(如发音方式、位置、浊音等)的监督训练来增强音频-发音模型的表示能力,实现从同步音频到发音模型的部分信息迁移,并通过后验分析展示了可解释的发音模式(如/t/的闪音化、/t/-/r/后缩或塞擦化、鼻音位置同化)。
Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio–articulatory modeling. Specifically, we extract representations from PhonoQ’s Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.
[41] PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models cs.CLPDF
Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh
TL;DR: 论文提出了PragMatch基准,用于评估大型视觉语言模型在多模态讽刺检测任务中是否真正理解语用不一致性,而非依赖表面线索。该基准包含3000个图像-文本对,通过系统掩蔽和注入实验揭示了LVLMs对词汇、OCR和风格等表面线索的敏感性。
Details
Motivation: 解决LVLMs在多模态讽刺检测中可能依赖表面相关性而非真正推理图像-文本关系的问题,区分语用不一致性与简单的跨模态不匹配。
Result: 实验表明LVLMs对词汇、OCR和风格等表面线索敏感,注入这些线索会显著改变模型预测,而底层图像-文本关系未变;PragMatch基准为评估超越表面对齐的多模态语用推理提供了系统测试平台。
Insight: 创新点在于构建了可控的PragMatch基准,通过系统掩蔽和针对性注入实验揭示LVLMs的捷径学习问题;客观来看,该方法为评估多模态模型的深层推理能力提供了可推广的框架。
Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.
[42] KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs cs.CL | cs.AIPDF
Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu
TL;DR: 论文提出KGCaRe方法,通过结合知识图谱自动构建与上下文检索来提升大语言模型在复杂条件问答任务中的表现。该方法从文档中提取结构化知识构建知识图谱,同时嵌入文档进行神经检索,通过LLM引导的迭代图遍历提取相关三元组路径,结合检索到的文本段落生成带解释的答案。
Details
Motivation: 解决大语言模型和检索增强生成在领域特定复杂条件问答任务中表现不佳的问题,假设通过结合非结构化文档和结构化知识图谱的知识能提升推理和答案准确性。
Result: 在两个复杂条件QA数据集上的实验表明,KGCaRe在Mistral、Mixtral、GPT-3.5和GPT-4o等多个LLM上均优于Vanilla LLM、Code Prompt、Text Prompt、Think-on-Graph、Vanilla RAG和HybridContextQA等基线方法,达到SOTA水平。
Insight: 创新点在于提出多提示提取策略构建知识图谱与神经检索的混合方法,以及LLM引导的迭代图遍历机制,通过路径形式的三元组提取和线索实体辅助的二次遍历增强上下文质量,实现符号推理与神经检索的有效融合。
Abstract: Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize that augmenting RAG with unstructured and structured knowledge, extracted from both documents and knowledge graphs (KGs), can improve reasoning and answer accuracy for such tasks. To test this, we propose KGCaRe, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs. KGCaRe constructs a KG from documents using a multi-prompt extraction strategy and stores it in a graph database. Simultaneously, the documents are embedded into a vector store to enable neural retrieval. KGCaRe performs innovative iterative graph traversal guided by the LLM to extract relevant triples, prune irrelevant information, and uses additional clue entities to traverse the graph again if the initial traversal does not provide satisfactory context to generate the answer. The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions with explanations. We evaluate KGCaRe on two complex conditional QA datasets. Our results on these datasets show that KGCaRe consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA, across multiple LLMs such as Mistral, Mixtral, GPT-3.5, and GPT-4o. We publicly release the software pipeline that we developed to implement the proposed KGCaRe approach.
[43] Fusion Training for Mathematical Generalization in Large Language Models cs.CL | cs.AIPDF
Congfeng Cao, Pengyu Zhang, Jelke Bloem
TL;DR: 本文系统研究了思考模式融合(TMF)训练中的动态特性,重点关注数学问题解决场景下思考模式与非思考模式之间的数据比例和训练计划的影响。研究发现两种模式间存在不对称交互,增加非思考监督会降低思考模式的准确性,且存在内在的张力。
Details
Motivation: TMF方法使大语言模型能同时支持简洁响应和长链推理,但其训练动态(如两种模式间的数据比例和训练计划)尚未得到充分探索,本研究旨在填补这一空白。
Result: 在构建的包含多种数据比例和三种训练计划的基准测试中,结果显示增加非思考监督会降低思考模式的准确性,且最优训练计划取决于数据比例,量化了两种模式间的负相关性。
Insight: 创新点在于系统揭示了TMF中思考与非思考模式间的不对称交互和内在张力,为设计有效的TMF训练设置提供了实用指导;客观分析认为其对训练动态的量化研究为多模式统一模型的优化提供了新视角。
Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two modes, remain underexplored. In this work, we present a systematic study of TMF by analyzing the effects of the training schedule and data ratio between thinking and non-thinking modes. Focusing on mathematical problem solving, we construct a benchmark with multiple thinking-to-non-thinking data ratios and three training schedules. Our results reveal an asymmetric interaction between the two modes: increasing the ratio of non-thinking supervision reduces the accuracy of the thinking mode. We further show that different training schedules modulate this trade-off and that the optimal schedule depends on the data ratio. Finally, we quantify a negative correlation between non-thinking and thinking mode supervision, highlighting an inherent tension between these two modes. These findings provide practical guidance for designing effective TMF training settings. All code and data are released to support further research at: \href{https://github.com/caocongfeng/Fusion-Bench.git}{\textbf{Fusion Bench}}.
[44] Consilience for Verifier-Free Test-Time Scaling cs.CL | cs.LGPDF
Lecheng Kong, Like Hui, Haitao Mao, Jun Huan
TL;DR: 本文针对无验证器的测试时扩展(VF-TTS)方法,特别是基于置信度的方法,在复杂任务中可能因过度自信而失败的问题,提出了一种名为‘consilience’的新选择框架。该框架通过评估置信度在推理过程中的时间不对称性,惩罚高初始置信度并严格要求最终确定性,从而提升模型在复杂任务上的表现。
Details
Motivation: 现有基于置信度的VF-TTS方法在复杂任务中会灾难性地失效,因为均匀的高置信度往往意味着探索不足,导致模型倾向于选择自信但错误的答案。这促使研究者重新思考如何设计更稳健的置信度评估机制。
Result: 在研究生级别数学问题和自由形式代码生成任务上的大量实验表明,consilience框架显著优于现有基线方法,验证了其对完成置信度的新视角的有效性。
Insight: 创新点在于提出了置信度时间不对称性的概念,并设计了一个组合度量来主动惩罚高初始置信度,同时严格要求最终确定性,这为无验证器的测试时扩展提供了一种新的、更稳健的选择策略。
Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and tasks. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning, as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.
cs.CV [Back]
[45] Learning an Interior Layout Policy in a Domain Specific Language Action Space cs.CV | cs.AIPDF
Yuhao Lu, Weichen Zhang, Wenyi Xiao, Haohui Chen, Yiyun Fei
TL;DR: 本文提出LayoutDSL,一个基于LLM的框架,用于在领域特定语言(DSL)动作空间中学习室内布局策略。该方法通过DSL提供布局的显式符号表示,将布局设计转化为可解释的决策序列,并利用监督微调和基于设计原则的强化学习进行优化,显著提升了布局的空间合理性和设计逻辑性。
Details
Motivation: 现有室内场景布局生成方法通常过度简化任务,将房间条件简化为粗糙的3D边界框并忽略门窗等结构元素,且许多方法将空间推理建模为直接坐标预测,阻碍了模型学习智能布局设计的底层推理逻辑。
Result: 大量实验表明,LayoutDSL在空间合理性和设计逻辑性上显著优于强基线方法和现有方法,但没有具体提及在哪个标准基准测试上或是否达到SOTA水平。
Insight: 创新点在于引入DSL作为结构化动作空间,将布局生成转化为可解释的符号决策序列,并结合监督学习与基于设计原则奖励的强化学习来优化策略,这为将符号推理与生成模型结合解决结构化设计问题提供了新思路。
Abstract: Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundamentally, many prior approaches formulate spatial reasoning as direct coordinate prediction, thereby casting interior layout design as continuous regression over raw geometric parameters, which hinders the model from learning the underlying reasoning logic of intelligent layout design. We propose \textbf{LayoutDSL}, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space. The DSL provides an explicit symbolic representation of layout information and serves as a structured action space for layout reasoning, where each action corresponds to an interpretable design decision. Under this DSL-based policy learning paradigm, we construct 3D-FrontDSL, a dataset of room-structure annotations paired with synthetic DSL action sequences for supervised fine-tuning. To promote a more generalizable and scalable policy with verifiable feedback, we design rewards grounded in interior design principles and physical plausibility, and optimize the policy via reinforcement learning. Extensive experiments demonstrate that LayoutDSL substantially improves spatial plausibility and design logicality over strong baselines and existing methods.
[46] PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings cs.CV | cs.CLPDF
Jagpal Singh Jhala
TL;DR: 本文提出了PragyaDoc,一个面向低资源环境的多语言医疗文档理解通用框架。该框架通过四层流水线(并行集成OCR提取层、几何-词汇融合层、确定性领域结构化层和双LLM医疗推理与本地化层)来解决印度因22种官方语言导致的医疗文档可访问性障碍,使非英语使用者(如农村人口、ASHA工作者和患者家属)能够理解英文医疗文档。
Details
Motivation: 解决印度因22种官方语言导致的医疗文档可访问性障碍,即大多数医疗文档仅以英文存在,而最需要这些信息的农村人口、ASHA工作者和患者家属等群体被排除在理解之外。
Result: 摘要中未提及具体的定量实验结果、基准测试或与SOTA的比较。
Insight: 创新点在于提出了一个四层通用框架,特别是结合了并行集成OCR、几何-词汇融合、确定性领域结构化以及双LLM进行医疗推理和本地化,以应对低资源多语言环境下的文档理解挑战。从客观角度看,该框架的系统性设计(从提取到推理)和针对低资源场景的优化是其主要创新之处。
Abstract: India’s 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information - rural populations, ASHA workers, and patient families - are functionally excluded from understanding it. This paper presents PragyaDoc, a Universal Document Intelligence Framework that addresses this gap through a four-layer pipeline: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer
[47] Performance of large language models in the optical diagnosis of colorectal polyps cs.CV | cs.AIPDF
Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan
TL;DR: 本研究评估了多种多模态大语言模型(MLLMs)在结直肠息肉光学诊断中的性能。研究使用包含白光和窄带成像(NBI)图像的PRIME数据集,测试了Claude Opus 4、Gemini 2.5 Pro、GPT-o3、GPT-4o和GPT-5等模型在息肉分类(如肿瘤性/非肿瘤性、浸润性/非浸润性)和巴黎分型预测上的准确性。结果表明,所有模型在区分肿瘤性与非肿瘤性息肉时F1分数均超过0.9,但Gemini 2.5 Pro在更精细分类任务中表现最佳,而Claude Opus 4和GPT-5在巴黎分型上正确率最高(41.7%)。尽管部分模型表现接近专家共识,但其敏感性和特异性尚未达到欧洲胃肠内镜学会(ESGE)的临床标准。
Details
Motivation: 结直肠息肉的准确光学诊断对于指导切除策略和后续监测至关重要,而多模态大语言模型在基于图像的诊断中显示出潜力。本研究旨在评估这些模型在息肉分类和组织学预测中的诊断准确性,以探索其临床应用的可能性。
Result: 在PRIME数据集的132个病例上,所有MLLMs在区分肿瘤性与非肿瘤性息肉时F1分数均>0.9。Gemini 2.5 Pro在区分浸润性/非浸润性息肉以及低级别/高级别腺瘤上F1分数最高(分别为0.560和0.492)。在巴黎分型任务中,Claude Opus 4和GPT-5的正确率显著高于其他模型,达到41.7%。总体而言,Claude Opus 4和Gemini 2.5 Pro在区分息肉亚型上准确性最高,表现最接近专家共识。
Insight: 论文的创新点在于系统性地评估了当前领先的多模态大语言模型在结直肠息肉内镜图像诊断这一具体医学任务上的性能,并进行了细致的分类任务对比。从客观角度看,研究揭示了MLLMs在基础分类任务(如肿瘤性/非肿瘤性)上已具备高准确性,但在更精细的病理分级和分型任务上性能仍有显著差距,这为未来开发专用于医学图像分析的模型或设计人机协同(human-in-the-loop)临床工作流程提供了明确的方向和基准。
Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran’s Q and McNemar’s Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.
[48] Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions cs.CV | cs.LGPDF
Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng
TL;DR: 本文评估了三种生成式视觉语言模型在三个机构的胸部X光片数据集上的表现,通过超过34.5万个预测,研究了模型与机构参考标准之间的一致性如何跨机构、发现项、预测方向和问题格式转移。研究发现,使用简单的固定估计器(如Beta-Binomial经验贝叶斯估计器)在估计接收机构的参考一致性时表现最佳,且跨机构的一致性需针对每个站点和接口重新评估。
Details
Motivation: 视觉语言模型在返回结构化胸部X光片发现时未提供置信度评分,导致接收机构无法判断单个预测的可信度;同时,模型与机构参考标准的一致性是否能在不同站点、发现项、预测方向和问题格式间转移尚缺乏测量。
Result: 在严格排除接收机构数据的评估中,自适应选择估计器未优于简单固定估计器:平均Brier分数为0.1083,而Beta-Binomial经验贝叶斯估计器为0.0853,仅使用目标数据的逻辑模型为0.0855;名义95%水平的后验预测计数区间覆盖率为87.0%,在最具挑战性的机构中更低。
Insight: 创新点在于系统评估了视觉语言模型在医疗影像上的跨机构一致性,并提出了基于小规模本地标签估计参考一致性的方法;客观分析表明,模型性能高度依赖于具体站点和接口,强调了在医疗应用中针对每个环境进行独立评估的必要性,而非依赖跨机构池化估计。
Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution’s reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family’s own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.
[49] Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models cs.CV | cs.ROPDF
Eduardo de los Santos, Andre S. Kelbouscas, Ricardo B. Grando, Bruna V. Guterres
TL;DR: 本文研究了数据集构成对基于紧凑型YOLO模型的嵌入式实时无人机野火检测系统性能的影响。通过评估四种训练配置(真实非增强、真实增强、混合非增强、混合增强),发现使用真实非增强数据集在召回率和平均精度均值之间取得了最佳平衡。
Details
Motivation: 开发基于视觉的无人机野火检测系统受限于多样化的真实世界训练图像有限,本文旨在探究合成数据混合与图像增强在资源受限的部署条件下是否能提升实际检测性能。
Result: 实验结果表明,在无人机野火检测任务中,真实非增强数据集获得了最佳整体性能平衡点;无论是与合成数据混合还是进行图像增强,均未产生更好的最终部署选择。
Insight: 研究指出,对于嵌入式无人机野火检测,数据集的真实性和领域对齐比通过合成扩展增加训练集规模更有价值,这为资源受限场景下的数据集构建策略提供了重要见解。
Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the limited availability of diverse real-world training images. This paper investigates the impact of dataset composition on embedded real-time UAV wildfire detection using compact YOLO models as a controlled validation family. Four training configurations were evaluated: real non-augmented, real augmented, hybrid non-augmented, and hybrid augmented, where the hybrid sets combine real wildfire images with AI-generated samples. The objective is to determine whether synthetic data mixing and image augmentation improve practical detection performance under resource-constrained deployment conditions. Experimental results show that the best overall operating point was obtained with the real non-augmented dataset, which achieved the strongest balance between recall and mean average precision for UAV-based wildfire detection. The results also show that neither hybridization with synthetic data nor augmentation produced a better final deployment choice. These findings suggest that, for embedded UAV wildfire detection, dataset realism and domain alignment are more valuable than increasing training set size through synthetic expansion.
[50] MVMD: A Multi-View Approach for Enhanced Mirror Detection cs.CV | cs.LGPDF
Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu
TL;DR: 本文提出了一种名为MVMD的多视角镜面检测方法,旨在解决3D重建中镜面导致的模型失真和碎片化问题。该方法通过跨视角和自注意力机制学习不同视角下物体与镜面内外反射之间的关联,包含三个关键模块:跨视角块追踪视角变化引起的镜内物体位移,视角内块检测镜内物体反射,以及细化块锐化镜面边界和增强细节。
Details
Motivation: 现有镜面检测方法仅关注单图像检测,忽略了多视角设置提供的丰富信息,导致在3D重建中镜面引入的挑战未得到有效解决。
Result: 实验结果表明,与单图像镜面检测技术相比,MVMD在准确率上提升高达2.6%,IoU提升高达11.1%,在多镜面环境中显著提高了3D重建的准确性。
Insight: 创新点在于首次提出多视角镜面检测方法,并构建了首个专为多视角场景设计的镜面检测数据库;通过注意力机制建模视角间和视角内的关联,有效利用多视图信息提升检测性能。
Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. As 3D reconstruction typically relies on multi-view images to capture different perspectives of a scene, detecting and labeling mirrors in multi-view images before reconstruction can effectively address this issue. However, existing methods focus solely on single-image detection, overlooking the rich information provided by multi-view setups. To overcome this limitation, we propose MVMD, a novel Multi-View Mirror Detection method, along with the first database specifically designed for mirror detection in multi-view scenes. The design of MVMD is grounded in the inherent associations between objects seen from different views and those reflected inside and outside of mirrors. These relationships are learned through cross- and self-attention mechanisms. MVMD consists of three key blocks: the Inter-Views Block tracks the shifts of objects within mirrors caused by changes in viewpoint; the Intra-View Block detects object reflections inside mirrors; and the Refinement Block sharpens mirror boundaries and enhances detected details. Experimental results show that our method improves accuracy by up to 2.6% and IoU by up to 11.1%, compared to single-image mirror detection techniques. This substantial improvement makes MVMD particularly effective for computer vision tasks, especially in enhancing the accuracy of 3D reconstruction in mirror-dense environments.
[51] XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection cs.CVPDF
Rehan Ahmad, Gousia Habib, Muhammad Shaban, Ishfaq Ahmad Malik
TL;DR: 本文提出了一种名为XEns-CKD的新型集成视觉变换器方法,用于基于超声图像进行慢性肾病(CKD)分期检测。该方法通过集成三个在不同训练参数下训练的ViT模型,实现了86.36%的总体分类准确率,并利用可解释人工智能技术(如LIME、LRP、Attention-Min和Attention-Max)生成注意力图,以识别和解释受CKD进展影响的肾脏区域,从而提高了模型的透明度和临床可信度。
Details
Motivation: 慢性肾病是一种隐匿性疾病,早期分期检测有助于患者了解肾功能状况并遵循医疗建议以延缓疾病进展。现有方法在CKD分期准确性和模型可解释性方面存在不足,因此需要一种既能准确分类又能提供临床可解释结果的方法。
Result: 在私有超声图像数据集上,集成模型实现了86.36%的总体分类准确率,并使用宏观敏感性、特异性、精确度、F1分数、Youden指数、马修斯相关系数和宏观平衡准确率等指标进行评估。与现有方法相比,该方法在将五个CKD阶段和正常肾脏状态分类时,准确率提高了4%。
Insight: 创新点在于结合了集成视觉变换器架构与多种可解释AI技术,通过融合Attention-Min和Attention-Max结果生成注意力图,有效可视化CKD进展中受影响的肾脏区域,这不仅提升了分类性能,还增强了模型在临床环境中的透明度和可信度。
Abstract: Chronic kidney disease (CKD) is a silent disease. Its progression may not significantly hamper a person’s daily routine. Human kidney function can be classified as normal or as one of the five stages of CKD. Early detection of the CKD stage can help patients understand the functional status of their kidneys and follow medical advice to slow CKD progression. In this paper, we propose XEns-CKD, a novel ensemble vision transformer-based scheme for CKD stage classification using ultrasound images. Three ViTs were trained on a private ultrasound image dataset using different training parameters. The performance of each ViT was evaluated using macro sensitivity, macro specificity, macro precision, macro F1-score, macro Youden index, the Matthews correlation coefficient (MCC), and macro balanced accuracy. The ensemble model achieved an overall classification accuracy of 86.36%. This work also emphasizes identifying and interpreting kidney regions affected by CKD progression. Explainable artificial intelligence techniques, including LIME, LRP, Attention-Min, and Attention-Max, were used to improve model transparency and clinical trust. An attention map combining the Attention-Min and Attention-Max results effectively identified and interpreted kidney regions affected during CKD progression from one stage to another. The attention map also highlighted the effects of CKD progression in these regions. Compared with existing methods, the proposed method classified the five CKD stages and normal kidney status with a 4% improvement in accuracy.
[52] What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems cs.CV | cs.AIPDF
Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li
TL;DR: 本文提出了一种用于图像创作对话系统的视觉对齐后续编辑建议框架。该框架通过三个阶段(基于真实数据的有监督微调、利用用户点击反馈的多目标强化学习、引入视觉验证器)来生成符合用户偏好、多样化且在当前图像上可执行的编辑建议。
Details
Motivation: 现有对话助手主要针对纯文本交互,而图像创作任务中的后续编辑建议需要反映用户偏好、提供多样化方向并确保在当前图像上可执行,这是一个未被充分探索的多模态推荐问题。
Result: 在自动和人工评估中,该框架显著优于基线模型。在涉及数百万用户的线上A/B测试中,最终框架将视觉不一致性从3.7%降至0.9%,并显著提升了推荐点击率(+32.70%)、图像采纳率(+16.32%)和用户平均对话轮次(+39.90%)。
Insight: 创新点在于构建了一个结合多模态策略、用户反馈强化学习和视觉一致性验证的三阶段框架,以解决图像依赖型对话中的后续编辑建议问题,并通过真实线上数据验证了其有效性。
Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).
[53] Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation cs.CV | cs.LGPDF
Nafis Fuad, Xiaodong Qian, Dongxiao Zhu
TL;DR: 本文提出了三种基于机制可解释性指导的视觉语言模型,用于从街景图像中连续估计厘米级洪水深度。其中,FloodLlama-Dense 作为全微调基线,而 FloodLlama-MI5 和 FloodLlama-MI6 则是通过可解释性分析识别出关键交叉注意力层进行稀疏微调的变体,大幅减少了可训练参数量。
Details
Motivation: 解决城市洪水对交通基础设施的威胁,目前缺乏能够提供厘米分辨率、实时、街道级洪水深度估计的实用系统。
Result: 在合成数据集上,FloodLlama-Dense 的 MAE 为 0.40 cm,RMSE 为 1.97 cm,Acc@5cm 为 97.59%。稀疏变体 FloodLlama-MI5/MI6 在减少 86-88% 可训练参数的同时,精度损失极小。在真实世界基准测试中,FloodLlama-MI6 达到了 98.62% 的准确率,显著优于已发布的 STURM-FloodDepth 基线(86.61%)。
Insight: 核心创新点在于利用机制可解释性分析(线性探测、logit lens、CKA、交叉注意力熵)来指导模型微调,识别出视觉表征重组(L13-L22层)和深度信息首次线性可解码(L23层)的两阶段适应模式,从而实现了高效的稀疏微调策略。这为模型压缩和高效微调提供了新思路。
Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution. This paper presents three vision-language models fine-tuned for continuous flood-depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline, and FloodLlama-MI5 and FloodLlama-MI6, interpretability-guided sparse variants that fine-tune only the top five and six causally relevant cross-attention layers identified through mechanistic interpretability analysis, respectively. Training uses an approximately 610,000-image subset of a 2.81-million-image synthetic corpus generated in Unreal Engine 5. The dataset combines single-vehicle subsets with 5 cm depth increments and mixed-vehicle subsets with 1 cm depth increments, spanning seven vehicle types, four weather conditions, and flood depths from 0 to 40 cm. FloodLlama-Dense achieves an MAE of 0.40 cm, an RMSE of 1.97 cm, and an Acc@5cm of 97.59%. Mechanistic interpretability analysis combining linear probing, logit lens, centered kernel alignment (CKA), and cross-attention entropy reveals a two-stage adaptation pattern: layers L13-L22 restructure visual representations, while depth first becomes linearly decodable at layer L23. FloodLlama-MI5 and FloodLlama-MI6 leverage this insight by fine-tuning only five or six of the eight cross-attention layers, achieving an 86-88% reduction in trainable parameters (6.55-7.86 million versus 54.4 million) with minimal accuracy loss. On a real-world benchmark, FloodLlama-MI6 achieves 98.62% accuracy, compared with 86.61% for the published STURM-FloodDepth baseline.
[54] Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark cs.CV | cs.AIPDF
Marios Petrov, Sahana Vinayak, Targol Bakhtiarvand, Moses Smith Guddah, Adham Atyabi
TL;DR: 本文针对基于功能近红外光谱(fNIRS)的自闭症谱系障碍(ASD)分类任务,提出了一个跨时间窗口迁移问题,以解决因血流动力学延迟和神经血管耦合差异导致的时域分布偏移。研究通过改变时间窗口长度和偏移量构建基准,评估了多种视觉架构和适应策略,发现少量受试者特异性微调可显著提升性能,并证明短至2.5秒的窗口仍包含可区分的判别信息。
Details
Motivation: 现有fNIRS分类方法假设时间对齐评估,但实际中由于受试者间血流动力学延迟和神经血管耦合的差异,最优观察窗口存在个体差异,导致时域分布偏移并降低分类性能。
Result: 在124名受试者的留一受试者交叉验证中,零样本跨窗口准确率接近随机水平(54-69%);仅需约5%的受试者特异性微调即可恢复至90-96%的准确率;无目标受试者数据时,领域对抗和自监督策略达到78-90%的准确率;且从短至2.5秒的窗口中仍可恢复判别信息。
Insight: 创新点在于将fNIRS分类中的时间窗口变化问题形式化为跨时间窗口迁移问题,并建立了系统的评估基准。客观分析表明,研究揭示了受试者间变异性是性能下降的主要障碍,并为在实际时域变异下部署fNIRS分类器提供了实用路线图,强调了少量个性化适应的重要性。
Abstract: Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assume temporally aligned evaluation. In practice, the optimal observation window varies across subjects due to differences in hemodynamic delay and neurovascular coupling, creating a temporal distribution shift that degrades performance. We formalize this as a \textit{cross-time-window transfer problem}, introducing a protocol that varies window length (2.5–10,s) and offset within biological motion trials. Using topographic map representations of fNIRS recordings, we benchmark three vision architectures under two zero-shot baselines and eight adaptation strategies under leave-one-subject-out cross-validation ($N{=}124$). Key findings: (1) zero-shot cross-window accuracy is near chance (54–69%); (2) ${\approx}5%$ subject-specific fine-tuning recovers 90–96%, while a subject-specific upper bound reaches 97–100%, identifying inter-subject variability as the dominant barrier; (3) domain-adversarial and self-supervised strategies achieve 78–90% without target-subject data; and (4) discriminative information is recoverable from windows as short as 2.5,s. These findings provide a practical roadmap for deploying fNIRS-based ASD classifiers under realistic temporal variability.
[55] Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators cs.CV | cs.AI | cs.LGPDF
Bowen Xue, Jiafeng Xiong, Xin Quan
TL;DR: 本文提出了一种称为潜在频率有效性(LFV)的方法,用于在视频变分自编码器(VAE)的潜在空间中进行快速频谱编辑。该方法通过学习VAE特定的紧凑频谱响应,仅在能提高解码目标保真度且不恶化往返漂移时部署,从而避免了传统解码-滤波-重新编码流程。LFV通过验证选择路径,从对角每频率校准器(C1)到全通道混合(CM),使跨通道容量成为可控制的每编辑资源。
Details
Motivation: 动机在于解决在视频VAE潜在空间中进行直接频谱编辑时的问题,即VAE可能将像素空间频带重新分配到潜在通道中,且潜在编辑可能破坏VAE的往返动态。目标是实现高效控制噪声、闪烁、平滑度和频率内容,而无需解码-滤波-重新编码过程。
Result: 在涵盖六个频谱家族的544个VAE-编辑单元中,LFV生成了423个廉价算子,其中277个由C1处理,146个(34.5%)需要通道混合。在主要的120单元径向扫描中,99/100的生成算子通过了源视频分组保留评估;在另外五个滤波器家族中,所有323个生成算子均通过保留评估。完全冻结的OpenVid拟合算子(包括验证选择路径系数)无需适应即通过所有20个测试的CogVideoX和HunyuanVideo生成域单元。所选响应匹配直接潜在滤波延迟,比像素滤波-重新编码快约3倍。
Insight: 创新点在于引入潜在频率有效性(LFV)概念,通过验证选择路径动态调整频谱响应,平衡编辑效率和保真度。从客观角度看,该方法揭示了不同VAE的独特机制,如CogVideoX的强通道耦合响应和Open-Sora的高频带稳定性边界,为潜在空间编辑提供了可控且高效的解决方案。
Abstract: Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode–filter–reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE-specific spectral response and deploys it only when it improves decoded-target fidelity without worsening round-trip drift. LFV follows a validation-selected path from a diagonal per-frequency calibrator (C1) to full channel mixing (CM), making cross-channel capacity a controllable per-edit resource. Across 544 VAE–edit cells spanning six spectral families, LFV emits 423 cheap operators: 277 are handled by C1, while 146 (34.5% of emitted operators) require channel mixing. On the primary 120-cell radial sweep, 99/100 emitted operators pass source-video-grouped held-out evaluation. Across five additional filter families, all 323 emitted operators pass held-out evaluation. Fully frozen OpenVid-fitted operators, including the validation-selected path coefficient, pass all 20 tested CogVideoX and HunyuanVideo generated-domain cells without adaptation. The selected response matches direct latent-filter latency and is about $3\times$ faster than pixel filter–reencode. The resulting maps reveal distinct VAE regimes, including strongly channel-coupled CogVideoX responses and a sharp Open-Sora high-band stability frontier.
[56] COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping cs.CV | cs.AIPDF
Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li
TL;DR: 本文提出了COMEX基准和两阶段学习框架,用于解决可解释的美学图像裁剪问题。该工作将任务重新定义为结构化裁剪-构图-解释问题,通过图像扩展和IO反转流程构建了包含33,161个四元组的数据集,并采用SFT+GRPO框架联合学习裁剪定位、构图理解和解释生成。
Details
Motivation: 现有方法将解释视为事后文本生成,忽视了构图这一连接裁剪决策与可解释推理的关键美学因素,因此需要建立结构化框架来统一处理裁剪、构图与解释。
Result: 在COMEX基准和先前数据集上的实验表明,该框架在裁剪质量、构图预测和解释忠实度方面均表现出色,评估指标全面领先,验证了方法的有效性和可迁移性。
Insight: 创新点在于将可解释美学裁剪重构为结构化三元任务,并构建了首个基于构图的可解释裁剪基准;提出的SFT+GRPO两阶段框架能同时优化裁剪精度与解释一致性,为视觉语言模型提供了新的评估体系。
Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.
[57] A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions cs.CVPDF
Jianlin Ye, Christos Kyrkou
TL;DR: 本文综述了基于无人机的交通监控中视觉车辆检测的最新进展,重点探讨了深度学习模型在不同城市环境下的应用。文章分析了当前面临的三大挑战:与交通控制系统的兼容性、实时处理需求以及不同环境下的鲁棒检测,并指出了未来研究方向应集中在优化检测模型、边缘计算和自适应控制集成上。
Details
Motivation: 无人机监控为智能交通系统提供了覆盖广、实时数据收集的创新方案,但面临高度变化、运动补偿和高分辨率图像处理等挑战,需评估准确性、延迟和与现有系统的协调性。
Result: 综述未提供具体定量结果,但指出深度学习显著提升了检测精度,现有解决方案在全面利用无人机数据进行事件响应和交通管理方面仍缺乏框架。
Insight: 创新点在于系统性地识别了无人机交通监控的关键挑战,并强调未来需结合优化模型、边缘处理和自适应控制来提高城市交通管理的响应能力。
Abstract: In Intelligent Transportation System (ITS), unmanned aerial vehicle (UAV)-based surveillance offers an innovative solution to traffic surveillance with wide coverage and real-time data collection capabilities. In comparison to fixed ground-based infrastructure, UAVs are able to respond to dynamic traffic but present challenges such as vehicle detection at varying altitudes, compensation for motion-induced image variations and efficient processing of high-resolution images. Deep learning has been largely beneficial on improving the detection accuracy; however, for practical deployment, a critical assessment of the accuracy, latency, and harmonization with current transportation systems needs to be carefully considered. This survey reviews recent advancements in the UAV-based traffic monitoring, with a primary focus being deep neural network models for traffic analytics in various urban settings. Three main challenges identified in the literature are ensuring compatibility with traffic control systems, achieving real-time processing to optimize traffic flow, and maintaining robust detection in different environmental conditions. Existing solutions often lack comprehensive frameworks for utilizing UAV captured data to respond to incidents and manage traffic effectively. Future research should focus on optimal detection models, edge processing, and adaptive control integration to improve the responsiveness of urban traffic management.
[58] BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference cs.CV | cs.AIPDF
Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou
TL;DR: 本文提出了一种名为BRACE的新方法,用于加速扩散变换器(DiTs)的推理过程。该方法通过引入基于重心有理函数和切比雪夫增强的预测机制,来替代传统的基于导数的多项式外推,从而在保持高图像生成质量的同时,显著减少计算开销。
Details
Motivation: 现有基于缓存的预测方法(如基于导数的多项式外推)在进行长步预测时不稳定,容易导致生成质量严重下降,尤其是在高加速比下。论文的动机是解决DiT特征轨迹中存在的尖锐不规则性和局部非平滑性问题,以实现更稳定、高效的推理加速。
Result: 大量实验表明,BRACE在各种DiT架构上均实现了最先进的质量-效率权衡,其计算开销可忽略不计,在加速推理的同时保持了高保真度的生成质量。
Insight: 论文的核心创新在于将预测范式从导数驱动的多项式外推,转变为特征驱动的重心有理函数预测。该方法利用局部滑动窗口缓存稀疏历史特征,并采用自适应切比雪夫权重来构建稳定的预测函数,直接聚合原始特征以确保数值稳定性,有效应对了特征轨迹中的不规则性。
Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache-then-forecast methods driven by derivative-based polynomials often cause severe quality degradation under high acceleration due to unstable long-step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state-of-the-art quality-efficiency trade-offs across various DiT architectures with negligible computational overhead.
[59] Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion cs.CV | cs.ROPDF
Nethmi Pathirana, Isuru Munasinghe, Dileeka Alwis
TL;DR: 本文提出了一种用于皮肤病变分类的多模态框架,该框架结合了基于Swin Transformer的图像特征和结构化临床元数据,通过集成视觉-上下文学习来提高诊断性能。实验表明,该模型在公开数据集上取得了高准确率和F1分数,并通过温度缩放进行后处理校准以降低预期校准误差,同时结合不确定性估计和可解释性分析,为自动化皮肤病变分类提供了一种有效且可信的方法。
Details
Motivation: 解决皮肤镜图像中因类别不平衡、类间相似性和类内变异性导致的自动分析挑战,旨在通过融合多模态信息来提升皮肤病变分类的诊断性能。
Result: 在公开数据集上,所提模型取得了92.55%的测试准确率和91.33%的宏F1分数,在少数类别上表现强劲;应用温度缩放后降低了预期校准误差,提高了预测可靠性。
Insight: 创新点在于将Swin Transformer提取的图像特征与临床元数据进行多模态融合,并系统性地结合了后处理校准、不确定性估计和定性可解释性分析,以构建一个更可靠、可解释的皮肤病变分类系统。
Abstract: Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due to class imbalance, inter-class similarity, and intra-class variability in dermoscopic images. This paper proposes a multimodal classification framework that combines Swin Transformer-based image features with structured clinical metadata to improve diagnostic performance through integrated visual-context learning. Experiments on a publicly available dataset show that the proposed model achieves a test accuracy of 92.55% and a macro F1-score of 91.33%, with strong performance across minority classes. Temperature scaling is applied as a post-hoc calibration method, resulting in a reduction in expected calibration error and improving prediction reliability, while uncertainty estimation is incorporated to further assess the confidence of model predictions. Qualitative explainability analysis further shows that the model focuses on lesion regions during inference. Therefore, the results demonstrate that multimodal fusion, combined with calibration and interpretability analysis, provides an effective and trustworthy approach for automated skin lesion classification.
[60] Multi-Branch Policy Optimization for Multimodal Large Language Models cs.CV | cs.AIPDF
Shuai Lyu, Yuning Gong, Ruiling Gao, Xiaoran Shang, Zhonghong Ou
TL;DR: 本文提出了一种名为多分支策略优化(MBPO)的树形框架,用于解决多模态大语言模型在强化学习中面临的感知不确定性问题。该方法通过在视觉-语言决策边界构建推理树,允许兄弟分支探索不同的视觉假设,并通过分支相对优势进行分段级信用分配,同时引入时间回放缓冲区以重用信息片段并控制策略陈旧性。
Details
Motivation: 针对多模态大语言模型在基于群体的强化学习方法中,通常依赖轨迹级信用分配(即对响应中的所有令牌应用单一优势)的不足,特别是在多模态推理中感知不确定性远高于纯文本设置,模型需要反复重新检查视觉信息以验证中间解释,且不同的视觉基础可能导致不同的推理路径,使得这种统一的信用分配尤其不充分,导致相对优势逐渐退化至零。
Result: 在多个多模态推理基准测试上的实验表明,MBPO优于代表性的基线方法,提高了学习信号质量和优化效率。
Insight: 创新点在于提出了树形框架MBPO,通过构建推理树实现分段级信用分配,解决了多模态推理中感知不确定性导致的信用分配退化问题,并引入时间回放缓冲区以提升效率,为多模态大语言模型的强化学习优化提供了新思路。
Abstract: Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially higher perceptual uncertainty than text-only settings, where the model must repeatedly re-examine visual information to verify intermediate interpretations, and different visual groundings can lead to divergent reasoning paths, making such uniform credit assignment particularly inadequate and causing relative advantages to progressively degenerate toward zero. To address these challenges, we propose Multi-Branch Policy Optimization (MBPO), a tree-based framework that constructs reasoning trees at vision-language decision boundaries, enabling sibling branches to explore diverse visual hypotheses and assigning segment-level credit through branch-relative advantages. We further introduce a temporal replay buffer to reuse informative segments while controlling policy staleness. Experiments on several multimodal reasoning benchmarks show that MBPO outperforms representative baselines, improving both learning signal quality and optimization efficiency. The code is publicly available at https://github.com/ShuaiLyu0110/MBPO.
[61] Predictive Failure Detection in Network Hardware Using Thermal Imaging and Deep Learning with Sensor Fusion cs.CVPDF
Ashly Joseph
TL;DR: 本文提出了一种基于深度学习的预测性维护策略,利用热成像和功率传感器数据来检测路由器、交换机和服务器等网络硬件的早期故障迹象。通过生成包含标注热图像和功率读数的模拟数据集,评估了多种CNN模型以及一个融合视觉和传感器时序信息的多模态CNN-LSTM模型。实验表明,基于感兴趣区域(ROI)的预处理和多模态融合显著提升了故障预测的准确性。
Details
Motivation: 解决数据中心中网络硬件意外故障导致服务中断和昂贵停机时间的问题,旨在通过非侵入式监测实现主动维护。
Result: 在模拟数据集上,未经预处理的CNN模型准确率较低(如ResNet-50为52%),而基于ROI的预处理使ResNet-50准确率提升至91%。CNN-LSTM多模态融合模型取得了最佳性能,准确率达到94%,精确率和召回率接近95%,验证了该方法在早期故障预测上的有效性。
Insight: 创新点在于将热成像视觉数据与功率传感器时序数据通过多模态CNN-LSTM进行融合,并结合领域特定的ROI预处理,显著提升了预测性能。这为基于非侵入式监测的预测性维护提供了可行的技术路径。
Abstract: Unplanned network hardware malfunctions can interrupt services and result in expensive downtime in data centers. A deep learning-based predictive maintenance strategy is presented that utilizes thermal imaging and power sensor data to detect early indicators of equipment breakdown in routers, switches, and servers. A simulated dataset was generated comprising annotated thermal pictures and power readings indicative of three operating states: Normal, Warning, and Critical. Three ImageNet-pretrained convolutional neural network (CNN) models ResNet-50, InceptionV3, and VGG16 were assessed together with a multi-modal CNN-LSTM fusion model that integrates visual and sensor time-series information. Experiments were performed with and without pre-processing procedures, including region-of-interest (ROI) extraction and normalization. In the absence of pre-processing, CNNs attained moderate accuracy (e.g., ResNet-50 at 52%), but ROI-based pre-processing significantly enhanced performance (ResNet-50 accuracy reaching 91%). The CNN-LSTM model attained the greatest accuracy of 94%, with precision and recall approaching 95%, illustrating the effectiveness of multi-modal fusion. The results validate that domain-specific pre-processing and sensor fusion substantially improve early failure prediction, providing a potential foundation for proactive maintenance of network hardware through non-intrusive monitoring.
[62] ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making cs.CVPDF
Ningxin Pan, Hanyu Li, Yehui Tang
TL;DR: 本文提出了COMPLEXITYWORLD基准,用于评估视觉语言模型在可验证的视觉决策任务上的能力。该基准包含390个任务,涵盖39个视觉领域和29个决策类别,每个任务都基于隐藏的结构化规范生成,并通过可执行的验证器来评分。实验发现,除了GPT-5.6-Sol模型达到75.6%的验证器接受率外,其他模型在直接推理下的表现均低于40%,揭示了视觉到决策的瓶颈问题。
Details
Motivation: 当前视觉语言模型在视觉感知方面进展迅速,但许多现实任务不仅需要识别图像内容,还要求模型利用视觉证据做出满足全局约束的完整决策。现有模型在此类复杂决策任务上的能力尚不明确,因此需要专门的基准进行评估。
Result: 在COMPLEXITYWORLD基准上,除GPT-5.6-Sol达到75.6%的验证器接受率外,其他评估模型在直接推理下的表现均低于40%。当决策信息以结构化形式明确提供时,性能有显著提升,但不同视觉呈现方式下的表现差异很大。
Insight: 论文的创新点在于构建了一个基于可验证决策的综合性视觉基准,强调结构化规范和可执行验证。客观来看,该研究揭示了视觉语言模型在从视觉输入到结构化决策推理中存在显著瓶颈,且性能对视觉呈现方式敏感,这为未来模型设计提供了重要方向。
Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global constraints. We introduce COMPLEXITYWORLD, a benchmark of 390 tasks across 39 domain-inspired visual worlds and 29 decision categories. Each task is generated from a hidden structured specification, rendered as a visual scene, and scored by an exe- cutable verifier that accepts any feasible solution. Under direct inference, all evaluated models ex- cept GPT-5.6-Sol remain below 40% verifier ac- ceptance rate (VAR), while GPT-5.6-Sol reaches 75.6%. Performance improves substantially when the same decision information is made explicit in structured form, yet varies sharply across equiva- lent visual presentations. Agent scaffolds provide smaller, model-dependent gains. Together, these results reveal a persistent visual-to-decision bot- tleneck that additional inference alone does not remove.
[63] LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents cs.CV | cs.LG | cs.MAPDF
Zijian Wang, Junnan Zhu, Rongzhen Li, Xiao Liu, Guohui Xiang
TL;DR: 本文提出了LAVE(Latent Visual Evidence-Enhanced Planning)框架,旨在解决长视频理解中视频工具使用智能体面临的工具观察瓶颈问题。该框架通过引入一个双通道观察接口(可见的文本轨迹通道和潜在的视觉证据通道),在无需额外训练、帧重放或修改原有编排的情况下,使智能体能够重用已完成的工具调用中未在文本中描述的潜在视觉证据,从而增强多步规划能力。
Details
Motivation: 现有视频工具使用智能体在规划时通常仅依赖文本化的工具观察结果,这会导致未在文本中描述的视觉证据被丢弃,形成工具观察瓶颈,限制了智能体从长且冗余的视频流中有效获取和重用稀疏视觉证据的能力。
Result: 在Video-MME、LongVideoBench和CG-Bench等多个基准测试上的广泛实验表明,LAVE能持续提升不同骨干网络的视频工具使用智能体的性能。在可比的帧预算下,LAVE将Video-MME的整体得分比最强基线提高了3.76分,证明了潜在视觉证据重用的有效性。
Insight: 核心创新点是提出了一个无需训练的双通道观察接口,通过存储并检索与当前规划状态相关但未被文本覆盖的潜在视觉证据(包括工具角色、源帧时间戳和视觉位置),并利用有界时间戳对齐的潜在更新和熵约束的帧时间路由进行整合,从而突破了传统纯文本接口的瓶颈,实现了对已有视觉计算的高效重用。
Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this challenge by iteratively invoking visual Tools at different temporal scales, but their Tool-Planner communication typically relies on textual observations. Such text-only interfaces provide lossy summaries of Tool computations, causing previously computed visual evidence not verbalized to be discarded and unavailable for subsequent planning. We identify this limitation as the Tool observation bottleneck and propose Latent Visual Evidence-Enhanced Planning (LAVE), a training-free framework for reusing latent visual evidence from completed Tool calls. LAVE introduces a dual-channel observation interface: the visible channel preserves the original textual trajectory, while the latent channel stores pre-verbal visual updates with their Tool roles, source-frame timestamps, and visual locations. During planning, LAVE retrieves evidence relevant to the current Planner state but not covered by textual observations, and integrates it through bounded timestamp-aligned latent updates with entropy-constrained frame-time routing. This enables video agents to reuse existing visual computation without additional training, frame replay, or modifications to the original orchestration. Extensive experiments on Video-MME, LongVideoBench, and CG-Bench show that LAVE consistently improves video tool-use agents across backbones. Under a comparable frame budget, LAVE improves the Video-MME overall score by 3.76 points over the strongest baseline, demonstrating the effectiveness of latent visual evidence reuse for multi-step video-agent planning.
[64] HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers cs.CVPDF
Dong Liu, Yanxuan Yu, Renata Borovica-Gajic, Ying Nian Wu
TL;DR: 本文提出HSMLA(分层Softmax多尺度线性注意力)机制,旨在解决视觉Transformer在高分辨率密集预测任务中因自注意力二次复杂度导致的计算开销问题。该方法结合了基于ReLU的线性注意力来捕获全局上下文,采用选择性Softmax细化关键局部特征,并通过深度卷积实现多尺度令牌表示,从而在多个密集预测任务上实现了显著的精度与效率权衡。
Details
Motivation: 视觉Transformer在高分辨率密集预测任务中面临自注意力二次复杂度带来的巨大计算开销,而现有线性注意力方法虽高效却牺牲了局部上下文建模能力,因此需要一种能兼顾全局与局部特征的高效注意力机制。
Result: HSMLA在多个密集预测任务上实现了优越的精度-效率权衡:推理速度最高提升4.2倍;在CT器官分割任务上达到87.3%的Dice系数且速度提升3.2倍;在病理全切片图像分析中达到94.2%的AUC且速度提升4.1倍。
Insight: 创新点在于将线性注意力与选择性Softmax细化相结合,以线性复杂度实现全局上下文建模的同时,通过Softmax增强关键局部特征;并引入多尺度令牌表示来提升特征丰富性,这为设计高效视觉Transformer提供了可借鉴的架构思路。
Abstract: Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficiency but sacrifices local context modeling. We propose \textbf{HSMLA (Hierarchical Softmax Multi-scale Linear Attention)}, which combines ReLU-based linear attention for global context, selective softmax refinement for critical local features, and multi-scale token representations via depthwise convolutions. HSMLA achieves superior accuracy-efficiency trade-offs: up to $4.2\times$ inference-time speedup across dense prediction tasks, $87.3%$ Dice with $3.2\times$ speedup on CT organ segmentation, and $94.2%$ AUC with $4.1\times$ speedup on pathology WSI.
[65] Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting cs.CV | cs.LGPDF
Lucas Gouveia Omena Lopes, William W. M. Lira, Alexandre M. Lima, Thales M. A. Vieira
TL;DR: 本文提出了一种基于几何方法的、类别无关的嵌入式车辆计数系统,用于高速公路数据采集。该系统通过背景减除和阈值处理检测运动物体,利用软件感应线圈检测器进行计数,无需对象模型或训练数据,可在树莓派等单板计算机上实时运行。
Details
Motivation: 解决传统基于深度学习的车辆检测方法依赖预训练检测器的问题,提出一种无需类别标注、计算资源要求低的几何方法,适用于实际部署中缺乏标注数据或计算资源受限的场景。
Result: 在四个视频上,自校准预校准规则的计数准确率达到83.3%至100%;现场部署中准确率为91%,而相同计算预算下的基线方法仅为37.5%。
Insight: 创新点包括类别无关的几何检测方法、自校准规则从统计中恢复车道几何信息,以及系统在低功耗、隐私保护和开放集类别场景下的适用性。
Abstract: Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing in practice: a detector already trained on the class one wants to count. We revisit a purely geometric traffic-sensing pipeline for Single Board Computers in which detection is class-agnostic: moving objects come from background subtraction and thresholding, and counting is decided by a geometric rule on an imaginary line across the road, a software inductive loop detector. With no object model, training set or per-object trajectory, it runs faster than real time on Raspberry Pi class hardware. Two counting rules are described: a constant average speed rule, whose expected accuracy is derived analytically as about 86% under a Gaussian speed distribution, and a self-calibrating pre-calibration rule that recovers the lane geometry from blob statistics and counts edges of lane occupancy, additionally yielding per-vehicle average speed at no extra cost. Over four videos the latter counts with 83.3%-100% accuracy; in a field deployment it reaches 91% against 37.5% for a blob-tracking baseline under the same compute budget. We report the observations of that period in detail: the resolution floor below which accuracy collapses, the frame rate floor at which vehicles alias past the counting line, the gap between short curated clips and long uncontrolled footage, and the trade-off between Python (easier to tune, 100% CPU) and C++ (40% CPU, thermally viable). These are properties of the sampling geometry, not of the hardware of the time, and still constrain edge deployments. We close by arguing where motion-based, class-agnostic detection remains the right tool: open-set classes with no annotated data, tight power budgets, privacy-constrained installations, and the cold start of mining training crops to bootstrap a learned detector.
[66] Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding cs.CV | cs.AI | cs.CLPDF
Yeeun Choi, Youngbeom Yoo, Joon-Young Lee, Hyolim Kang, Seon Joo Kim
TL;DR: 本文提出MERIT框架,用于解决超长视频理解中直接端到端处理不可行的问题。该框架采用两阶段范式:先构建与查询无关的记忆,再基于检索进行推理。通过简单的多键匹配机制和推理时邻域过滤,实现了高效且高性能的视频理解。
Details
Motivation: 针对当前多模态大语言模型无法直接处理超长视频(数小时至数天)的挑战,现有方法在记忆构建阶段过度复杂且未考虑下游查询。本文旨在优先保证记忆构建阶段的高召回率,并将查询特定的高层关系组合推迟到推理时处理。
Result: MERIT在三个长视频基准测试(EgoLifeQA、LVBench和Video-MME (Long))上取得了最先进的性能。
Insight: 创新点在于提出了一种简单的多情景多键表示,通过键匹配机制实现细粒度记忆的精确检索,并引入推理时邻域过滤机制,在无需全局记忆构建的高计算开销下捕获更广泛的语义上下文。客观来看,其将复杂关系建模延迟到推理阶段的思路,平衡了效率与效果。
Abstract: When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).
[67] CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting cs.CV | cs.AI | cs.CLPDF
Quang Minh Dinh, Tuan Kiet Doan
TL;DR: CosmosAlign是一个基于预训练世界基础模型Cosmos3-Nano构建的生成式交通视频预测框架。它通过两阶段LoRA适应策略对齐条件分布和训练描述,并在推理时使用无需训练的基于共识的样本选择和运动自适应混合技术来提升预测质量。该框架在AI City Challenge 2026 Track 5基准测试中以76.49的最终得分排名第一。
Details
Motivation: 论文的动机是观察到将大型预训练世界模型成功适应下游预测任务主要依赖于分布对齐,而非增加模型容量。因此,旨在通过有效的适应策略,将预训练的世界基础模型专门用于生成式交通视频预测任务。
Result: 在AI City Challenge 2026 Track 5基准测试中,CosmosAlign取得了76.49的最终得分,在最终排行榜上排名第一,达到了SOTA水平。
Insight: 创新点包括:1)提出两阶段LoRA适应策略,分别对齐条件模式分布和通过LLM重述管道使训练描述与模型原生结构化提示接口对齐;2)在推理阶段引入无需训练的基于共识的中位数样本选择和运动自适应静态区域混合程序,以提升预测质量。从客观角度看,其核心洞察在于强调分布对齐而非模型扩展对于适应预训练基础模型的关键作用,并提供了一套系统性的适应和推理优化方法。
Abstract: Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video forecasting framework built upon the pretrained Cosmos3-Nano world foundation model. Our approach is motivated by the observation that successfully adapting large pretrained world models to downstream forecasting tasks depends primarily on distribution alignment rather than increased model capacity. To this end, we propose a two-stage LoRA adaptation strategy that first aligns the conditioning-mode distribution with the target forecasting task, and then aligns the training captions with the model’s native structured prompting interface through an LLM-based re-captioning pipeline. During inference, we further improve prediction quality using a fully training-free procedure consisting of consensus-based medoid sample selection and motion-adaptive blending of static scene regions. CosmosAlign achieves a final score of 76.49 on the AI City Challenge 2026 Track 5 benchmark, ranking first on the final leaderboard. Our code is publicly available at https://quangminhdinh.github.io/CosmosAlign/.
[68] SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models cs.CV | cs.AIPDF
Ziqiao Yu
TL;DR: 论文提出了一种名为SpikeWorld的稀疏脉冲模型,该模型在部署时冻结所有训练参数,仅通过外部路径(基于累积固定损失库的动作校正和基于特定路径残差矩阵的状态预测细化)进行快速状态适应,从而在不更新模型权重的情况下提升预测性能和策略奖励。
Details
Motivation: 解决在部署后,当动态模型和语义模型共享参数时,难以利用自监督信号进行持续适应的问题:冻结参数会阻碍适应,而更新权重则需要优化器状态并可能破坏已学习的表示。
Result: 在联合优化下,动作下一状态MSE提升了17.10%,多模态预测、语义准确性和图像-文本检索也有所改善。在保留的剪切和衰减数据流上,组合外部状态分别将总体预测提升了5.48%和30.01%,其固定库动作路径分别将跟踪性能提升了24.20%和3.94%。在包含450条新Meta-World轨迹的六臂研究中,冻结策略的奖励提升了7.90点。
Insight: 核心创新在于部署时冻结所有训练参数,仅通过轻量级的外部路径(不依赖标签、教师输出、奖励或真实状态值)进行快速适应,实现了模型表示的稳定性和适应能力的结合。该方法表明,将线性识别等简单适应机制与冻结的多模态脉冲检查点集成,可以显著提升性能。
Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficult when dynamics and semantics share parameters: freezing prevents adaptation, whereas weight updates require optimizer state and may alter the learned representation. Here we introduce SpikeWorld, a 1.45M-parameter sparse spiking model jointly trained for heterogeneous sensory prediction, semantics, image-text binding and action-conditioned dynamics. At deployment, all trained parameters are frozen. Delayed next-state residuals update two external paths: cumulative fixed-bank losses select the bounded action correction, while route-specific residual matrices refine next-state prediction. Neither path uses labels, teacher outputs, rewards, success signals or the true shift value. Joint optimization improves action next-state MSE by 17.10% while also improving multimodal prediction, semantic accuracy and image-text retrieval. On held-out shear and attenuation streams, the combined external state improves aggregate prediction by 5.48% and 30.01%; its fixed-bank action path improves tracking by 24.20% and 3.94%, respectively. In a six-arm study comprising 450 new Meta-World trajectories (75 per arm), SpikeWorld raises frozen-policy reward by 7.90 (95% CI [2.48, 14.06]); the 13.33-point success difference is descriptive (CI [0, 40]). For identical sensory inputs, model parameters and inherited semantic outputs remain bitwise unchanged. A 16-byte RLS estimator obtains the highest non-oracle reward on linear attenuation, showing that the contribution is not superior linear identification, but its integration with a frozen multimodal spiking checkpoint. Reference code is publicly available at https://github.com/Oooorca/SpikeWorld.
[69] Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties cs.CVPDF
Ali Bahri, Hongliang Li, Soufiane Lamghari, Jie Chuai, Zhitang Chen
TL;DR: 本文提出ViWi框架,通过结合视觉与WiFi射频描述符,从图像中估计物体的体积力学属性(如杨氏模量、泊松比和密度)。该方法利用材料槽聚合共享材料的体素证据,并引入基于电磁模拟的RF描述符来补充视觉信息,以解决仅凭视觉的模糊性问题。
Details
Motivation: 仅凭视觉估计体积力学属性存在本质模糊性,因为视觉相似的物体可能具有不同的材料组成和物理行为;现有方法独立预测各体素属性,忽略了物体的分段恒定材料结构,且缺乏解决视觉模糊性的显式机制。
Result: 在体积力学属性和质量估计基准测试中,ViWi在六个每体素指标中的四个上超越了先前的最先进方法,而其仅视觉变体在所有质量估计指标上均有提升。
Insight: 创新点在于引入以物体为中心的材料槽表示来聚合共享材料体素,并结合WiFi频段电磁模拟生成的RF描述符提供全局组成线索,从而实现了更准确且物理一致的体积属性估计,超越了仅依赖视觉的方法。
Abstract: Estimating volumetric mechanical properties, including Young’s modulus, Poisson’s ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. Across volumetric mechanical-property and mass-estimation benchmarks, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all mass-estimation metrics. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.
[70] BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning cs.CVPDF
Saim Rehman, Muhammad Shafique
TL;DR: 本文提出了BRUCE,一个用于评估科学视觉语言推理模型在图像质量退化下鲁棒性的基准框架。该框架通过向输入图像施加多种扰动(如模糊、低对比度),并引入RCI和T-RCI两个新指标来量化模型性能随扰动强度增加而恶化的速度。研究在化学和数学推理任务上进行了评估,并对预测失败进行了可解释的细粒度分析。
Details
Motivation: 现有的视觉语言模型在现实世界中常因输入图像质量低或变化而面临鲁棒性问题,而当前评估框架主要关注干净任务的准确性,缺乏对模型在鲁棒性维度上推理稳定性如何退化的系统性分析。
Result: 研究在多个数据集(化学和数学推理任务)上评估了BRUCE框架,并利用提出的RCI和T-RCI指标量化了模型性能随视觉损坏严重程度增加而恶化的速率,实现了对模型鲁棒性的系统性评测。
Insight: 创新点在于提出了一个系统性的多模态推理脆弱性评估框架BRUCE,并引入了RCI和T-RCI两个量化鲁棒性退化的新指标。该框架还将预测失败归类为OCR依赖推理、空间推理、符号推理和语义失败四个高层推理域,并进一步细分为特定损坏类型的子类,从而实现了可解释的失败分析。
Abstract: Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs’ robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation, a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics – Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) – to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.
[71] LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks cs.CV | cs.AI | cs.LGPDF
Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha
TL;DR: 本文提出LoRSA,一种用于生物医学下游任务的参数高效微调框架。它通过联合学习一个密集的低秩组件和一个动态结构化稀疏的低秩组件,将适应能力组织为全局和残差两条路径,旨在提升模型在未见成像域上的泛化能力。
Details
Motivation: 现有参数高效微调方法(如单一低秩更新)将所有任务特定变化限制在一个狭窄的参数子空间中,这可能阻碍模型同时表示全局共享的任务结构和泛化到未见域所需的局部残差方向。
Result: 在基于DINOv3-Base的四类乳腺密度分类任务中,LoRSA在内部验证集上保持竞争力,并在MammosighTR和RSNA两个未见外部域上取得了最佳的外部宏F1分数,分别比最强竞争方法提高了2.15和3.09个百分点。
Insight: 核心创新是将适应能力分解为全局协调的密集低秩组件和提供互补残差校正的结构化稀疏低秩组件。分析表明两个组件的更新方向具有高度互补性,这种全局-残差路径的组织方式能有效提升模型的外部域泛化性能。
Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. This restriction may prevent the model from simultaneously representing globally shared task structure and localized residual directions required for generalization to unseen imaging domains. We introduce LoRSA, a global–residual adaptation framework that jointly learns a dense low-rank component and a dynamically structured-sparse low-rank component. The dense component captures globally coordinated task adaptation, while the structured component provides complementary residual corrections whose support evolves during training. We characterize the representational capacity, approximation properties, rank structure, and singular-subspace complementarity of this decomposition. We evaluate LoRSA for four-class breast-density classification using DINOv3-Base, with VinDr-Mammo as the source domain and MammosighTR and RSNA as unseen external domains. LoRSA remains competitive on the internal validation set and achieves the best external macro-F1 on both target datasets, improving upon the strongest competing method by 2.15 percentage points on MammosighTR and 3.09 percentage points on RSNA. Weight-matrix analysis further shows that approximately $92%$ of the energy of each adaptation component lies outside the bilateral singular subspace of the other, indicating that the two components learn largely complementary update directions. These results suggest that organizing adaptation capacity into distinct global and residual paths can improve the external-domain generalization of parameter-efficiently adapted biomedical vision models.
[72] Multi-Task Consistency-based Detection of Adversarial Attacks cs.CV | cs.AIPDF
Cong Chen, Jean-Philippe Monteuuis, Jonathan Petit
TL;DR: 本文提出了一种基于多任务一致性的对抗攻击检测方案,通过利用复杂视觉系统中多个视觉任务(如目标检测和实例分割)推理输出之间的不一致性来检测对抗性扰动。该方法设计了一致性评分指标来衡量任务间的不一致性,并选择最佳模型对以有效检测不一致性。在BDD100k验证数据集上针对PGD攻击的评估表明,该防御方案在考虑的攻击者模型下实现了99.9%的ROC-AUC检测性能。
Details
Motivation: 深度神经网络在视觉感知系统中广泛应用,但其对对抗攻击的脆弱性引发了实际应用(如自动驾驶)中的担忧。现有防御方法往往成本效率低下,难以在资源受限的场景中部署。
Result: 在BDD100k验证数据集上,针对PGD攻击对多个视觉模型进行评估,实验结果显示该防御方案在考虑的对抗攻击模型下实现了99.9%的ROC-AUC检测性能,表现出高效性。
Insight: 创新点在于利用多任务感知系统(如目标检测与实例分割)输出间的一致性作为检测对抗攻击的信号,通过设计一致性评分指标和优化模型对选择,实现了高效且低成本的防御方案,为资源受限应用提供了新思路。
Abstract: Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attacks has prompted concerns regarding their practical applications, specifically in the context of autonomous driving. Existing defenses often suffer from cost inefficiency, rendering their deployment impractical for resource-constrained applications. In this work, we propose an efficient and effective adversarial attack detection scheme leveraging the multi-task perception within a complex vision system. Adversarial perturbations are detected by the inconsistencies between the inference outputs of multiple vision tasks, e.g., object detection and instance segmentation. To this end, we developed a consistency score metric to measure the inconsistency between vision tasks. Next, we designed an approach to select the best model pairs for detecting inconsistencies effectively. Finally, we evaluated our defense against PGD attacks across multiple vision models on the BDD100k validation dataset. The experimental results demonstrated that our defense achieved a ROC-AUC performance of 99.9% detection within the considered attacker model.
[73] XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting cs.CV | cs.GRPDF
Zhongpai Gao, Benjamin Planche, Meng Zheng, Anwesa Choudhuri, Chaoyi Zhou
TL;DR: XClipGS是一种用于医学体积高斯泼溅的精确半空间裁剪方法,解决了传统高斯泼溅在渲染时无法精确裁剪高斯基元的问题。该方法通过局部仿射模型将半空间限制的高斯积分分解为普通2D足迹和条件高斯CDF,实现了无学习参数的闭式像素级裁剪算子,并利用多距离参考视图监督内部结构。
Details
Motivation: 传统高斯泼溅代理在交互式渲染医学体积扫描时,裁剪平面会暴露未受外部视图训练约束的解剖结构,且只能整体保留或丢弃相交的基元,导致渲染不精确。
Result: 在八个CT和MRI体积数据上,XClipGS在未用于训练的平面偏移上取得了最高PSNR(33.56 dB vs ClipGS的32.34 dB),渲染速度超过650 FPS,远高于实时要求。在体素轴切面视图上,平均带SSIM从0.809提升至0.860,泄漏减少约40倍。
Insight: 创新点在于将裁剪问题分解为渲染时裁剪算子和隐藏内部监督两个独立问题,提出闭式像素级裁剪算子,无需学习参数或辅助网络,并引入剪裁/未剪裁切面对协议与差异参考误差指标,有效聚焦平面附近错误。
Abstract: Gaussian-splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external-view training and intersects primitives that conventional splatting can only keep or drop whole. We present XClipGS (eXact Clipping), which treats these as two separate problems: the render-time clip operator and supervision of the hidden interior. Under the local affine model used by EWA splatting, the ray integral of a half-space-restricted Gaussian factorizes exactly into its ordinary 2D footprint and a conditional Gaussian CDF whose argument is affine in pixel coordinates. The resulting closed-form per-pixel operator introduces no learned clipping parameters or auxiliary network and remains differentiable with respect to the primitive and plane. We use multi-distance reference views with varied clipping-plane axes and offsets to supervise the interior through the same operator. We also introduce a paired clipped/unclipped cut-face protocol with difference-referenced cut error (CDE) and culled-side leakage (Leak), because global image metrics dilute errors near the plane. On eight CT and MRI volumes with plane offsets not used for training, XClipGS attains the highest PSNR on every volume (33.56 versus 32.34 dB for ClipGS) while rendering at over 650 FPS, far above real time, versus 278 FPS. On voxel-axis cut-face views, it raises average band SSIM from 0.809 to 0.860 and leaks roughly 40 times less. Without retraining, it also achieves the best average across all four metrics on arbitrary-normal planes; on a fixed interior, it matches RaRa’s face fidelity with about 16 times less leakage. Project page: https://gaozhongpai.github.io/XClipGS/
[74] DINO-3DRA: Leveraging 2D Foundation Model Semantics for 3D Cerebral Aneurysm Segmentation cs.CVPDF
Jiayang Lu, Fengming Lin, Alejandro F. Frangi, Ali Sarrami-Foroushani
TL;DR: DINO-3DRA是一种用于3D旋转血管造影(3DRA)中脑动脉瘤分割的双路径框架。它通过创新的空间混合与残差融合机制,将预训练的2D视觉基础模型(DINOv3)的语义特征有效注入3D U-Net主干网络,以解决3D分割中存在的类别极度不平衡、与血管形态相似以及缺乏大规模3D预训练数据的问题。
Details
Motivation: 解决3DRA中动脉瘤分割面临的三大挑战:极端的类别不平衡、动脉瘤与血管的形态相似性,以及缺乏大规模3D预训练模型。同时,探索如何有效利用从海量2D图像中学习到的、富含结构先验知识的2D视觉基础模型来提升3D分割性能,避免简单的逐切片迁移方法破坏解剖结构连续性并导致优化不稳定。
Result: 在多中心3DRA数据上,DINO-3DRA实现了最先进的动脉瘤分割性能(Dice系数:0.758;HD95:2.75 mm),比nnU-Net提升了13%,且仅需570万个可训练参数。在未对CADA和SHINY-ICARUS数据集进行微调的情况下,该方法消除了基线架构中观察到的所有灾难性失败案例,展现了跨异构成像协议的强大泛化能力。
Insight: 论文的核心创新点在于提出了一个结构化的跨维度语义迁移框架,通过Room-Lite空间混合和校准残差融合机制,将冻结的2D基础模型特征桥接到3D网络中,从而有效利用了2D模型的丰富先验知识并保持了3D解剖结构的连续性。这为将2D基础模型的能力迁移到数据稀缺的3D医学影像任务提供了一种有效范式,其增益主要源于结构化的特征迁移而非损失函数设计。
Abstract: Accurate aneurysm segmentation in 3D rotational angiography (3DRA) is hindered by extreme class imbalance, morphological similarity to vessels, and absent large-scale 3D pretraining. 2D vision foundation models encode dense structural priors from 1.7 billion images, yet naïve slice-wise transfer fragments anatomical continuity and destabilises optimisation. We propose DINO-3DRA, a dual-path framework achieving effective cross-dimensional semantic transfer by injecting frozen DINOv3 features into a 3D U-Net backbone via Room-Lite spatial mixing and calibrated residual fusion. On multi-centre 3DRA data, DINO-3DRA achieves state-of-the-art aneurysm segmentation (Dice: 0.758; HD95: 2.75 mm; +13% over nnU-Net) with only 5.72M trainable parameters. Ablation studies confirm that gains arise from structured cross-dimensional transfer rather than loss design alone, with bridged foundation features improving anatomical continuity between aneurysms and parent vessels. Without fine-tuning on CADA and SHINY-ICARUS, DINO-3DRA eliminates all catastrophic failure cases observed in baseline architectures, demonstrating robust generalisation across heterogeneous imaging protocols.
[75] How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems cs.CV | cs.HC | cs.IR | cs.MMPDF
Henri Vanhuynegem, Weitao Xu, Yiran Shen, Guohao Lan
TL;DR: 该论文提出了VQABench,这是首个系统性地将客户端输入预处理作为可控变量来评估基于云端视觉语言模型(VLM)的视觉问答(VQA)系统的基准测试。研究评估了12种预处理技术在三个VQA数据集和四个商业VLM上的表现,总计进行了95,168次API调用,揭示了预处理效果受模型、API范式、提供商计费规则和任务形式影响,并非总是有益。
Details
Motivation: 随着VLM成为移动VQA系统的实用后端,由于难以在移动设备上运行,系统常将推理卸载到云端VLM。客户端视觉输入预处理成为关键系统变量,影响答案质量、负载大小、令牌成本和延迟,但现有预处理技术对商业云VLM的成本-质量影响缺乏研究。
Result: 在三个VQA数据集和四个商业VLM上的评估表明,预处理并非普遍有益;其效果取决于目标模型、API范式、提供商令牌计费规则和任务形式。不当的预处理策略可能增加部署成本或延迟,同时降低答案准确性。
Insight: 创新点在于首次系统性地将客户端预处理作为可控变量进行基准测试,揭示了预处理在云VLM VQA系统中的复杂影响。客观分析认为,该研究为VQA系统的实际部署提供了重要指导,明确了预处理何时有效、何时无效及其原因,有助于优化成本-质量权衡。
Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users’ questions about the physical world. Since modern VLMs remain difficult to run on mobile and edge devices, VQA systems increasingly offload inference to cloud-based VLMs. This gives mobile devices access to stronger computation, but it also makes visual input preparation a key system variable: how the image is prepared before offloading affects not only answer quality but also payload size, token cost, and system latency. Proprietary APIs expose little control over model internals or serving behavior, leaving client-side preprocessing as the main practical optimization space for downstream developers. Many such techniques have been proposed for visual offloading, yet their cost-quality impact on commercial cloud VLMs has never been studied. To fill this gap, we present VQABench, the first systematic benchmark that treats client-side input preprocessing as a controlled variable for cloud-VLM-based VQA. We evaluate 12 preprocessing techniques across three VQA datasets and four commercial VLMs from three providers, totaling 95,168 API calls. Our results show that preprocessing is not universally beneficial: its effectiveness depends on the target model, API paradigm, provider token-accounting rule, and task formulation. A poorly selected preprocessing strategy can increase deployment cost or latency while degrading answer accuracy. Overall, our benchmark clarifies when preprocessing helps, when it fails, and why, providing insights to guide future research and real-world deployment of VQA systems.
[76] LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering cs.CV | eess.IVPDF
Qian Yao, Jun-Jie Huang, Yongjun Wang, Luming Yang
TL;DR: LHSDet是一种基于视觉问答的高分辨率AI生成图像检测器,通过微调视觉语言模型并重新设计视觉编码器,有效捕捉AI生成图像中的低层纹理和高层伪影,结合语义级文本分支实现多模态特征融合,从而提升检测性能。
Details
Motivation: 现有AI生成图像检测方法通常对图像进行下采样,忽略了高分辨率图像中的关键低层纹理细节,且未知生成模型的不断涌现使得大规模预训练数据集难以获取,因此需要一种能充分利用多模态信息的新方法。
Result: 实验结果表明,LHSDet在包括扩散模型和自回归模型在内的多种生成模型上实现了高检测精度和鲁棒性能,在相关基准测试中表现出色。
Insight: 创新点在于将AI生成图像检测任务重新定义为视觉问答问题,并采用三分支架构(低层视觉分支、高层视觉分支和语义级文本分支)提取互补的多模态特征,特别是重新设计的视觉编码器能更好地捕获AI生成图像的特有伪影。
Abstract: Driven by advances in diffusion models and autoregressive models, the fidelity and resolution of AI-generated images now rival those of real images. However, existing AI-generated image detection methods often downsample the images, inevitably overlooking critical low-level texture details in high-resolution AI-generated images, therefore limiting their detection performance. In addition, the ceaseless emergence of unknown generative models makes large-scale pre-training datasets inaccessible. To address these challenges, we propose a novel high-resolution AI-generated image detector, termed LHSDet. Specifically, we formulate the AI-generated image detection task as a Visual Question Answering problem, leveraging a fine-tuned vision-language framework to fully exploit the complementary information between visual and textual modalities. Recognizing that the default visual encoder of existing vision-language models is not tailored for AI-generated image detection, we redesign a visual encoder to better capture both the low-level and high-level artifacts inherent in AI-generated images. Furthermore, we incorporate a semantic-level textual branch to enable multi-modal feature fusion and detection. Consequently, LHSDet employs a triple-branch architecture to extract complementary multi-modal features: a low-level visual branch that aggregates non-overlapping patches for local texture cues, a high-level visual branch based on SigLIP2 for global perception feature extraction, and a semantic-level textual branch that generates captions using BLIP-2. Extensive experimental results demonstrate that LHSDet achieves high detection accuracy and robust performance across diverse generative models, including both diffusion and autoregressive models.
[77] Vision-Language Grounding as Bidirectional Concept Correspondence cs.CV | cs.AI | cs.CLPDF
Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna
TL;DR: 本文提出了一种新的视觉-语言基础任务——双向概念对应,旨在同时识别图像中的实例级区域和文本中的视觉指称性片段,并建立它们之间的对应关系。该方法通过可学习的桥接令牌统一了文本分割、图像分割和跨模态对齐,并在多个数据集上实现了显著的性能提升。
Details
Motivation: 现有视觉-语言基础方法通常将任务简化为单向定位问题,即给定文本短语定位图像区域,忽略了确定文本中哪些部分具有视觉指称性及其与图像实体的对应关系这一基本挑战。本文旨在解决这一更基础的通信问题。
Result: 在长描述数据集上,ConCor-1模型将对应关系的F1分数提升了48%;在零样本LVIS基准测试中,F1分数提升了29%,显著优于基线方法。
Insight: 创新点在于将基础任务重新定义为双向概念对应问题,统一了短语基础、指称表达式基础和开放词汇检测等任务;通过桥接令牌机制将文本分割、图像分割和跨模态对齐整合为单一预测问题,提升了模型的泛化能力和性能。
Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as $\textit{bidirectional concept correspondence}$ over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce $\textbf{ConCor-1}$, a grounding model built on top of a pretrained vision-language model. It uses learnable $\textit{bridge tokens}$ to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that $\textbf{ConCor-1}$ consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.
[78] SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange cs.CV | cs.MM | cs.SDPDF
Jaemo Jeong, Junho Yoon, Hyunju Kim, Dongman Lee
TL;DR: 本文提出了SCoPE,一种无需训练的音视频事件感知框架,通过稀疏跨模态先验交换来解决现有方法中存在的错误共激活问题。该方法让所有查询标签竞争共享证据,并利用每个模态引导另一个模态的事件选择,从而在多个基准测试上显著提升了性能。
Details
Motivation: 动机在于解决无需训练的音视频事件感知方法中存在的错误共激活问题,即相关标签共享证据导致错误标签的得分可能不低于正确标签,从而影响预测准确性。
Result: 在LLP数据集上,使用相同的冻结CLIP+CLAP骨干网络,SCoPE相比已报道的AV$^2$A方法,将Type@seg提升了7.45分,Event@seg提升了5.04分。该固定配置无需修改即可迁移到OV-AVEBench和VGGSound-AVEL100k数据集。
Insight: 创新点在于提出了一个让所有查询标签竞争共享证据的机制,并通过跨模态引导来优化事件选择,这为解决多标签分类中的错误共激活问题提供了一种新颖的训练免费解决方案。
Abstract: Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual features with text-encoded event names. However, related labels share evidence. An incorrect label can then score at least as high as a correct one. We call this a false co-activation (FCA). No scalar cutoff can reject the incorrect label while keeping every correct one. Class-specific thresholds may prevent that label from becoming a final prediction, but the FCA remains in the underlying score vector. We introduce SCoPE, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other. We derive an exact condition for when this competition removes an FCA in a two-label fit. With identical frozen CLIP+CLAP backbones on LLP, SCoPE improves Type@seg by 7.45 points and Event@seg by 5.04 points compared with the reported AV$^2$A values. The same fixed configuration transfers unchanged to OV-AVEBench and VGGSound-AVEL100k.
[79] Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification cs.CVPDF
Yang Shu
TL;DR: 本文揭示了多模态LLM评审团中存在的源盲锚定攻击面,即引用不可信的同行判断会误导模型决策。研究发现,精心构造的错误引用比自然错误陈述更能推翻正确判决,并提出了面板共识验证方法作为防御机制,能有效阻断大部分伪造攻击并显著降低其危害。
Details
Motivation: 多模态LLM评审团依赖同行交叉验证,但引用的同行判断本身可能不可信,这构成了潜在的安全漏洞。论文旨在揭示这种文本层面的攻击面,并探索防御方案。
Result: 在两个数据集和七个VLM评审模型上,伪造的错误引用使正确判决被推翻的概率比自然错误陈述高1.5-2.7倍(95%置信区间排除对等性)。提出的面板共识验证方法能阻断84.9%的伪造攻击,将净危害降低97.5%。
Insight: 创新点在于发现了多模态评审系统中源盲锚定这一新型攻击面,并通过实验证明标签本身并非效应主因;同时提出的共识验证机制通过独立盲投交叉验证,为多模态协作评估提供了低成本防御方案。
Abstract: Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19–26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5–2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9% of fabricated attacks, cuts their net harm by 97.5%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.
[80] SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning cs.CVPDF
Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, Gaoang Wang
TL;DR: 本文提出了SportsGrounder框架,旨在解决密集体育视频推理中模型难以处理细粒度、小规模、高交互且视觉同质实体(如穿相同队服的球员、球)的问题。该框架利用开放词汇视觉专家辅助交错式定位,通过领域引导的对象提议和交错式定位融合机制实现精确的空间定位,并结合动作感知监督模块和混合偏好优化来提升模型对动作的准确表征和抗干扰能力。
Details
Motivation: 当前大型多模态模型在处理密集体育视频时,由于缺乏细粒度视觉细节,往往过度依赖文本先验来猜测答案,尤其是在区分视觉相似的动作和球员时存在困难。
Result: 在新构建的密集体育视频问答数据集(基于SoccerNet和FineSports)上的大量实验表明,SportsGrounder显著提升了细粒度推理能力,并达到了最先进的准确率。
Insight: 创新点包括:利用开放词汇视觉专家辅助交错式定位以增强空间感知;提出交错式定位融合机制整合显式边界框坐标与隐式视觉语义,保持严格时间对齐;设计动作感知监督模块直接正则化模型隐藏状态,迫使网络学习准确的动作表征而非依赖语言偏差;采用混合偏好优化以更好地区分欺骗性干扰项。
Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal contexts. Current Large Multimodal Models (LMMs) inherently struggle with such dense visual complexities. Due to the lack of fine-grained visual details, these models often over-rely on textual priors to guess answers, especially when distinguishing visually similar actions and players. To address this, we propose \textbf{SportsGrounder}, a framework that leverages an open-vocabulary visual expert to aid interleaved grounding specifically for dense sports video reasoning. To achieve precise spatial localization, we extract domain-guided object proposals and introduce an Interleaved Grounding Fusion (IGF) mechanism. The IGF frame-by-frame integrates explicit bounding box coordinates and implicit visual semantics with global grid features. This design preserves strict temporal alignment and prevents sequence length explosion. Furthermore, we design an Action-Aware Supervision (AAS) module that directly regularizes the model’s hidden states, forcing the network to learn accurate motion representations rather than relying on language bias. Optimized with Mixed Preference Optimization (MPO) to better distinguish deceptive distractors, our extensive experiments on newly curated dense sports VQA datasets (derived from SoccerNet and FineSports) demonstrate that SportsGrounder significantly improves fine-grained reasoning and achieves state-of-the-art accuracy.
[81] LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection cs.CVPDF
Shangye Song, Tianzhi Zhu, Syed Ariff Syed Hesham, Xin He, Yun Liu
TL;DR: 本文提出了一种用于伪装目标检测(COD)的新框架LAD-COD,该框架通过语言对齐的双重视觉融合(LADVF)将自上而下的语义目标指导与自下而上的分层视觉特征对齐,以解决伪装目标与背景视觉相似度高、边界模糊的难题。
Details
Motivation: 现有基于大型多模态模型(LMMs)的语言到掩码范式,其生成的目标嵌入主要指导掩码解码器,而负责保留低对比度边界和精细局部结构的密集视觉特征缺乏显式指导,这限制了伪装目标检测的性能。
Result: 在CAMO、COD10K和NC4K三个基准数据集上的实验表明,LAD-COD在所有12个数据集-指标对比中都取得了最佳报告值,达到了当前最先进(SOTA)水平。
Insight: 核心创新在于提出了语言对齐的密集感知框架,通过可训练的分层视觉分支捕获伪装敏感信息,并利用LADVF将目标嵌入从稀疏提示扩展到查询补丁级语言对齐特征,以门控残差方式与分层特征融合,实现了语义指导与精细结构保留的平衡。
Abstract: Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as top-down guidance for identifying which weak visual cues are relevant to the target. Recent segmentation systems built on large multimodal models (LMMs) demonstrate this possibility through instruction-conditioned target embeddings that guide mask decoding. However, in this language-to-mask paradigm, the generated target embedding conditions mainly the mask decoder, leaving the dense visual features that must preserve low-contrast boundaries and fine local structure without explicit guidance. We propose Language-Aligned Dense perception for COD (LAD-COD), a framework that aligns top-down semantic target guidance with bottom-up hierarchical visual features. Instead of fully adapting a large generic image encoder, LAD-COD learns a trainable hierarchical visual branch that captures camouflage-sensitive texture, boundary, and contextual information. To align these features with the target embedding, LAD-COD applies Language-Aligned Dual Visual Fusion (LADVF), which extends the embedding beyond sparse prompting to query patch-level language-aligned features and to gate their residual integration with the hierarchical features. This design allows semantic information to guide localization while preserving the fine structural details needed for camouflage segmentation. Experiments on CAMO, COD10K, and NC4K show that LAD-COD obtains the best reported value in all 12 dataset-metric comparisons.
[82] LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode Manufacturing cs.CVPDF
Wenbo Sui, Daniel Lichau, Harold Phelippeau, Zhao Liu
TL;DR: 本文提出了LIBAD,一个针对锂离子电池电极制造过程的多模态异常检测基准数据集,该数据集采集自真实的卷对卷生产线,包含对齐的双面可见光成像、高分辨率X射线成像以及适用于在线检测的低分辨率X射线成像。针对现有方法在该场景下表现不佳的问题,作者提出了DA-Core方法,这是一种基于记忆库的方法,通过在核心集选择中联合考虑正常特征的特征空间覆盖度和局部密度,以更紧凑的记忆库更好地保留细粒度的正常变化。
Details
Motivation: 现有的多模态工业异常检测研究主要集中在离散产品上,并使用强相关的RGB和3D观测数据,而连续过程制造和弱相关传感模态的研究不足。锂离子电池电极制造过程具有高度同质化的材料外观,且缺陷证据在不同模态间可能不一致,这为异常检测带来了挑战。
Result: 在LIBAD基准上,作者在适用于在线检测的可见光和低分辨率X射线设置下评估了代表性方法,发现它们存在有限的迁移性和持续较高的假阳性率。提出的DA-Core方法在核心集比例为0.05时,将FPR95从标准最远点采样的60.4%降低到54.3%。在该比例下,DA-Core也优于在0.20比例下获得的最佳标准核心集结果,同时将推理时间减少了43.9%。
Insight: 论文的主要创新点在于:1)引入了首个针对连续过程制造(锂离子电池电极制造)的多模态异常检测基准LIBAD,其特点是高度同质化外观和跨模态异常不一致性;2)提出了DA-Core方法,在构建记忆库时显式地联合优化特征空间覆盖度和局部密度,以更有效地捕捉正常数据的细粒度变化。这表明在设计过程制造的异常检测方法时,需要同时考虑正常特征的数据分布和模态间关系本身。
Abstract: Multimodal industrial anomaly detection has largely focused on discrete products using strongly correlated RGB and 3D observations, leaving continuous process manufacturing and weakly correlated sensing modalities underexplored. We introduce LIBAD, the first multimodal anomaly detection benchmark for Li-ion battery electrode manufacturing. Collected from real roll-to-roll production lines, LIBAD provides aligned double-sided visible-light imaging, high-resolution X-ray radiography, and inline-compatible low-resolution X-ray radiography. Electrode patches in LIBAD exhibit highly homogeneous material appearance, while defect evidence can be strong in one modality but weak or absent in another, resulting in pronounced cross-modal anomaly inconsistency. Benchmarks of representative methods under the inline-compatible visible-light and low-resolution X-ray setting exhibit limited transferability and consistently high false-positive rates. We therefore propose DA-Core, a memory-based method that jointly considers feature-space coverage and local density of normal features during coreset selection, allowing compact memory banks to better preserve fine-grained normal variations. With a coreset ratio of 0.05, DA-Core reduces FPR95 from 60.4% to 54.3% compared with standard farthest point sampling. At this ratio, DA-Core also outperforms the best standard coreset result (obtained at 0.20) while reducing inference time by 43.9%. These results suggest that both the data distribution of normal features and the modality relationship itself require explicit consideration when designing anomaly detection methods for process manufacturing.
[83] Distilling Physical Priors into Streaming World Models cs.CVPDF
Liangliang Zhao, Junying Wang, Danni Yang, Yifan Chang, Bin Fu
TL;DR: 该论文提出了PhyS框架,通过三个阶段将物理先验知识蒸馏到流式世界模型中,以提升模型在长时预测中保持物理一致性的能力。具体包括:构建包含12万真实物理交互视频的数据集PhyS-120K,通过物理感知的监督微调将物理先验注入双向DiT教师模型,再蒸馏到轻量级因果DiT中进行流式生成,最后使用在线强化学习和时间信用路由(TCR)进一步优化物理一致性。
Details
Motivation: 现有流式世界模型在长时预测中常违反基本物理约束,且传统基于双向DiT蒸馏的方法存在物理先验获取有限及蒸馏过程中进一步损失的问题。
Result: 在PhysicsIQ基准上,PhyS相比Wan2.1-14B教师模型提升了18.2%,相比Self Forcing、Rolling Forcing和Causal Forcing分别提升了23.7%、14.8%和31.4%;在VideoPhy、VideoPhy2和PhyGenBench等物理感知视频基准上也取得了改进。
Insight: 创新点包括:构建大规模结构化标注的物理交互视频数据集以增强物理先验学习;提出三阶段蒸馏框架结合监督微调与在线强化学习;引入时间信用路由(TCR)机制解决时序信用分配问题,通过重叠时间窗口评估物理一致性并分配优势信号。
Abstract: Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical constraints. A common approach distills pretrained bidirectional DiTs into few-step causal generators. However, this paradigm suffers from two fundamental limitations: generic bidirectional teachers acquire limited physical priors from visually oriented pretraining, and the limited priors suffer further loss during bidirectional-to-causal distillation. We present PhyS, a three-stage framework for distilling physical priors into streaming world models. To acquire physical priors from real-world interactions, we construct PhyS-120K, a dataset of 120K real-world physical-interaction videos spanning rigid-body dynamics, soft-body deformation, fluid phenomena, and phase transitions. Each video is annotated with structured descriptions of object properties and causal state transitions. Physics-aware supervised fine-tuning injects the physical priors into a bidirectional 14B DiT teacher, which we then distill into a lightweight 1.3B causal DiT for few-step autoregressive streaming generation. Finally, we use online reinforcement learning to incentivize the distilled model to generate physically plausible rollouts and further propose Temporal Credit Routing (TCR) to address temporal credit assignment. TCR evaluates physical consistency over overlapping temporal windows and routes the resulting group-relative advantages to temporally aligned denoising actions. On PhysicsIQ, PhyS improves the Wan2.1-14B teacher by 18.2% and the Self Forcing, Rolling Forcing, and Causal Forcing by 23.7%, 14.8%, and 31.4%, respectively. Results also improve the physics-aware video benchmarks VideoPhy, VideoPhy2, and PhyGenBench. The dataset, code, and more sample videos are available on our Project Page.
[84] AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining cs.CVPDF
Vasileios Tzouras, Paraskevas Pegios, Lazaros Nalpantidis
TL;DR: 本文提出了AgriField-40K,一个从17个公共资源中收集的、涵盖多种作物、杂草和田间条件的以农田为中心的数据集。基于此,作者提出了AgriMAE,一种参数高效的持续预训练基线方法,该方法通过仅训练轻量级适配器来适应在自然图像上预训练的掩码自编码器。实验表明,AgriMAE能持续提升下游任务性能,在使用少至9倍的可训练参数时,其表现可匹配甚至超越全量微调。
Details
Motivation: 解决基于田地的农业计算机视觉任务严重依赖昂贵标注和大规模预训练模型成本高昂的适应性问题,旨在为农业视觉任务提供一个实用的持续预训练资源和方法。
Result: AgriMAE在多个下游任务上进行了迁移评估,其性能持续提升,在使用少至9倍可训练参数的情况下,其表现可匹配甚至超越全量微调。
Insight: 创新点在于构建了大规模、多样化的农业领域数据集AgriField-40K,并提出了参数高效的持续预训练方法AgriMAE,该方法仅训练轻量级适配器,显著降低了适应成本。此外,探索了语义特征重建作为替代的预训练目标,这为领域自适应提供了新的思路。
Abstract: Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation of large pretrained models. We introduce AgriField-40K, a field-centric dataset curated from 17 public resources and covering diverse crops, weeds, and field conditions. Building on this, we present AgriMAE, a parameter-efficient continual pretraining baseline that adapts a masked autoencoder pretrained on natural images by training only lightweight adapters. We further explore semantic feature reconstruction as an alternative pretraining objective and evaluate transfer across multiple tasks. AgriMAE consistently improves downstream performance and can match or even outperform full fine-tuning while using up to $9\times$ fewer trainable parameters, showing that AgriField-40K is a practical resource for continual pretraining in agricultural vision. Project page: https://dtu-pas.github.io/agrifield40k/
[85] AdaDINO: Pair-Aware In-Backbone Adaptation of Frozen DINO for Efficient Remote Sensing Change Detection cs.CVPDF
Xu Zhang, Xinqing Li, Jianpeng Xie, Zeshuai Zhu, Xin He
TL;DR: 本文提出了AdaDINO框架,用于高效的遥感变化检测。该框架通过在冻结的DINO编码器主干中引入双时相交互,解决了现有基于视觉基础模型(VFM)的方法在独立编码图像后进行比较,导致主干网络无法感知跨时相关系的不足。核心组件包括变化感知门控局部适配(CGLA)、批共享块选择(BSCS)和CGLA先验引导细化(CPGR)解码器。
Details
Motivation: 动机在于解决DINO等视觉基础模型(VFMs)预训练用于单图像表示,而遥感变化检测需要对双时相对进行推理之间的不匹配问题。现有方法通常在独立编码两幅图像后才进行比较,使得VFM主干网络无法感知跨时相(cross-temporal)关系。
Result: 在四个遥感变化检测基准测试上的实验表明,AdaDINO相比基于VFM的基线方法取得了有竞争力或更优的性能,尤其在类别无关的SYSU-CD数据集上提升最大。在移除了62.5%的FFN隐藏宽度后,AdaDINO在SYSU-CD上仍能达到85.29%的F1分数,同时实现了1.41倍的吞吐量加速。
Insight: 主要创新点在于提出了“主干内成对适配”框架,通过CGLA模块在冻结块后耦合双时相流,并以相反符号注入共享的时相残差,从而增强真实变化响应并保留对中点。BSCS通过保留批共享的通道块子集来减少FFN计算,实现了高效的稀疏激活。CPGR解码器则复用编码器侧的变化响应进行由粗到细的预测,提升了效率与性能。
Abstract: Vision foundation models (VFMs) such as DINO are pretrained for single-image representation, whereas remote sensing change detection requires reasoning over a bi-temporal pair. Existing VFM-based methods usually encode the two images independently and compare them only afterward, leaving the VFM backbone unaware of cross-temporal relations. To bridge this mismatch, we present AdaDINO, a pair-aware in-backbone adaptation framework that equips a frozen DINO encoder with bi-temporal interaction for efficient change detection. Its core component, Change-aware Gated Local Adaptation (CGLA), couples the two streams after selected frozen blocks and injects a shared temporal residual into them with opposite signs, enhancing genuine change responses while preserving the pair midpoint. Batch-Shared Chunk Selection (BSCS) further reduces feed-forward network (FFN) computation by retaining a batch-shared subset of channel chunks that can be executed as a compact dense FFN. A CGLA-Prior-Guided Refinement (CPGR) decoder reuses encoder-side change responses for coarse-to-fine prediction. Experiments on four remote sensing change detection benchmarks show that AdaDINO achieves competitive or superior performance against VFM-based baselines, with the largest gain on the category-agnostic SYSU-CD dataset. With 62.5% of the FFN hidden width removed, AdaDINO still achieves an F1 score of 85.29% on SYSU-CD while delivering a 1.41$\times$ throughput speedup. The code will be released.
[86] Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence cs.CVPDF
Ling Lin, Yang Bai, Congcong Zhu, Jiangming Shi, Meng Wang
TL;DR: 本文提出了一种优势引导的门控框架,用于提升多模态大语言模型在视觉空间理解与推理任务中的开放端推理稳定性。该框架将逐步推理建模为有限时域决策过程,通过蒙特卡洛价值评估在推理树上提供中间监督信号,并利用步优势门和轨迹优势门动态筛选高价值推理步骤与高质量完整推理轨迹。
Details
Motivation: 针对多模态大语言模型在开放端推理过程中容易产生决策错误和错误累积,导致答案质量不稳定的问题,本文旨在通过动态干预和纠正推理偏差来提升推理的鲁棒性。
Result: 在构建的Reasoning-Tree-160k数据集上进行两阶段学习,大量实验表明该优势引导门控框架有效提升了基准MLLMs在视觉空间理解与推理任务上的性能。
Insight: 创新点在于将逐步推理形式化为决策过程并引入蒙特卡洛价值评估进行中间监督,同时通过多分支采样生成推理树来训练门控机制,并结合共享参数初始化与任务特定头部实现跨任务鲁棒性与多样性。
Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step-Advantage Gate and Trajectory-Advantage Gate, which dynamically select high-value reasoning steps and high-quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi-branch sampling, and combine shared-parameter initialization with task-specific heads to achieve cross-task robustness and diversity. During inference, the model greedily selects high-value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning-Tree-160k dataset and performed two-stage learning on it. Extensive experiments demonstrate that this advantage-guided gating framework effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks. The code is open to the public for research: https://github.com/LingLin-ll/Advantage-Guided-Gate.
[87] MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cs.CVPDF
Fulong Liu, Liang Xu, Chengqun Yang, Yuhao Zhang, Yichao Yan
TL;DR: 本文提出了MRBench,一个全面的人体运动-文本检索基准数据集,旨在解决现有基准中运动类型单一、分布不平衡以及文本描述过于简化的问题。MRBench包含来自动捕、真实视频、合成视频和运动生成模型的3,390个运动样本,涵盖118个细粒度类别,每个运动配有多粒度的描述文本。
Details
Motivation: 现有运动-文本检索基准存在运动同质化、分布不平衡以及文本描述过于简单重复的问题,这阻碍了对跨域和跨粒度对齐能力的可靠评估。
Result: 在MRBench上对代表性检索基线的广泛评估揭示了显著的跨数据集泛化差距和对查询粒度的敏感性。作者提出的轻量级粒度感知模型在不损害标准描述性能的前提下,提升了混合粒度检索的效果。
Insight: 创新点在于构建了一个包含异构运动、广泛平衡类别覆盖以及可靠、可区分、多粒度描述的综合基准。方法上,通过基于LLM的简洁和细粒度描述为额外的粒度特定运动提取器和文本适配器提供伪监督,并通过粒度感知的分数融合策略整合全局和适配的相似度,同时严格保持所有描述级别间的分数可比性。
Abstract: Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distributions, and oversimplified, repetitive texts, which hinder the reliable measurement of cross-domain and cross-granularity alignment. We thus introduce MRBench, a comprehensive motion-text retrieval benchmark featuring heterogeneous motions, broad and balanced category coverage, and reliable, discriminative, multi-granular descriptions. MRBench is constructed through a meticulously designed multi-stage data curation pipeline, which filters and balances candidates, verifies unambiguous semantic alignment, and generates motion-grounded descriptions at multiple granularities. The resulting benchmark contains 3,390 motions drawn from motion capture, in-the-wild videos, synthetic videos, and motion generative models, covering 118 fine-grained categories. Each motion is paired with concise, standard, and fine-grained descriptions, yielding 10,170 captions. Extensive evaluations of representative retrieval baselines on MRBench reveal a substantial cross-dataset generalization gap and pronounced sensitivity to query granularity. We propose a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model. LLM-based concise and fine-grained captions provide pseudo-supervision for extra-branch granularity-specific motion extractors and text adapters. For inference, granularity-aware score fusion integrates global and adapted similarities while strictly maintaining score comparability across all description levels. The resulting model improves mixed-granularity retrieval without compromising standard-caption performance. We believe that our MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.
[88] Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation cs.CV | cs.AIPDF
Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang
TL;DR: 本文提出了一种基于证据驱动的法医推理(EFR)框架,用于检测和定位多模态媒体篡改(DGM4)。该框架通过引入锚定-验证推理链,强制进行模态隔离感知和跨模态比较,并利用可验证的奖励系统和模态解耦优势路由机制,确保证据与结论的一致性。实验表明,EFR在实现最先进性能的同时,能生成结构化的法医推理记录,明确将解释与证据绑定。
Details
Motivation: 现有方法在检测多模态伪造内容时仅提供黑盒检测结果,缺乏决策依据,限制了其在法医实践中的可靠性。多模态大语言模型虽能提供可解释性,但应用于DGM4时面临解释与证据位置脱节、证据-结论一致性难以优化等困难。
Result: 实验显示,EFR在DGM4任务上实现了最先进的性能,并生成了结构化的法医推理记录,明确将解释与证据绑定。
Insight: 创新点包括锚定-验证推理链强制模态隔离感知和空间对应,可验证奖励系统优化证据-结论一致性,以及模态解耦优势路由机制缓解多任务训练中的信用分配问题,这些设计增强了检测的可解释性和可靠性。
Abstract: Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection results without any decision rationale, limiting their reliability in forensic practice. Multi-modal Large Language Models (MLLMs) offer a natural path toward explainability, but applying them to DGM4 raises two difficulties. First, models tend to generate explanations disconnected from predicted evidence locations, producing unverified attribution. Second, enforcing evidence-conclusion consistency requires active optimization, yet uniform training signals fail to distinguish localization tokens from classification tokens, making multi-head joint training unreliable. We propose a multi-modal manipulation detector based on an Evidence-Grounded Forensic Reasoning (EFR) framework. EFR introduces an Anchor-and-Verify reasoning chain that enforces modality-isolated perception before cross-modal comparison, with conclusion coordinates as explicit anchors to which downstream evidence must spatially correspond. A verifiable reward system then enforces evidence-conclusion consistency during training, while a Modality-Decoupled Advantage (MDA) routing mechanism mitigats credit misassignment across prediction tasks. Experiments show that EFR achieves state-of-the-art performance while producing structured forensic reasoning records that explicitly bind explanations to evidence.
[89] PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection cs.CVPDF
Kutub Uddin, Nusrat Tasnim, Khalid Malik
TL;DR: 本文提出PE-Mamba,一种用于AI生成图像检测的新框架。它基于预训练的PE-Core视觉Transformer,通过轻量级LoRA适配,并引入了双向选择性聚合器、软最大值加权聚合器和Sigmoid门控融合三个互补组件,以更好地聚合跨层特征。该方法在多个基准测试中取得了SOTA性能,且仅需训练极少参数。
Details
Motivation: 现有基于视觉Transformer的检测器通常采用加权求和策略聚合中间层表示,忽略了特征从浅层纹理线索到深层语义表示的内在有序语义递进,导致检测性能受限。
Result: 在UniversalFakeDetect数据集上达到96.6% mACC和99.5% mAP,在AIGCDetect数据集上达到95.3% mACC和98.1% mAP,超越了18个现有检测器,展现出优异的泛化能力,同时仅训练总参数的1.3%(LoRA部分仅0.13%)。
Insight: 创新点在于设计了双向选择性扫描聚合器来模拟特征层次的有序语义递进与上下文精炼,并结合全局聚合与自适应门控融合,实现了对跨层特征更精细、动态的建模,同时通过预训练主干与LoRA适配实现了高效参数微调。
Abstract: AI-generated image (AIGI) detection has become increasingly challenging due to the rapid advancement of generative models and the diminishing gap between synthetic and authentic content. Existing vision transformer-based detectors commonly rely on weighted-sum strategies to aggregate intermediate representations across transformer layers, often overlooking the inherently ordered semantic progression of hierarchical features from shallow texture cues to deep semantic representations. In this work, we propose \textbf{PE-Mamba}, a novel framework built upon a pre-trained PE-Core vision transformer with lightweight LoRA adaptation that introduces three complementary components for cross-layer feature aggregation and fusion. First, a bidirectional selective aggregator (BSA) processes layer-wise classification tokens through forward and backward selective scans, where the forward scan progressively accumulates shallow-to-deep forensic evidence, and the backward scan performs deep-to-shallow contextual refinement to reinterpret low-level cues in light of high-level semantic context. Second, a softmax-weighted aggregator (SWA) computes a learned global summary of all layer tokens as a complementary aggregation path. Third, a sigmoid-gated blend (SGA) adaptively fuses the BSA and SWA outputs via a learnable scalar gate, allowing the model to dynamically balance directional sequential evidence and global layer-wise aggregation. Extensive experiments on UniversalFakeDetect (96.6% mACC, 99.5% mAP) and AIGCDetect (95.3% mACC, 98.1% mAP) demonstrate that \methodname{} outperforms 18 detectors with superior generalization across diverse generative models, while training only 1.3% of total parameters (0.13% for LoRA alone).
[90] Evidence-RL: Towards Evidence-intensive Visual Reasoning cs.CV | cs.AIPDF
Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen, Cheng Yang
TL;DR: 本文提出Counterfactual Evidence Disentanglement (CED)方法,用于训练视觉语言模型(VLMs),使其回答基于具体的图像证据而非语言先验或捷径。该方法通过消除对象级证据区域并比较支持度下降来审计证据依赖,结合GRPO进行强化学习奖励。在九个基准测试和四个骨干网络上,该方法优于现有基于RL的后训练方法。
Details
Motivation: 现有感知感知的后训练方法通过全局扰动或注意力代理鼓励使用图像,但无法测试答案是否因果依赖于支持它的局部证据。
Result: 在九个公共基准测试和四个骨干网络上,CED优于先前的基于RL的后训练方法,针对性分析验证了其以对象为中心的信号有效性。
Insight: 创新点在于提出反事实证据解耦(CED)进行训练时证据审计,无需问题特定的证据标注或增加推理开销,使用弱对象级提议即可实现对象中心的证据依赖奖励。
Abstract: Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.
[91] EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking cs.CVPDF
Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang, Boyang Sun, Marc Pollefeys
TL;DR: 本文提出了EgoTrack3D,一个用于从第一人称(自我中心)RGB视频中重建和维持动态3D场景表示的模块化框架。该框架通过将2D分割掩码提升到全局3D坐标系,结合基于点的运动评分机制和基于体素的合并启发式方法来关联物体轨迹,旨在解决视角快速变化和部分遮挡带来的挑战。
Details
Motivation: 从第一人称视频理解3D场景对机器人和自主导航至关重要,但现有方法主要处理显式交互或假设静态场景,难以捕捉复杂动态。本文旨在解决在动态环境中进行持续3D物体跟踪的通用问题。
Result: 在Aria Digital Twin (ADT) 数据集上,EgoTrack3D在正确位置百分比(PCL)指标上相对于最强基线提升了11%,达到了最先进水平。即使在用稀疏3D边界框估计替代密集深度图的降级条件下,系统仍能保持准确的空间表示。
Insight: 主要创新点在于模块化框架设计,结合了2D到3D的提升、基于点的运动评分和基于体素的合并策略。客观来看,其将密集深度依赖解耦为稀疏估计,并集成交互引导的动态关联,增强了在噪声观测下的鲁棒性,为实际部署提供了实用方案。
Abstract: Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static scenes, limiting their ability to capture complex dynamics. We introduce EgoTrack3D, a modular framework that reconstructs and maintains a dynamic 3D scene representation directly from egocentric RGB video. The framework lifts 2D segmentation masks into a global 3D coordinate frame, using a point-based motion scoring mechanism alongside a voxel-based merging heuristic to associate object tracks. EgoTrack3D maintains accurate representations over time, achieving an 11% improvement in percentage of correct locations (PCL) relative to the strongest baseline on the Aria Digital Twin (ADT) dataset, while addressing the more general setting of persistent 3D tracking for both static and dynamic objects. Furthermore, to demonstrate the system’s robustness under degraded conditions that simulate real-world deployment constraints, we replace dense depth maps with sparse 3D bounding box estimation and integrate interaction-guided dynamic association, enabling EgoTrack3D to maintain accurate spatial representations despite noisy observations.
[92] ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models cs.CVPDF
Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh, Ramtin Pedarsani
TL;DR: 本文提出了一种名为ZOMP的零阶多模态提示调优方法,用于在仅能前向传播的受限场景下(如边缘设备或专有模型部署)高效微调CLIP等视觉语言模型。该方法通过跨模态低秩重参数化、梯度校正动量项和预算索引秩调度三项技术,在冻结的CLIP模型中同时优化视觉和文本分支的深度提示,实现了查询高效且无需反向传播的调优。
Details
Motivation: 解决在仅能前向传播的受限环境下(如内存受限的边缘设备或专有模型部署),传统基于反向传播的微调方法不可行,而现有零阶提示调优方法要么仅调优单模态提示,要么搜索空间过大导致收敛需要数千次前向传播,查询效率低下的问题。
Result: 在匹配的5,000次查询预算下,在13个视觉语言基准测试中,ZOMP在少样本准确率和查询效率方面均持续优于先前的无需反向传播的提示调优方法,并在基类到新类、跨数据集迁移和分布外设置中表现出更好的泛化能力。
Insight: 创新点在于联合利用多模态性和低秩结构,通过跨模态低秩重参数化(将两个分支通过共享因子绑定以保持低搜索维度)、梯度校正动量项(稳定噪声零阶估计)和预算索引秩调度(随查询预算增加解锁容量)三项技术,实现了实用且查询高效的无需反向传播的提示调优路径。
Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and proprietary model deployments. Prior BP-free, zeroth-order prompt-tuning methods avoid this requirement but often tune prompts in a single modality or optimize over a search space large enough that convergence requires thousands of forward passes, which is impractical under realistic query budgets. We propose ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation. ZOMP combines three ingredients: a cross-modal low-rank reparameterization that ties the two branches through a shared factor and keeps the effective search dimensionality small, a gradient-correction momentum term that stabilizes the noisy zeroth-order estimate, and a budget-indexed rank schedule that unlocks capacity as the query budget is spent. Across 13 vision-language benchmarks under a matched 5,000-query budget, ZOMP consistently outperforms prior BP-free prompt-tuning methods in both few-shot accuracy and query efficiency, and it generalizes better across base-to-new, cross-dataset transfer, and out-of-distribution settings. Our results show that jointly exploiting multimodality and low-rank structure is an effective route to practical, query-efficient BP-free prompt tuning.
[93] When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery cs.CV | cs.GRPDF
Y Huynh, Duc Thanh Nguyen, Thao Minh Le, Mohamed Abdelrazek
TL;DR: 本文提出ASV3D框架,旨在利用一张额外的辅助图像来改进单视图3D物体重建。该框架包含两种适应策略:无需重新训练的零样本适应方案,以及通过对比学习进一步优化视觉保真度和跨视图一致性的优化适应方案。
Details
Motivation: 单视图3D重建面临的关键挑战是缺乏来自其他视角的关键信息,导致难以完成完整的3D结构。现有方法缺乏有效机制将额外视图整合到单视图重建流程中。
Result: 在基准数据集和真实世界数据集上,ASV3D应用于两种最先进的单视图3D重建流程,均能一致地提升重建精度和鲁棒性,在定量指标和人类偏好评估中均优于基线方法。
Insight: 创新点在于提出了一个灵活的适应框架,能够利用测试时的额外图像来增强单视图重建,而无需重新训练基础模型。特别是通过对比学习优化跨视图一致性,为解决单视图信息不足问题提供了新思路。
Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to complete 3D structures. Using an additional view may help to resolve the issue. However, there is no mechanism that can integrate the extra view into the single-view 3D reconstruction principle. We address this challenge by proposing ASV3D, a framework for adapting single-view 3D object reconstruction to test-time data with support from one additional image. We introduce two adaptation strategies: (i) a zero-shot adaptation scheme that leverages the auxiliary image to improve the reconstruction quality of an object without retraining, and (ii) an optimised adaptation scheme that further enhances visual fidelity and cross-view consistency via contrastive learning. We apply our ASV3D to improve two state-of-the-art single-view 3D reconstruction pipelines on both benchmark and real-world datasets. Results demonstrate that our approach consistently improves reconstruction accuracy and robustness under unconstrained multi-view inputs, outperforming the baselines in both quantitative metrics and human preference. We publish our code and the real-world object dataset in our project page at https://github.com/YNhuHuynh/ASV3D/tree/main.
[94] Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching cs.CV | cs.AIPDF
Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
TL;DR: 该论文提出了一种基于全体积多任务潜在流匹配的跨模态医学图像合成方法。该方法利用大规模预训练的3D变分自编码器提供紧凑的体素外观潜在表示,将图像翻译问题转化为条件流匹配问题,从而实现了对完整医学影像体积的处理,并仅需一个模型即可完成多种跨模态(如MRI→CT)和模态内(如MRI→MRI)的翻译任务。
Details
Motivation: 解决现有跨模态医学图像合成方法的两大局限:一是大多基于2D切片或3D小块而非完整体积进行处理,二是每个翻译任务都需要单独训练一个模型。其根本原因在于缺乏足够强大的体素先验,导致生成模型必须在有限配对数据下同时学习解剖结构外观和跨模态映射这一不适定问题。
Result: 在三个多中心数据集上的实验表明,该方法在所有任务上,全体积处理均优于基于小块的方法,且单个多任务模型的性能与针对每个任务单独训练的基准模型相当。更重要的是,联合训练解锁了任务特定方法无法实现的能力:对训练中未见过的解剖区域实现零样本泛化(SSIM与全监督模型相差仅0.15以内),以及实现从未被直接监督的、跨数据集的组合式翻译。
Insight: 核心创新在于将目标解耦:利用预训练的3D VAE提供强大的体素解剖先验,将翻译任务简化为潜在空间的条件流匹配,这使得处理完整体积变得可行。同时,采用多任务联合训练,用一个模型取代N个专用网络,不仅保持了性能,还显著提升了模型的泛化能力和组合推理能力,为构建可扩展、能超越训练分布泛化的合成系统提供了新路径。
Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task. Both stem from a single cause, the absence of a sufficiently strong volumetric prior, which forces generative models to learn anatomical appearance and cross-modality mapping simultaneously, an ill-posed problem at the scale of available paired datasets. We propose to decouple these objectives. A large-scale pretrained 3D variational autoencoder provides a compact latent representation of volumetric appearance, reducing translation to a conditional flow-matching problem. This compression makes whole-volume processing tractable, while a resolution-aware sampling strategy preserves native anatomical scale. We train a single model jointly across inter-modality (MRI$\to$CT, CBCT$\to$CT) and intra-modality (MRI$\to$MRI) tasks over three multi-center datasets. Across all tasks, whole-volume processing outperforms its patch-based counterpart, and the multi-task model matches task-specific baselines while replacing $N$ networks with one. Crucially, joint training unlocks capabilities inaccessible to task-specific approaches: zero-shot generalization to anatomical regions unseen during training, within 0.15 SSIM of the fully supervised model, and compositional cross-dataset translation along paths never directly supervised. These results suggest that combining a strong volumetric prior with multitask training is a scalable route toward synthesis systems that generalize beyond their training distribution. Code is available at https://github.com/arco-group/Whole-Volume-Latent-FM.
[95] Wiener Representation Filtering for VLM Hallucination Suppression cs.CV | cs.LGPDF
Ameen Ali, Tamim Zoabi, Lidor Brami, Lior Wolf
TL;DR: 本文提出了一种无需训练、后处理的表示编辑技术,用于抑制视觉语言模型(VLM)中的物体幻觉现象。该方法通过对语言骨干网络的表示空间进行轻量级离线校准,估计协方差结构,并推导出维纳型估计器,在推理时直接修正选定深层的前馈输出投影,从而减少幻觉。
Details
Motivation: 视觉语言模型在开放式描述和视觉问答中表现出色,但经常描述图像中不存在的物体、属性或关系,即物体幻觉问题。本文旨在提出一种无需额外训练或微调的后处理方法来抑制这种幻觉。
Result: 在LLaVA-1.5、MiniGPT-4、Gemma3和mPLUG-Owl2等模型上的实验表明,该方法在CHAIR、POPE和MME基准上持续减少了物体幻觉,同时保持了描述的流畅性和整体响应质量。在TempCompass视频理解基准和基于离散扩散的语言模型上的进一步实验也证明了其泛化能力。
Insight: 创新点在于将隐藏状态建模为真实成分和幻觉相关成分的叠加,并推导出闭式最优增益的维纳型估计器,通过特征分解实现满足稳定性准则的模式级衰减。该方法无需梯度更新或微调,仅需一次离线校准,推理时模型运行速度不变,是一种高效的后处理表示滤波技术。
Abstract: Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object hallucination. We propose a {training-free, post-hoc representation editing technique} that operates in the representation space of the language backbone. The method performs a lightweight, one-time offline calibration on a modest paired dataset to estimate the required covariance structures, using only forward passes and empirical second-order statistics with no gradient updates or fine-tuning, after which the correction is absorbed directly into the model’s existing weights. By modeling hidden states as a superposition of truthful and hallucination-associated components, we derive a Wiener-type estimator whose optimal gains are given in closed form from the covariances of paired truthful and hallucinated representations. An eigendecomposition yields mode-wise attenuation that respects a stability criterion, i.e., the filter responds continuously to estimation noise. The correction is applied once to the feed-forward output projections of selected deeper layers, at inference time, the model runs unchanged and at the same speed. Experiments on LLaVA-1.5, MiniGPT-4, Gemma3, and mPLUG-Owl2 demonstrate consistent reductions in object hallucination on CHAIR, POPE, and MME while maintaining caption fluency and overall response quality. We further demonstrate the generality of our approach on the TempCompass video understanding benchmark and on discrete diffusion language models for grounded dialogue, showing that representation filtering reduces hallucinations even in temporal video reasoning and multi-step, sequence-wide denoising settings.
[96] VTO: Visual Tool Orchestration for Video Anomaly Detection cs.CV | cs.AIPDF
Rui Wang, Yeteng Wu, Xianling Zhang, Mengshi Qi
TL;DR: 本文提出VTO,一个用于视频异常检测的视觉工具编排框架。该框架采用过程监督的强化学习方法,通过引入基础模型驱动的认知评估器提供细粒度语义反馈,以优化多步推理策略。
Details
Motivation: 传统深度学习方法在视频异常检测中泛化能力差,而现有基于监督微调的多模态智能体系统难以处理复杂工具编排,标准强化学习则因粗粒度奖励导致过早终止。
Result: 在精心构建的VAD-Tool基准测试上,VTO显著优于基线方法,在工具调度任务中实现了高达10.2%的绝对准确率提升。
Insight: 核心创新在于提出了过程监督认知对齐机制,通过基础模型提供上下文感知的语义反馈,显式惩罚逻辑截断并奖励完整因果链,从而引导智能体优化多步推理策略和工具动态编排。
Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse scenarios. While multimodal agents offer a promising tool-learning paradigm for VAD, current systems relying on supervised fine-tuning struggle with complex orchestration, and standard reinforcement learning often causes premature termination due to coarse-grained outcome rewards. To address these challenges, we propose VTO, a process-supervised reinforcement learning framework. Moving beyond static tool usage, VTO enables the agent to dynamically explore and interact with the environment. Specifically, we introduce a foundation model-driven cognitive evaluator to provide context-aware semantic feedback, which is seamlessly integrated into a Process-Supervised Cognitive Alignment that delivers fine-grained, step-wise supervision. By explicitly penalizing logical truncation and rewarding complete causal chains, the agent optimizes its multi-step reasoning policy for interrelated tool orchestration. To support our proposed framework, we meticulously crafted VAD-Tool, a hierarchical visual tool set comprising 12 specialized vision tools spanning from entity tracking to high-stakes hazard detection, and established the corresponding benchmark for rigorous multi-step reasoning evaluation. Extensive experiments on VAD-Tool demonstrate that VTO significantly outperforms baselines, achieving up to a 10.2% absolute accuracy improvement in tool scheduling. Code and data are available at https://github.com/MICLAB-BUPT/VTO.
[97] Ego-OSCAR: Egocentric Open source Stereo CAptuRe System cs.CV | cs.AR | cs.ROPDF
Gunjan Paul, Senthil Palanisamy, Satpal Singh Rathore, Pratyush Kumar Patnaik, Shubhanshu Khatana
TL;DR: 本文介绍了Ego-OSCAR,这是一个开源的、低成本的、头戴式立体惯性捕捉设备,用于在自然环境中收集自我中心数据。该设备集成了硬件同步的全局快门立体相机、六轴IMU、嵌入式Linux单板计算机用于设备端视频编码,以及实时微控制器用于用户反馈和看门狗功能。同时,作者还发布了完整的软件栈和约550小时/相机的带同步IMU的自我中心立体视频数据集,并提供了自由形式的动作标注和每帧3D手部重建。
Details
Motivation: 旨在提供一个低成本、可众包的自我中心数据采集平台,降低大规模收集自我中心数据的门槛,而不是追求与高端研究级系统(如Project Aria)相媲美的单机保真度。
Result: 每个设备的材料成本低于200美元,使用了商业可用组件和3D打印部件;发布了包含约550小时/相机带标注的立体视频数据集,覆盖日常室内环境,并提供了动作标注和3D手部重建。
Insight: 创新点在于设计了一个开源的硬件-软件-数据集完整生态系统,强调低成本、可扩展性和易用性,通过分布式贡献者网络收集数据,并提供了丰富的标注(如自由形式动作描述和3D手部重建),为自我中心视觉研究提供了可访问的基础设施。
Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced
[98] Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering cs.CV | cs.AIPDF
Yusra Tariq, Rakesh Chandra Joshi
TL;DR: 本文提出了一种用于医学视觉问答(VQA)的频域双分支融合模块。该方法通过在频域中根据输入问题自适应地选择全局低频结构和细粒度高频细节,以更好地利用视觉和文本表征中的互补频率信息,从而提升医学VQA的性能。
Details
Motivation: 现有在空间域操作的多模态融合方法可能无法充分利用视觉和文本表征中存在的互补频率信息,而医学VQA需要对病变纹理、边界锐度等细微视觉证据与临床语言进行精确对齐。
Result: 在PMC-VQA数据集上进行预训练,并在VQA-RAD和SLAKE基准上进行微调,结果表明,具有频率感知的多模态融合提升了医学VQA的性能,同时保持了轻量高效的架构。
Insight: 核心创新点是提出频域双分支融合模块,将问题作为条件进行频谱滤波,自适应融合不同频率的视觉信息。此外,从冻结的BiomedCLIP编码器中提取早期纹理敏感层和最终语义层的互补特征,并通过对称InfoNCE目标与问题表征对齐,提供了更丰富的频谱进行滤波。
Abstract: Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language. Existing multimodal fusion approaches operating in the spatial domain may not fully exploit complementary frequency information present in visual and textual representations. We introduce a dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation. To provide a richer spectrum for filtering, we extract complementary features from early texture-sensitive and final semantic layers of a frozen BiomedCLIP encoder and align both with the question representation using a symmetric InfoNCE objective prior to staged joint training with a BioBART decoder. We pretrain the proposed model on PMC-VQA and fine-tune it on the VQA-RAD and SLAKE benchmarks, demonstrating that frequency-aware multimodal fusion improves medical VQA performance while maintaining a lightweight and efficient architecture.
[99] Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation cs.CVPDF
Haozhe Wang, Jintao Cheng, Weibin Li, Xiaoyu Tang
TL;DR: 本文提出了一种无需训练、即插即用的测试时原型适应(TPA)方法,用于提升开放词汇语义分割(OVSS)的性能。TPA在输出层面操作,利用少量未标注的部署域图像,通过轻量级的转导适应阶段,从宿主模型的预测中提取高置信度的锚点补丁,并聚合其冻结的DINO特征以构建每个类别的原型库。在推理时,通过计算与原型库的余弦相似度,生成辅助分数并与宿主模型的逻辑值线性融合,从而在不修改宿主模型前向传播或权重的情况下,显著提高分割精度。
Details
Motivation: 现有OVSS方法通常通过重新设计CLIP的内部注意力或注入辅助视觉基础模型的特征来改善其空间行为,但这些方法需要访问宿主模型的内部计算,且针对特定前向传播定制,缺乏通用性和灵活性。本文旨在开发一种无需训练、在输出层面操作的通用插件,以兼容多种宿主模型,并利用未标注的部署域数据自适应提升性能。
Result: TPA在五个代表性的OVSS宿主模型(涵盖注意力重设计和VFM注入两种设计)、三个CLIP骨干网络、八个基准测试集以及多种内部VFM选择上进行了评估。在单一超参数设置且无需针对每个宿主模型进行调整或参数更新的情况下,TPA一致性地提高了分割准确率,并且在大多数基准测试中,仅需约10%的未标注部署域图像即可有效构建原型库。
Insight: 论文的创新点在于提出了一种训练无关、输出层面的通用适配方法,通过转导学习利用未标注数据构建类别原型,增强了模型在部署域的泛化能力。从客观角度看,该方法避免了模型内部修改,具有较好的兼容性和可扩展性,且通过轻量级适应过程实现了显著的性能提升,为开放词汇分割的测试时适应提供了新思路。
Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP’s spatial behavior either by redesigning its internal attention or by injecting features from auxiliary vision foundation models; both require access to the host’s internal computation and are tailored to its specific forward pass. In this work, we propose Test-time Prototype Adaptation (TPA), a training-free plug-in that operates at the output level, leaving the host’s forward pass and weights unmodified. By leveraging a lightweight transductive adaptation phase, TPA identifies confident anchor patches from the host’s own output predictions on a small pool of unlabeled deployment-domain images, and aggregates their frozen DINO features into per-class prototypes; at inference, a single cosine similarity lookup against this frozen bank provides an auxiliary score fused linearly with the host’s logits. TPA composes with five representative OVSS hosts spanning attention-redesign and VFM-injection designs, across three CLIP backbones, eight benchmarks, and multiple internal VFM choices. Under a single set of hyper-parameters and without per-host tuning or parameter updates, TPA consistently improves segmentation accuracy, with as few as approximately 10% of unlabeled deployment-domain images sufficing for effective bank construction on most benchmarks.
[100] Open-World Semantic Segmentation with Sensitivity Modeling cs.CV | cs.AIPDF
Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
TL;DR: 本文提出了一种用于开放世界语义分割的三解码器统一编码器-解码器架构。该方法通过三个互补的解码器分别处理已知类别分割、未知区域检测以及捕捉语义不确定性的细粒度纹理异常和激活不稳定性,从而在分割已知类别的同时,无需额外监督地检测和分组新颖或异常内容。
Details
Motivation: 现代视觉系统需在’开放世界’中运行,而传统语义分割模型基于’封闭世界’假设,会对新内容产生过度自信的错误分类。本文旨在解决开放世界语义分割这一联合任务。
Result: 在Cityscapes和BDD-Anomaly数据集上的实验表明,该方法在保持竞争力的封闭集精度的同时,提升了异常分割和新类别发现能力。在BDD-Anomaly上,AUROC提升了2.4%,FPR@95TPR降低了2.5个百分点,优于基线模型。
Insight: 核心创新是引入了第三个’敏感性解码器’,专门捕捉语义原型和对比范数无法可靠检测的、指示语义不确定性的细粒度纹理不规则性和激活不稳定性。三个解码器分别从logit空间、嵌入空间和编码器尺度提供了真正互补的异常信号。
Abstract: Modern vision systems must operate in “open-world” settings, where models must recognize known categories and detect unseen or anomalous content. Conventional semantic segmentation models operate under a “closed-world” assumption, often producing overconfident misclassifications on novel content. We address open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design. The first decoder performs closed-set segmentation using Gaussian prototypes for known categories. The second uses contrastive feature learning to isolate unknown regions in embedding space. The third, our key contribution, is a sensitivity decoder that captures fine-grained texture irregularities and activation instabilities indicative of semantic uncertainty, which neither semantic prototypes nor contrastive norms can reliably detect. The three decoders provide genuinely complementary signals: class-level OOD distance in logit space, global feature energy in embedding space, and local activation instability across encoder scales. Experiments on Cityscapes and BDD-Anomaly show that our method improves anomaly segmentation and novel-class discovery while maintaining competitive closed-set accuracy, with gains of +2.4% AUROC and a 2.5 pp. reduction in FPR@95TPR on BDD-Anomaly over the baseline.
[101] Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No cs.CV | cs.MMPDF
Ji Huang, Barry Devereux, Hui Wang
TL;DR: 本文提出了一种无需训练的时间定位方法FV-Action,通过将时间戳回归任务转化为多粒度二元问题扫描,显著提升了多模态大语言模型在视频时间定位任务上的性能。该方法在多个基准测试中超越了现有训练方法和零样本模型,实现了最先进的训练无关结果。
Details
Motivation: 当前强大的视觉语言模型在事件识别上可靠,但在预测时间戳时表现不佳,常出现高置信度的错误预测。作者认为问题在于任务接口设计,而非感知能力本身,因此探索了无需修改模型权重的改进方案。
Result: 在Charades-STA基准上达到56.8%的R@0.5,超越了同骨干网络的本地定位流程和该基准上所有训练无关方法;在TACoS上零样本评估超过所有经过时间定位训练的模型,并在ActivityNet Captions和QVHighlights上优于直接预测方法。
Insight: 创新点在于将时间定位重构为基于首词概率排名的层次化二元问题扫描,避免了回归任务中的置信度误导。分析揭示了错误可分解为感知轴和几何轴,后者可通过输出窗口与事件宽度比解析预测,为模型诊断提供了新视角。
Abstract: Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as $3.8%$ R@0.5 on Charades-STA, and $77$ to $80%$ of their wrong predictions carry low output entropy: the models are confidently wrong, and entropy-based error detection stays below a random classifier. We show that this failure lives in the task interface, not in perception. Holding the weights fixed, replacing timestamp regression with a coarse-to-fine scan of binary questions, whose first-token probabilities are consumed only as a ranking, raises R@0.5 by $28$ to $50$ points across four frozen backbones. The residual failures decompose into two measurable axes: a perception axis that moves with the backbone, and a geometry axis that is analytically predictable from the ratio of the output-window and event widths. FV-Action, the training-free method built on this analysis, reaches $56.8%$ R@0.5 on Charades-STA, above the same backbone’s native grounding pipeline and the strongest training-free result on this benchmark; it surpasses every TVG-trained model evaluated zero-shot on TACoS, and improves over direct prediction on ActivityNet Captions and QVHighlights, with no temporal supervision at any stage.
[102] Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation cs.CVPDF
Uri Z. Kialy, Gil Ben-Artzi
TL;DR: 本文提出了电路微调(CFT)框架,旨在实现计算高效的视觉Transformer(ViT)适应。该方法利用电路发现技术,在训练前选择关键模块进行微调,从而显著减少训练计算量(FLOPs)和训练时间,同时不增加额外参数或推理开销。
Details
Motivation: 现有参数高效微调(PEFT)方法虽能减少参数量,但计算效率仍不足,训练步骤多且成本高。本文旨在解决PEFT中计算效率低下的问题,通过预选模块来加速微调过程。
Result: 在VTAB-1k、Swin、CBIS-DDSM和Gemma-3 on CUB-200等多个基准测试中,CFT平均仅需约20个epoch达到峰值准确率,相比基线方法(44-96个epoch)减少了2.3-6.6倍的训练FLOPs和最高16倍的挂钟时间,且不增加参数。
Insight: 创新点在于将电路发现应用于训练前模块选择,并通过近零初始化探针头来隔离主干网络对目标分布的响应,避免了传统归因方法对特定分类器的依赖,实现了无需学习率预热的高效微调。
Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PEFT, it does not imply \textit{compute efficiency}: parameter-sparse methods can still incur full-model training cost per step, and typically need long schedules to reach peak accuracy. We introduce Circuit Fine-Tuning (CFT), a compute-efficient framework that uses circuit discovery—conventionally used to explain trained models—to select modules for fine-tuning before training. Whereas attribution is conventionally formulated against a trained task head, we formulate it against a near-zero-initialized probe head, which isolates the response of the backbone to the target distribution rather than the preferences of a particular classifier. CFT then fine-tunes only the recovered subgraph. CFT needs no learning-rate warmup and reaches peak accuracy in ${\sim}20$ epochs on average—versus $44$–$96$ for strong PEFT baselines—yielding $2.3$–$6.6\times$ fewer training FLOPs and up to $16\times$ less wall-clock time, while adding zero parameters and no inference operations. Experiments across a standard visual transfer benchmark (VTAB-1k), hierarchical backbones (Swin), domain-shifted medical imaging (CBIS-DDSM), and a vision-language model (Gemma-3 on CUB-200) demonstrate the effectiveness of CFT. Code is available at https://github.com/UriKialy/CFT
[103] PARAGraph: Pathology-Anatomy-Aware Hierarchical Graph for Diabetic Retinopathy Grading cs.CVPDF
Ziyang Zhang, Yuankai Huo, Yalin Zheng, He Zhao
TL;DR: 本文提出了一种名为PARAGraph的病理-解剖感知分层图框架,用于糖尿病视网膜病变(DR)的严重程度分级。该框架将眼底图像表示为包含病灶、区域和解剖语义节点的三层层次图,通过引入基于视盘-黄斑的归一化坐标系和双融合策略,增强了模型的临床可解释性和对病灶分割噪声的鲁棒性。
Details
Motivation: 现有深度学习方法通常将DR分级视为图像级分类任务,未能显式建模病灶类型、空间关系等临床证据,限制了其可解释性和可靠性。
Result: 在Messidor-2、APTOS和DDR三个基准数据集上的实验表明,PARAGraph取得了优于现有最先进方法(SOTA)的DR分级性能。
Insight: 创新点在于构建了结合医学先验(如视盘-黄斑坐标系)的分层图表示,以及通过双融合策略(全局视觉上下文与决策级分支)来缓解病灶分割噪声,从而实现了临床可解释且鲁棒的预测。
Abstract: Diabetic retinopathy (DR) remains a leading cause of vision loss among working-age adults worldwide, making reliable severity grading clinically important. Despite strong performance, most deep models formulate DR grading as image-level classification and do not explicitly model clinically grounded evidence, such as lesion types and spatial relations. In this paper, we propose PARAGraph, a Pathology-Anatomy-Aware Hierarchical Graph framework for DR grading. PARAGraph represents each image as a three-level hierarchical graph with lesion-level nodes, intermediate category and region nodes, and global anatomical and semantic nodes. To incorporate medical priors into nodes, we construct an optic disc-fovea-anchored coordinate frame that provides a scale- and rotation-normalized retinal reference system. Within this frame, lesion nodes are encoded with category, normalized area, and anatomical coordinates. To mitigate noisy lesion segmentation, PARAGraph uses a dual-fusion strategy that introduces global visual context into a graph semantic node and a decision-level prediction branch, improving robustness when lesion evidence is unreliable. Extensive experiments on Messidor-2, APTOS, and DDR show that PARAGraph achieves consistent DR grading performance over state-of-the-art methods. Interpretability and robustness analyses further demonstrate that its predictions are clinically grounded, closely associated with lesion evidence and robust to lesion segmentation noise.
[104] VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression cs.CV | q-bio.GNPDF
Xin Luo, Yicheng Tao, Haoxuan Zeng, Suyuan Wang, Chenzi Ouyang
TL;DR: 论文提出了VOICE,一种多模态基础模型,旨在从H&E病理图像中预测单细胞基因表达。该模型通过对比学习对齐病理图像和转录组嵌入,并采用直接回归与检索式预测的双分支架构,根据基因的形态可预测性进行融合,从而在未见过的患者、切片和部分重叠基因面板上实现泛化。
Details
Motivation: 空间转录组学成本高、基因覆盖有限且样本量小,而H&E成像廉价且大规模可用,因此从形态学直接预测单细胞表达成为将分子分析应用于大规模组织档案的实用方法。
Result: 在Xenium数据上,VOICE在七个评估指标上一致优于先前的单细胞表达预测方法,展示了在未见患者、切片和部分重叠基因面板上的泛化能力。
Insight: 创新点包括结合病理与转录组基础模型的对齐训练、双分支(直接回归与检索)预测架构,以及按基因可预测性动态融合分支的机制,这有助于处理缺乏形态信号的基因。
Abstract: Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.
[105] Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information cs.CVPDF
Xianghan Meng, Wei He, Zhiyuan Huang, Chun-Guang Li
TL;DR: 本文提出了一种名为DeepMORSE的新方法,用于利用文本信息进行图像聚类。该方法通过一个模态共享的自表达模型来发现跨模态结构,并同时学习符合模态特定子空间并集的表示,避免了直接强制跨模态对齐可能带来的问题。
Details
Motivation: 现有方法通常为每张图像检索对应的文本,并通过直接强制跨模态一致性(如最大化预训练视觉语言模型中的图像-文本相似度)来优化多模态表示。这种策略在对齐异构表示时,没有显式建模每个模态的内部结构,可能导致不可靠的对齐或扭曲对聚类至关重要的模态特定结构。
Result: 在六个广泛使用的图像聚类基准测试(包括UCF-101、DTD-47和ImageNet-Dogs)上进行了评估,性能提升超过3%。此外,在下游任务(如图像检索和零样本分类)上实现了最先进的性能,无需任何任务特定损失或后处理,证明了所学表示具有很强的可迁移性。
Insight: 核心创新点是提出了模态共享的自表达模型来发现跨模态结构,并同时学习符合模态特定子空间并集的表示。理论分析表明,该模型的自表达系数能抑制类间噪声,促进子空间保持解,且小批量优化过程引入了对自表达模型的隐式正则化。
Abstract: Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Vision-Language Models (VLMs). Existing approaches typically retrieve a textual counterpart for each image and then refine multimodal representations by directly enforcing cross-modal agreement, e.g., maximizing image-text similarity inherited from pretrained VLMs. However, such a strategy aligns heterogeneous representations across modalities without explicitly modeling the intrinsic structure within each modality and thus might yield unreliable alignment or distort modality-specific structures that are crucial for clustering. In this paper, we propose a simple but principled approach, termed deep modality-shared self-expressive model (DeepMORSE), which discovers cross-modal structures via a modality-shared self-expressive model and simultaneously learns structured representations that conform to a union of modality-specific subspaces. Moreover, we theoretically justify that the modality-shared self-expressive coefficients suppress inter-class noise towards a subspace-preserving solution, and show that mini-batch optimization procedure introduces an implicit regularization onto the self-expressive model. We evaluate our DeepMORSE on six widely used image clustering benchmarks and observe performance improvements exceeding 3% on the UCF-101, DTD-47, and ImageNet-Dogs datasets. In addition, we demonstrate the strong transferability of the learned representations by achieving state-of-the-art performance on downstream tasks such as image retrieval and zero-shot classification—without requiring any task-specific losses or post-processing. The code is available at: https://github.com/mengxianghan123/DeepMORSE.
[106] Agentic AI-powered flexible fiber-bundle endoscopy for high-resolution NIR-II fluorescence imaging in vivo cs.CVPDF
Yanzhao Shi, Yuanhua Liu, Sixin Xu, Wayne Jason Li, Yuyuan Chen
TL;DR: 本文提出了一种AI驱动的柔性光纤束内窥镜平台,通过光学-计算协同设计克服了传统光纤束内窥镜在近红外二区(NIR-II)成像中的空间分辨率低、蜂窝伪影和纤芯间串扰等限制。该平台优化了超薄光纤束以减少串扰引起的图像模糊,并开发了一种名为GAME(Agent-Guided Mixture-of-Experts)的智能处理流程,用于去除伪影和恢复图像,实现了超越奈奎斯特-香农采样极限的四倍分辨率提升。
Details
Motivation: 传统光纤束内窥镜自20世纪50年代首次报道以来,一直受限于低空间分辨率、蜂窝伪影和纤芯间串扰,尤其在能提供更优成像对比度和组织穿透深度的近红外二区(NIR-II)波段,串扰问题更为突出。
Result: 该平台在活体小鼠解剖结构以及数字微镜器件(DMD)投影的人体胃管和淋巴系统的NIR-II成像中证明了其效用,GAME流程为从细胞、小鼠到人体样本的多种生物医学图像提供了统一的恢复入口,并实现了超越奈奎斯特采样极限的四倍分辨率提升。
Insight: 创新点在于光学硬件(优化超薄光纤束以减少串扰)与计算算法(GAME智能专家混合处理流程)的协同设计。GAME流程利用视觉-语言模型动态路由输入到合适的恢复专家,为跨光谱范围(可见光到NIR-II)和多样本类型的高保真图像传输与重建提供了灵活高效的解决方案。
Abstract: Fiber-bundle endoscopy offers a compact and flexible route for clinical fluorescence imaging through natural human orifices, but since its first report in the 1950s, it has remained limited by low spatial resolution, honeycomb artifacts, and inter-core crosstalk. The crosstalk becomes more pronounced at near-infrared-II wavelengths (NIR-II, 1000-3000 nm), a spectral window that offers superior contrast, resolution, and tissue penetration depth for biomedical imaging. Here, we present an AI-powered flexible endoscopy platform that overcomes these constraints through optical-computational co-design: optimizing ultrathin fiber bundles to mitigate crosstalk-induced image blur and enable high-fidelity image transmission across the visible-to-NIR-II spectral range, and developing an Agent-Guided Mixture-of-Experts (GAME) pipeline for honeycomb-artifact removal and image restoration. GAME provides a single restoration entry point for diverse biomedical images acquired with our endoscope, spanning cell, mouse and human samples. It dynamically routes each input to suitable restoration experts via a vision-language model, facilitating image reconstruction with a fourfold resolution improvement beyond the NyquistShannon sampling limit. The utility of our endoscope is demonstrated through in vivo NIR-II imaging of anatomical structures in mice, as well as imaging of the digital micromirror device (DMD)-projected human gastric tube and lymphatic system, paving the way for future clinical translation.
[107] DoRF++: Spherical Representation Learning over Doppler Radiance Fields for Robust Wi-Fi Sensing cs.CV | eess.SPPDF
Navid Hasanzadeh, Shahrokh Valaee
TL;DR: 本文提出DoRF++,一种用于Wi-Fi感知的球形表示学习方法。该方法将Wi-Fi信道状态信息(CSI)提取的多普勒速度投影建模为人体运动的稀疏虚拟相机视图,并利用神经辐射场(NeRF)概念推断潜在3D运动序列,最终在单位球面上生成球形运动表示,结合球形Transformer进行分类,以提升跨用户手势识别的鲁棒性。
Details
Motivation: 针对Wi-Fi感知在真实世界变异性下泛化能力不足的挑战,研究旨在利用多普勒速度投影(直接反映人体运动速度)实现更鲁棒、跨用户泛化能力更强的人类活动识别(HAR),以推动IEEE 802.11bf标准下的先进WLAN感知应用。
Result: 在自收集的手势数据集上,DoRF++在跨用户泛化准确率上显著优于最先进的基于Wi-Fi的HAR方法,特别是在单多天线接收接入点(AP)设置下的困难手势识别场景中。
Insight: 创新点在于将计算机视觉中的NeRF概念引入Wi-Fi感知,将多普勒投影建模为虚拟视图并学习有效多普勒方向,进而提出球形表示学习框架(DoRF++),利用球形Transformer处理球面数据,增强了模型对运动表示的几何一致性和泛化能力。
Abstract: Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-free, and privacy-preserving activity and gesture recognition has grown rapidly. Recent studies have shown that Doppler velocity projections extracted from CSI, which directly reflect human-motion velocity, enable more robust human activity recognition (HAR) and stronger generalization across users and unseen conditions. Nevertheless, reliable generalization under real-world variability remains a major challenge, hindering the adoption of Wi-Fi sensing in real-world applications. To address this challenge, we introduce Doppler Radiance Fields (DoRF), bringing the concept of neural radiance fields (NeRF) from computer vision into Wi-Fi sensing. DoRF models Doppler velocity projections extracted from Wi-Fi CSI as sparse and diverse virtual-camera views of human motion. It then infers a latent 3D motion sequence whose projections along learned effective Doppler directions explain the CSI-derived Doppler observations. The recovered motion is subsequently projected onto an equiangular grid of directions on the unit sphere, producing a spherical representation of the underlying motion. Since DoRF naturally defines the Doppler representation on spheres, we further introduce DoRF++, a spherical-learning design that applies spherical Transformers for activity classification. Experiments on our collected hand-gesture dataset show that DoRF++ significantly outperforms state-of-the-art Wi-Fi-based HAR methods in cross-user generalization accuracy, especially for difficult gestures in settings with a single multi-antenna receiver access point (AP).
[108] InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions cs.CVPDF
Shun Okamoto, Satoshi Iizuka, Kazuhiro Fukui
TL;DR: InstructionCrafter是一个基于扩散模型的框架,用于根据文本任务指令生成高质量、一致的视觉指令图像序列。它通过空间层冻结训练和指令感知适配器,分离了时间与指令对齐的优化与单帧视觉质量,在保持单帧细节的同时学习指令语义和步骤间关系。
Details
Motivation: 现有文本到图像生成方法难以同时满足步骤忠实性、跨图像一致性和单帧视觉质量这三个关键属性,主要由于独立采样破坏一致性、在低质量视频上微调降低单帧质量,以及冻结主干网络缺乏多步骤理解。
Result: 在两个基准数据集上的广泛实验表明,该方法在步骤忠实性、跨图像一致性和单帧视觉质量方面实现了最先进的整体性能,同时显著减少了噪声、模糊和虚假字幕。
Insight: 创新点在于提出了空间冻结训练策略以保留生成先验并减少可训练参数,并设计了两种轻量级适配器(一致性适配器和上下文感知时间适配器)来增强模型对指令上下文的理解和跨帧关系的显式传播。
Abstract: Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple properties, specifically step faithfulness, cross-image consistency, and per-frame visual quality. Existing text-to-image generation approaches rarely meet all three properties, owing to independent sampling that breaks consistency, finetuning on low-quality video that degrades per-frame quality, and frozen backbones that lack multi-step understanding. In this work, we propose InstructionCrafter, a diffusion-based framework with the key idea of separating the optimization of temporal and instructional alignment from per-frame visual quality via (1) spatial-freeze training and (2) instruction-aware adapters. Built on a pretrained video diffusion backbone, InstructionCrafter freezes the spatial layers that control per-frame detail and updates only temporal and text-conditioning pathways to learn instruction semantics and inter-step relations, which preserves the generative prior for per-frame quality and reduces trainable parameters by about 50 percent compared with full finetuning. We also introduce two lightweight adapters that enhance the model’s understanding of instructional context. The Consistent Adapter aggregates textual cues from the entire instruction sequence and from neighboring steps to keep object identity and attributes consistent across frames, and the Context-Aware Temporal Adapter converts cross-attention outputs into biases for temporal self-attention, explicitly propagating inter-frame relations. Extensive experiments on two benchmark datasets demonstrate state-of-the-art overall performance on step faithfulness, cross-image consistency, and per-frame visual quality while significantly reducing noise, blur, and spurious subtitles. Our code and trained models will be publicly available.
[109] RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting cs.CVPDF
Zecheng Ren, Yafei Hu, Jianing Zhao, Ruichen Cong, Qun Jin
TL;DR: 本文提出了RenderMatte,一个基于FLUX.1 Kontext模型进行全参数微调的trimap引导图像抠图框架,旨在解决开放世界场景中前景外观和透明度模式高度多样导致的语义模糊和细粒度透明度变化问题。该方法通过alpha边缘目标在微调中保持潜在流匹配信号并加强像素空间边界监督,并引入了基于组相对alpha对齐的后训练优化策略。
Details
Motivation: 开放世界场景中真实前景具有高度多样化的外观和透明度模式,现有方法在处理语义模糊和细粒度透明度变化(尤其是在稀疏、脆弱且难以监督的边界区域)时面临挑战,需要一种能够精确估计alpha通道的方法。
Result: 实验表明,该方法在所有基准测试上都达到了最先进的(SOTA)性能,证明了其在开放世界场景中实现高保真抠图的可扩展路径。
Insight: 主要创新点包括:1)通过全参数微调FLUX.1 Kontext并利用图像编辑先验进行结构保持的alpha预测;2)在监督微调中引入alpha边缘目标以协同优化潜在空间和像素空间;3)提出了基于组相对alpha对齐的后训练优化,使用抠图专用奖励(alpha精度、边界保真度、trimap合规性、合成一致性)对同一trimap条件下采样的多个matte进行比较和优化;4)构建了具有精确发丝级alpha标注和多样化背景合成的大规模合成数据集RenderMatte,以克服精确边缘标注的缺乏。
Abstract: Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editability of downstream creation workflows. However, precise alpha estimation in open-world scenes remains challenging because real foregrounds exhibit highly diverse appearances and opacity patterns. This makes existing methods struggle with semantic ambiguity and fine-grained opacity variation, especially in sparse boundary regions that are fragile and difficult to supervise. To address this gap, we present RenderMatte, a trimap-guided matting framework that adapts FLUX.1 Kontext through full-parameter fine-tuning, leveraging image editing priors for structure-preserving alpha prediction. During supervised adaptation, an alpha-edge objective preserves the latent flow-matching signal while strengthening pixel-space boundary supervision. We further introduce group-relative alpha alignment for post-training. It compares multiple mattes sampled under the same trimap condition using matting-specific rewards for alpha accuracy, boundary fidelity, trimap compliance, and compositional consistency. To overcome the lack of precise edge annotations, we construct the RenderMatte dataset, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets. It features exact strand-level alpha annotations and diverse background composites. Experiments show state-of-the-art performance across all benchmarks, demonstrating a scalable path toward high-fidelity matting in open-world scenes.
[110] RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion cs.CVPDF
Meng Wang, Hongxia Yu, Wenzhe He, Xingdong Song, Huilong Pi
TL;DR: RayLift是一个用于基于相机的3D语义场景补全(SSC)的框架,旨在解决现有方法将立体深度估计视为确定性几何约束所导致的深度不确定性和局部对应误差传播问题。该框架利用立体几何作为度量参考,并结合互补的射线证据来自适应地恢复可靠的3D结构,通过互补上下文编码器、深度射线证据提升模块和语义感知体素集成器实现。
Details
Motivation: 现有基于相机的SSC方法通常将立体深度估计视为确定性约束,导致深度不确定性和局部对应误差直接传播到体素表示中,影响场景理解的可靠性。
Result: 在SemanticKITTI和SSCBench-KITTI-360基准测试上的大量实验表明,RayLift取得了有竞争力的性能,并持续优于现有方法。
Insight: 创新点在于将立体几何作为度量参考而非硬约束,并联合建模几何差异、深度置信度和空间不确定性来自适应采样和加权候选表面位置;同时利用冻结的3D视觉基础模型提取几何感知先验以丰富场景上下文,并通过显式建模空间支持将射线证据注入体素特征。
Abstract: Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.
[111] Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation cs.CVPDF
Qiang Hu, Yuxuan Luo, Yingjie Guo, Hao Wang, Qimei Wang
TL;DR: 本文提出了一种名为DPO-Clin的新型后训练框架,旨在提升医学报告生成(MRG)的可靠性。该框架通过实体级临床诊断模块分离临床发现与语言特征,并引入检索增强的多模态DPO变体M²DPO来实现细粒度视觉-语言对齐,同时利用反事实修改构建偏好数据以降低潜在风险。
Details
Motivation: 现有基于直接偏好优化(DPO)的MRG方法通常采用简单的偏好构建策略,直接将模型生成的报告与真实报告配对,这无意中将关键临床发现与临床无关的语言特征混为一谈,并且缺乏明确的视觉-语言对齐。
Result: 在两个公开的胸部X光数据集(MIMIC-CXR和IU X-Ray)和一个内部内窥镜数据集上的广泛实验表明,DPO-Clin在临床感知指标上显著优于监督微调(SFT)基线,并且超越了现有的基于DPO的MRG方法,在不同基线架构和多样医学成像模态上表现出强大的泛化能力。
Insight: 创新点包括:1)实体级临床诊断模块实现语言对齐的报告偏好对构建,隔离临床差异与语言变化;2)检索增强的多模态DPO变体M²DPO,通过视觉上下文切换触发文本偏好反转,实现细粒度跨模态对齐;3)针对正确但高度不确定的预测实体进行反事实修改,构建针对性偏好数据以缓解潜在风险。
Abstract: Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct Preference Optimization (DPO) has emerged as a promising post-training paradigm to enhance the performance of Supervised Fine-Tuned (SFT) MRG models, existing DPO-based MRG methods typically adopt a naive preference construction that directly pairs model-generated reports with ground truth reports. This strategy inadvertently entangles critical clinical findings with clinically irrelevant linguistic characteristics, and fundamentally lacks explicit vision-language alignment. To address these challenges, we propose DPO-Clin, a novel post-training framework that focuses preference optimization on clinical findings and cross-modal alignment. First, we introduce the Entity-level Clinical Diagnostic (ECD) module to perform a precise entity-level factual diagnosis. ECD guides the generation of linguistically-aligned report preference pairs, isolating clinical discrepancies from linguistic variations. Second, to achieve fine-grained cross-modal alignment, we develop M$^2$DPO, a retrieval-augmented multi-modal DPO variant that enforces textual preference inversion triggered by visual context switches. Third, we locate correct yet highly uncertain predicted entities and apply counterfactual modifications to construct targeted preference data for latent risk mitigation, thereby further enhancing the model reliability. Extensive experiments on two public chest X-ray datasets (MIMIC-CXR and IU X-Ray) and an in-house endoscopy dataset demonstrate that DPO-Clin significantly improves the SFT baselines on clinical-aware metrics. Furthermore, it achieves superior performance over existing DPO-based MRG methods, exhibiting robust generalizability across distinct baseline architectures and diverse medical imaging modalities.
[112] Towards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation cs.CVPDF
Ajeet Kumar Verma
TL;DR: 该博士研究提出了一种基于测试时适应(TTA)的统一框架,用于解决真实世界条件下视频超分辨率(VSR)的泛化问题以及无参考视频质量评估(VQA)。具体包括:1)一个TTA驱动的无参考VQA方法,为VSR提供感知质量引导;2)一个基于Transformer的屏幕内容超分辨率架构,以保持文本清晰度和结构保真度;3)一种无需高分辨率真值的区域感知TTA策略,可选择性优化文本与非文本区域。
Details
Motivation: 现有视频超分辨率方法在面对由异构设备、编解码器和网络环境产生的未知退化时,泛化能力不足。研究旨在通过测试时适应(TTA)范式,在不重新训练或依赖高质量监督的情况下,提升模型的鲁棒性和感知质量。
Result: 在多个基准测试上的实验结果表明,该方法在感知质量和可读性方面取得了一致的提升。
Insight: 创新点在于将测试时适应(TTA)统一应用于视频质量评估和超分辨率任务,实现了在未知失真下的自适应增强;提出的区域感知TTA策略能够针对性地处理文本和非文本区域,无需高分辨率真值监督,这在屏幕内容处理中具有实用价值。
Abstract: This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video super-resolution (VSR) methods struggle to generalize under unknown degradations arising from heterogeneous devices, codecs, and network environments. We address this challenge through test-time adaptation (TTA), a unified paradigm that improves robustness and perceptual quality without retraining or high-quality supervision. Specifically, we: 1) propose a TTA-based framework for no-reference video quality assessment (VQA), where adapted quality predictions provide perceptual guidance for VSR under unseen distortions; 2) develop a transformer-based architecture for screen-content super-resolution that preserves text clarity and structural fidelity; and 3) introduce a region-aware TTA strategy that selectively refines text and non-text regions without requiring high-resolution ground truth. Experimental results across diverse benchmarks demonstrate consistent improvements in perceptual quality and readability. We also outline ongoing work toward fully adaptive video enhancement systems capable of generalizing across unseen domains.
[113] A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems cs.CV | cs.AI | cs.CRPDF
Sangiya Pararajasingham
TL;DR: 本文提出并比较了五种结合特征提取与分类的流程(PM、LPM、HPM、SM、HM),用于在统一框架内同时解决人脸识别系统中的伪装(disguise)和欺骗(spoofing)检测问题。这些方法基于预处理、特征提取、特征过滤和分类的两阶段流程,并在包含FEI、Disguised Faces Database和NUAA数据库的115个受试者数据集上进行训练与评估。
Details
Motivation: 人脸识别系统面临两种通常被分开处理的失效模式:欺骗(冒名者使用授权用户的照片或视频)和伪装(合法用户因配饰、胡须、光照或姿态变化导致外观与注册模板不同而被拒绝)。论文旨在通过一个单一框架同时解决这两个问题。
Result: 在涵盖混合外观、正面人脸、暗光照、左右转向姿态和照片欺骗尝试的六种测试条件下进行评估。基于HOG的流程(HPM)在不同条件下表现最一致,在混合外观伪装上准确率达94.59%,在姿态和光照变体上为81.5-93.2%,在欺骗检测上为91.67%。基于LBP的流程(LPM)在欺骗检测上准确率最高(93.2%),但对姿态变化的鲁棒性较弱。
Insight: 论文的创新点在于将伪装和欺骗检测整合到一个统一的特征提取与分类框架中,并比较了多种经典特征表示(如PCA、LBP、HOG、SURF、Harris角点)的性能。客观分析表明,不同特征在欺骗敏感性和伪装鲁棒性之间存在可衡量的权衡,这为后续的深度学习和跨数据库扩展提供了动机。
Abstract: Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of an authorized user, and disguise, where a legitimate user is rejected because their appearance differs from their enrolled template due to accessories, facial hair, illumination, or pose. This paper proposes and compares five combined feature-extraction and classification pipelines that address both problems within a single framework: PM (PCA and Minimum Euclidean Distance, MED), LPM (Local Binary Patterns with PCA and MED), HPM (Histogram of Oriented Gradients with PCA and MED), SM (Speeded-Up Robust Features with MED), and HM (Harris corner features with MED). Each pipeline follows a common two-phase process comprising pre-processing, feature extraction, feature filtering, and classification. The methods were trained on 115 subjects drawn from the FEI, Disguised Faces Database, and NUAA databases and evaluated on six test conditions covering mixed appearances, frontal faces, dark illumination, left- and right-turned poses, and photo-spoofing attempts. The HOG-based pipeline (HPM) achieved the most consistent performance across conditions, with 94.59% accuracy on mixed-appearance disguise, 81.5-93.2% across pose and illumination variants, and 91.67% on spoofing, while the LBP-based pipeline (LPM) achieved the highest spoofing-detection accuracy (93.2%) but weaker robustness to pose change. These results reveal a measurable trade-off between spoof sensitivity and disguise robustness among classical feature representations, motivating the deep-learning and cross-database extensions discussed in the concluding sections.
[114] ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints cs.CVPDF
Xiaoyang Bai, Zhenyang Li, Weiwei Xu, Edmund Y. Lam, Yifan Peng
TL;DR: 本文提出了一种名为ERF-GS的事件-RGB融合高斯溅射框架,用于从分离的事件和RGB视角重建快速运动场景。该方法将事件信息集成到高斯溅射管线的优化和致密化阶段,以利用高帧率事件传感器的优势,有效处理自然视频中低帧率、严重运动模糊和复杂布局的挑战。
Details
Motivation: 现有基于传统帧视频的动态3D场景重建方法(如NeRF、3DGS)在处理快速运动物体(如体育赛事、动物摄影)时存在困难,因此需要融合高帧率事件数据来提升重建性能。
Result: 在包含模糊RGB帧和分离RGB-事件视角的Neu3D和Nvidia数据集变体上,ERF-GS的表现优于4DGS基线方法和并发的E-D3DGS方法。
Insight: 创新点在于将事件信息深度集成到高斯溅射的优化与致密化流程中,并采用脱离RGB输入的纯事件学习设计,这使得方法能够从模拟数据推广到复杂的自然视频场景,提升了在快速运动和模糊条件下的重建鲁棒性。
Abstract: Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D scene reconstruction with improved visual precision and scalability. However, the reconstruction of fast-moving objects remains a challenge; existing methods based on conventional frame-based videos often struggle in scenarios such as sports events and animal videography. We propose an event-RGB fusion Gaussian splatting (ERF-GS) framework that integrates event information into both optimization and densification stages of the Gaussian splatting pipeline, taking advantage of novel event sensors with high frame-rate. Unlike many other event-assisted scene reconstruction methods, ERF-GS was developed using realistic simulation settings and realizes event-based learning detached from RGB inputs. This design enables its application beyond straightforward synthetic data into the realm of natural video with complex layout, low frame rates and severe motion blur. Our experiments show that ERF-GS outperforms both the 4DGS baseline and the concurrent E-D3DGS on different variants of the Neu3D and Nvidia datasets which include blurry RGB frames and disjoint RGB-event viewpoints. Our code is available at https://github.com/andrewbxy/ERF-GS.
[115] Goal-oriented Navigation Instruction Generation with Tour Video Priors cs.CVPDF
Fangdi Li, Juncheng Liao, Changxu Cheng, Jiazhi Wang, Senda Chen
TL;DR: 本文提出了VideoNIG任务,旨在从第一人称游览视频、初始观察和文本/视觉目标中生成面向目标的导航指令,而无需依赖图或地图等中间表示。作者构建了一个包含6万条视频和3.7万条多模态提示的模拟器基准,并设计了一个结合文本相似度、空间一致性测试和下游导航执行的诊断评估协议。为解决此任务,论文提出了一个两阶段课程学习框架,通过动作预热和复杂度渐进训练来提升模型性能。实验表明,现有MLLMs在此任务上表现不佳,而所提方法显著提升了指令质量,并验证了生成指令对端到端导航的可执行性。
Details
Motivation: 现有导航指令生成研究主要作为视觉语言导航的辅助任务,而从紧凑的环境先验生成指令需要精细的空间推理,尤其是在目标路线不简单跟随演示游览时,这对当前多模态模型仍具挑战性。
Result: 在包含连续室内环境和渐进难度提示的模拟器基准上进行了广泛实验,结果表明现有MLLMs在VideoNIG任务上表现挣扎,而所提出的课程学习方法在互补的诊断指标上显著提升了指令质量。
Insight: 创新点在于提出了一个不依赖中间表示的、目标导向的视频接地导航指令生成任务,并构建了相应的基准和诊断评估协议;方法上采用了两阶段课程学习框架,将学习分解为基础运动感知和长时域导航推理,通过动作对齐和渐进难度训练有效提升了模型的空间推理能力。
Abstract: Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary task for vision-andlanguage navigation (VLN), focusing on data augmentation or multi-task learning. However, generating navigation instructions from compact environmental priors requires meticulous spatial reasoning, especially when the target route does not simply follow the demonstrated tour, and remains challenging for current multimodal models. In this work, we introduce VideoNIG, a goal-oriented video-grounded NIG task that generates navigation instructions from ego-centric tour videos, an initial observation, and a textual or visual goal, without relying on intermediate representations such as graphs and maps. We instantiate VideoNIG in a controlled simulator benchmark with 60K tour videos across continuous indoor environments and 37K multimodal prompts with progressive difficulty levels. We further introduce a diagnostic evaluation protocol that combines text similarity, choice-based spatial consistency tests, and downstream navigation execution. To address this task, we propose a two-stage Curriculum Learning framework that decomposes the learning into foundational motion perception and long-horizon navigation reasoning. Specifically, we first employ Action Warmup for spatial action-view alignment, followed by Complexity Progression using trajectories with increasing exploratory difficulty. Extensive experiments show that existing MLLMs struggle with VideoNIG, while our approach significantly improves instruction quality across complementary diagnostic metrics. Finally, integrating VideoNIG-generated instructions with a VLN agent demonstrates the executability of this task formulation for end-to-end navigation.
[116] Population-Scalable Multi-Agent World Modeling cs.CV | cs.AI | cs.LGPDF
Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li
TL;DR: 本文提出了一种名为Khora的可扩展多智能体世界模型,旨在解决现有方法在训练和推理时智能体数量固定的限制。该模型通过解耦世界状态演化与视觉渲染,并引入与智能体数量无关的渲染机制,支持在推理时扩展到任意数量的智能体而无需重新训练。
Details
Motivation: 现有世界模型在多智能体环境中面临可扩展性挑战,因为它们通常假设训练和推理时的智能体数量固定,这限制了模型在推理时对智能体数量的适应能力。
Result: 定性实验表明,该方法能够泛化到未见过的智能体数量,同时保持视觉质量和多智能体一致性。此外,实现了一个实时交互系统来展示可扩展的开放世界模拟。
Insight: 核心创新在于将跨视图一致性建立在共享的世界状态上,该状态的演化不预设智能体数量,而智能体特定的观察则通过统一的渲染接口查询该状态生成。这种设计避免了在昂贵的视频生成器内进行密集的观察流交互,实现了与查询视图数量近似线性的实际扩展。
Abstract: World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed number of agents during training and inference, which ties the model to a pre-determined agent population and limits inference-time scalability. Our key insight is that cross-view consistency should arise from a shared world state whose evolution does not assume a predefined number of agents, while agent-specific observations should be generated by querying this state through a unified rendering interface. Based on this insight, we propose Khora, a scalable multi-agent world model that supports inference-time expansion to arbitrary numbers of agents without retraining. Our framework decouples world-state evolution from visual rendering and introduces a population-agnostic rendering mechanism for incorporating other agent information. This design maintains cross-view consistency through the shared world state rather than through dense interactions among observation streams inside the expensive video generator, enabling approximately linear practical scaling with the number of queried views. Qualitative experiments demonstrate that our approach generalizes to unseen numbers of agents while maintaining visual quality and multi-agent consistency. We further implement a real-time interactive system to demonstrate scalable open-world simulation.
[117] MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling cs.CV | cs.LG | cs.MMPDF
Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu
TL;DR: MotionCraft是一个可控的视频超分辨率框架,它将视频恢复任务构建为受世界模型启发的运动感知潜在状态预测。该框架结合了鲁棒的运动融合、一个平衡局部与目标非局部交互的潜在世界变换器,以及一个紧凑的条件解码器,旨在流式约束下提供时间一致的高质量重建。
Details
Motivation: 现有视频超分辨率方法在局部细节保真度、长程时空建模、感知真实性和效率之间存在权衡,例如卷积对齐技术在大运动或复杂退化时效果不佳,而基于变换器的方法计算成本高。本文旨在解决这些权衡问题,提出一个可控且高效的框架。
Result: 经验评估表明,MotionCraft在重建和感知性能方面表现强劲,同时能够在时间平滑度和重建保真度之间实现可预测的权衡。
Insight: 创新点在于将视频超分辨率视为运动感知的潜在世界建模,并集成了自适应稀疏注意力机制以平衡计算效率与长程依赖建模。此外,框架提供了明确的用户可访问控制接口,增强了可控性。
Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
[118] VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cs.CVPDF
Dong Xing, Jiaxin Chen, Hang Yang, Peixun Liu, Qiushi Yang
TL;DR: 本文提出VADER框架,一种无需训练的方法,通过视觉焦点重分配和选择性证据擦除两个互补模块,自适应地缓解视频大语言模型中的幻觉问题,提升事件级定位和时间一致性。
Details
Motivation: 现有免训练方法采用全局固定的视觉干预或通过输入扰动构建对比分支,前者无法适应视频依赖的融合路径,后者可能被跨帧冗余补偿,因此需要一种能自适应视频内容并有效缓解幻觉的方法。
Result: 在多个VideoLLM上,VADER显著提升了事件级定位和时间一致性;在LLaVA-Video-7B上,其在EventHallusion基准上达到72.60%的准确率。
Insight: 创新点在于提出视频自适应的干预策略,通过诊断层间视觉到文本的证据流动态决定干预位置和强度,并结合选择性证据擦除构建难以被跨帧冗余补偿的对比分支,从而更有效地抑制幻觉。
Abstract: Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence. Existing training-free methods typically apply a globally fixed visual intervention or construct a contrastive branch through input perturbation. The former cannot accommodate video-dependent fusion paths, while the latter can be compensated by cross-frame redundancy. We therefore propose Video-Adaptive Debiasing via Evidence Reweighting (VADER), a training-free framework with two complementary modules. Visual Focus Reallocation (VFR) automatically instantiates an intervention policy for each video-question input: it diagnoses layer-wise visual-to-text evidence flow, determines where to intervene, and derives how strongly to reallocate pre-softmax attention from system-token to video-token blocks. Selective Evidence Erasure (SEE) independently masks high-importance visual tokens in every frame, constructing a prior-biased branch that is difficult to compensate through neighboring frames. Contrastive decoding then down-weights predictions that remain confident after selective evidence erasure. Across multiple VideoLLMs, VADER yields substantial improvements on event-level grounding and temporal consistency; on LLaVA-Video-7B, it reaches 72.60% accuracy on EventHallusion.
[119] REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering cs.CV | cs.AIPDF
Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li
TL;DR: 本文提出了REVEAL框架,一种基于评分标准引导的智能体,用于在长视频问答中显式验证证据的充分性。该框架通过自适应视觉相似性预处理构建离线-在线视频记忆,并利用自动构建的评分标准库来指导证据检索与验证,从而在无需额外训练的情况下超越现有SOTA方法。
Details
Motivation: 现有长视频问答方法通常依赖固定长度的时间分块和静态离线记忆,这会割裂连续事件且无法在实时推理中适应;同时,这些方法过于关注检索相关性而忽视了证据充分性,导致在关键时序、因果或细粒度动作证据缺失时仍过早停止回答。
Result: 在广泛的实验中,REVEAL无需任何额外训练,在多个基准测试上一致超越了闭源和开源的最先进方法,证明了显式验证证据充分性能够检索到先前方法遗漏的关键线索,从而实现更可靠的长视频推理。
Insight: 创新点在于引入了评分标准引导的显式证据充分性验证机制,以及自适应视觉相似性预处理构建的离线-在线视频记忆结构,这使系统能够动态维护问题条件化记忆并指导针对性重检索,从而提升了长视频问答的准确性和可靠性。
Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline memory banks, which not only fragment coherent continuous events but also fail to adapt during real-time reasoning. Moreover, whether using multi-scale summaries or multimodal knowledge graphs, current approaches prioritize retrieval relevance while overlooking evidence sufficiency, often stopping to answer once only semantically relevant clues are retrieved, even when key temporal, causal, or fine-grained action evidence is still missing. To tackle these challenges, we propose REVEAL, a rubric-guided agent framework. As a foundation, we introduce an adaptive visual-similarity-based preprocessing pipeline that groups visually coherent adjacent frames into natural event units to construct an offline-online video memory—capturing global video context offline while dynamically maintaining question-conditioned memory online. Built upon this structured memory, REVEAL uses an automatically constructed rubric library to explicitly verify whether retrieved evidence satisfies sufficiency criteria, pinpoints missing clues upon verification failure, and directs targeted re-retrieval for complementary information. Without any extra training, REVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments. These results show that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasoning.
[120] VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling cs.CVPDF
Yuqi Zhang, Cheng Chen, Yuyu Guo, Wenjie Yang, Lingchen Meng
TL;DR: 本文提出VLZip框架,通过统一视觉和文本压缩来解决视觉语言模型(VLM)处理超长交错图像-文本序列时自注意力二次复杂度带来的挑战。该方法将视觉和文本片段分层蒸馏为紧凑的层特定“软前缀”并注入解码器隐藏状态,从而大幅缩短注意力序列并保留细粒度全局上下文。论文还引入了LongVLBench基准,并在实验中实现了高达120K令牌的训练和280K令牌的推理,内存可扩展至2M令牌,为长上下文多模态AI建立了高效新标准。
Details
Motivation: 现有方法在处理超长交错图像-文本序列时,要么采用激进的令牌剪枝导致信息丢失,要么使用高效但精度较低的架构,且往往忽略文本组件的重要性,因此需要一种能同时压缩视觉和文本信息并保持高保真推理的解决方案。
Result: 在引入的LongVLBench基准(源自视频叙事)上,VLZip在长上下文多模态推理中取得领先性能,训练令牌数提升至120K(比基线增加6倍),推理可超过280K令牌且内存显著减少,并展示了处理高达2M令牌的内存可扩展性,在现有方法失效的极端上下文长度下表现出色。
Insight: 创新点在于统一视觉和文本的分层压缩机制,通过生成层特定的“软前缀”注入解码器以缩短注意力序列,同时保留全局上下文;客观分析认为,该方法在兼顾效率和精度方面提供了新思路,并通过专门基准LongVLBench弥补了领域评估不足的问题。
Abstract: Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific “soft prefixes” and injects them into each decoder layer’s hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at https://github.com/ShareLab-SII/VLZip.
[121] Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception cs.CVPDF
Jingyun Chen, Fengchun Liu, Linghan Cai, Songhan Jiang, Shenjin Huang
TL;DR: 本文提出了AdaptivePath,一种用于全切片病理图像视觉推理的主动感知框架。该框架将证据获取建模为序列决策过程,包含导航器、形态解释器、审议器和仲裁器四个模块,通过分层稀疏观察实现高效的病理图像分析。
Details
Motivation: 现有全切片图像方法要么将密集采样补丁压缩为全局表示,要么使用预训练视觉语言模型结合启发式区域选择,导致预测与形态学联系减弱或缺乏病理学训练的观察策略。
Result: AdaptivePath在WSI和区域病理VQA基准测试中实现了零样本最先进性能,在六个TCGA队列的癌症亚型分类中达到80.14%准确率。在盲法诊断效用研究中,病理学家使用该框架选择的观察序列实现了82.9%的准确率。
Insight: 创新点在于将病理图像分析建模为主动感知的序列决策问题,通过问题无关的异常驱动导航避免昂贵的轨迹标注,结合交替表示学习和策略优化训练观察策略,实现了可追溯的高效视觉推理。
Abstract: Whole-slide visual reasoning requires identifying sparse diagnostic evidence in gigapixel pathology slides and integrating observations across spatial scales. Existing WSI methods either compress densely sampled patches into global representations or use pretrained vision-language models with heuristic region selection, weakening links between predictions and morphology or lacking pathology-trained observation policies. We present AdaptivePath, an active-perception framework that formulates WSI evidence acquisition as sequential decision making. The Navigator learns question-agnostic abnormality-driven navigation from pathologist-reviewed labels to select observation locations and spatial extents, avoiding costly question-specific trajectory annotations. We train this policy through alternating representation learning and proximal policy optimization, followed by fine-tuning with geometric and appearance consistency objectives to stabilize focus trajectories. During inference, the Navigator hierarchically acquires sparse observations from low to high magnification under a limited ROI budget. A Morphology Interpreter converts observations into question-conditioned evidence, while the Deliberator evaluates evidence and revises intermediate answers across magnifications. The Arbiter integrates deliberation history to produce final answers. AdaptivePath achieves state-of-the-art zero-shot performance on WSI and region pathology VQA benchmarks and reaches 80.14% accuracy for cancer subtype classification across six TCGA cohorts. In a blinded diagnostic-utility study, pathologists using AdaptivePath-selected observation sequences achieve 82.9% accuracy. These results demonstrate that learned active perception enables effective and traceable visual reasoning over gigapixel pathology slides.
[122] JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views cs.CVPDF
Jinhua Cui, Anhong Wang, Kai Hu, Donghan Bu, Peihao Li
TL;DR: 本文提出了一种名为JSGS的方法,用于处理混合质量JPEG图像在3D高斯泼溅(3DGS)重建中的问题。该方法利用JPEG文件中的亮度和色度量化表构建视图特定的JPEG观测算子,对渲染视图进行编码和解码,以实现与输入图像的域匹配比较,并通过频率带损失和块不一致性指导来优化高斯原语。
Details
Motivation: 标准3D高斯泼溅假设所有输入图像都能准确采样场景辐射,但混合质量JPEG图像中的压缩伪影(如块状和振铃效应)会破坏跨视图共享的高斯更新,导致重建质量下降。
Result: 在七个场景和三种混合质量调度下,JSGS在所有调度中实现了最低的平均LPIPS和最高的平均SSIM,同时渲染速度约为150 FPS,达到了当前最佳水平(SOTA)。
Insight: 创新点在于利用JPEG内部量化表构建观测算子进行域匹配监督,并通过频率带损失加权和块不一致性指导来正则化高斯原语,这为处理压缩图像在3D重建中的问题提供了新思路。
Abstract: Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS). JSGS uses luminance and chrominance quantization tables stored in each JPEG file to construct a view-specific JPEG observation operator. This operator encodes and decodes each rendered view for domain-matched comparison with the corresponding decoded input image. The luminance quantization table supplies continuous weights within a fixed middle frequency band. A loss in the low frequency band anchors coarse structure, while the weighted middle frequency loss redistributes supervision among the selected DCT coordinates. The resulting block disagreement also guides the Gaussian Controller to regularize small primitives with high opacity in disagreement regions. Across seven scenes and three mixed-quality schedules, JSGS achieves the lowest mean LPIPS and the highest mean SSIM under every schedule while rendering at approximately 150 FPS. Code: https://github.com/Jayden-Cui/JSGS.
[123] Semi-Dense Matching Uncertainty Is Not Just Local Confidence cs.CVPDF
Khoa Hoang, Hoang-Tuan Nguyen, Huong Ninh, Hai Tran, Long Q. Tran
TL;DR: 本文提出了一种轻量级的事后整体不确定性估计框架,用于半稠密匹配任务。该框架采用仅含9个可学习参数的双分量校准拉普拉斯混合模型,旨在同时捕捉局部精细化噪声和粗分配失败的宽尾分布。通过引入基于粗分配成功后验概率的几何重拟合模块(CoRe方法),该方法能有效提升下游几何估计的准确性。
Details
Motivation: 现有半稠密匹配方法通常在粗到精范式下实现性能与计算成本的平衡,但难以提供良好量化的不确定性估计,尤其忽略了灾难性的粗分配失败,导致误差分布被截断并严重误导几何估计。
Result: 大量实验表明,该方法在各种预训练匹配器和鲁棒估计器上,以极小的计算开销,持续提升了下游几何精度。
Insight: 创新点在于提出了一个参数极少的双分量混合模型来显式建模匹配不确定性的两种不同来源(局部噪声与粗分配失败),并设计了利用粗分配成功后验概率作为软对应权重的几何重拟合模块(CoRe),这是一种轻量且事后可应用的不确定性校准与利用方法。
Abstract: Reliable semi-dense matching is essential for modern geometric vision systems. Designed under a coarse-to-fine paradigm, it achieves an optimal balance between performance and computational cost. However, existing methods often struggle to provide well-quantified uncertainties, where catastrophic coarse-assignment failures are ignored, leading to truncated error distributions and severely misjudged geometric estimations. In this paper, we propose a lightweight, post-hoc overall uncertainty estimation framework that introduces a two-component calibrated Laplace mixture model with only 9 learnable parameters. The objective is to explicitly capture both the sharp local refinement noise and the broader tail of coarse-assignment failures. We introduce the Coarse-success posterior Refit (CoRe) method, a geometric refitting module that utilizes the posterior probability of coarse-assignment success as soft correspondence weights. Extensive experiments show that our method consistently improves downstream geometric accuracy across various pretrained-only matchers and robust estimators with minimal computational overhead. Our code is available at https://github.com/khoavpt/Probabilistic-matching.
[124] UniSpace: Unified Visual Representation and Scalable Multimodal Modeling cs.CV | cs.AIPDF
Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He, Feize Wu
TL;DR: 本文提出了UniSpace,一种统一视觉表示和可扩展多模态建模方法。通过引入“补丁重参数化”技术,在预训练语义ViT的冻结Transformer块中同时保留语义理解和细粒度视觉细节,从而在单一视觉空间中实现理解、生成和编辑任务。
Details
Motivation: 现有语义视觉编码器的最终token丢弃了细粒度视觉细节,导致像素重建质量差,限制了其在图像生成和编辑等重建敏感任务中的应用。本文旨在探索是否能在预训练语义ViT构建的单一视觉表示空间中统一建模理解、生成和编辑。
Result: 所提出的统一表示在保持多模态理解能力的同时,实现了高保真图像重建和有利的重建-生成权衡。进一步扩展为80亿参数的UniSpace模型,在系统级评估中展示了实用的文本到图像生成和基于指令的图像编辑性能。
Insight: 核心创新点是补丁重参数化,它在保留原始语义通路的同时,增加了重建感知的补丁嵌入,为同一冻结ViT块提供细粒度视觉信息。这揭示了预训练ViT的冻结块并非本质上无法保留视觉细节,而是原始补丁参数化驱动了语义抽象;通过重参数化,预训练ViT可作为可扩展多模态建模的统一视觉接口。
Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce \emph{Patch Reparameterization}, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction–generation trade-off. We further scale this representation into \emph{UniSpace}, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.
[125] TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases cs.CV | cs.AIPDF
Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
TL;DR: 该论文提出了TomaMMU,一个大规模番茄叶部病害多模态理解数据集,以及TomaBench基准测试,用于系统评估视觉语言模型在番茄病害理解上的能力。数据集包含超过2.8万张高质量图像和21.3万个人工标注的视觉问答对,基准测试则通过三个层次(基础感知、病理理解、专家诊断)的七个任务来评估模型从低级视觉识别到高级诊断推理的综合性能。
Details
Motivation: 旨在解决当前视觉语言模型在农业领域,特别是番茄叶部病害的细粒度识别和基于事实的诊断推理方面存在的显著能力差距,缺乏一个系统性的评估基准。
Result: 在评估了14个最先进的视觉语言模型后,发现它们在具有挑战性的多项选择题和开放式问题上表现不佳,存在显著差距。然而,在TomaMMU数据集上进行简单的微调后,模型在挑战性多选题上的准确率大幅提升至96.09%,超过了近期的视觉语言模型。
Insight: 创新点在于构建了一个大规模、高质量、任务层次化的农业领域多模态理解基准,其系统性的三级任务分类法(基础感知、病理理解、专家诊断)为评估模型从感知到推理的完整能力提供了新框架。结果表明,领域针对性微调是缩小当前通用VLM与专业领域需求之间差距的有效且前景广阔的方向。
Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation. Building on this foundation, TomaBench organizes seven agricultural tasks into a hierarchical three-level taxonomy spanning Basic Perception, Pathology Understanding, and Expert Diagnosis, which together enable systematic evaluation from low-level visual recognition to high-level diagnostic reasoning. The tasks assess visual symptom recognition, taxonomic relationships, and diagnostic reasoning, offering a comprehensive view of how well models grasp plant pathology. Our results pronounced gaps in fine-grained recognition and factually grounded reasoning with 14 state-of-the-art VLMs, consistently underperforming on both challenging MCQs and open-ended questions. These results suggest that current VLMs struggle to translate visual perception into reliable diagnostic knowledge, motivating the need for targeted domain adaptation. Simple fine-tuning on TomaMMU substantially narrows this gap, boosting accuracy on challenging MCQs to 96.09%, outperforming recent VLMs, and pointing toward promising directions for future work. All data and code is available in https://huggingface.co/datasets/enalis/TomaMMU.
[126] Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation cs.CV | cs.AIPDF
Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl
TL;DR: 本文系统研究了在三维放射学报告生成任务中,如何高效地分配视觉token预算,以平衡计算效率与临床细节保留。通过评估多种视觉编码器、token压缩投影器和大型语言模型的组合,发现解剖学引导的感兴趣区域裁剪是最一致的策略,而提升输入分辨率的效果强烈依赖于投影器的选择。
Details
Motivation: 将视觉-语言模型应用于完整3D CT体积时,视觉编码器会产生大量视觉token,成为计算瓶颈。如何在输入视野、空间分辨率和视觉到语言的投影之间分配视觉token预算,以在压缩序列减少计算的同时保留临床相关细节,是一个关键的开放设计问题。
Result: 在匹配的LLM token预算下,最佳配置在CT-RATE和Merlin两个大型CT报告数据集的测试集上达到了最先进的临床宏观F1分数,分别为49.5和49.0。
Insight: 论文的创新点在于系统地探索了视觉token预算的分配策略,揭示了在3D放射学报告生成中,解剖学引导的裁剪是稳健的压缩方法,而分辨率提升的效益高度依赖于投影器的设计,这为构建高效且准确的医学视觉-语言模型提供了重要的设计指导。
Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of vision tokens per scan, making the visual sequence passed to the large language model (LLM) a primary computational bottleneck. Vision-to-language projectors can compress this sequence to reduce computation, but may discard clinically relevant detail; conversely, effective compression can accommodate higher-resolution inputs while keeping the downstream token count fixed. How this vision-token budget should be allocated across input field of view, spatial resolution, and vision-to-language projection therefore remains an open design question. We systematically evaluate four heterogeneous VEs (CNN- and ViT-based), five token-reducing projectors at up to 64x compression alongside a non-reducing MLP projector baseline, and five instruction-tuned LLMs (1.7B–4B) on two large-scale CT report datasets (CT-RATE and Merlin). At matched LLM token budgets, anatomy-guided region of interest cropping is the most consistent strategy, improving clinical macro F1 in 19 of 20 settings by +3.7 points on average for the 3D ViT Primus encoder and +1.1 for the slice-based 2D ViT Curia encoder. Increasing input resolution further is strongly projector-dependent: the PerceiverResampler, paired with higher-resolution Curia features, yields the strongest configuration in the resolution study on both datasets. Our best configurations achieve state-of-the-art clinical macro F1 on the test sets, reaching 49.5 on CT-RATE and 49.0 on Merlin. Code and models will be published upon publication.
[127] AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval cs.CV | cs.CL | cs.IRPDF
Haoyu Zuo, Yibo Yan, Xin Zou, Shuliang Liu, Yi Cao
TL;DR: 本文提出了AnchorFold,一种无需训练的‘聚焦-折叠’框架,用于压缩视觉文档检索(VDR)中的多向量索引。该方法通过递归注意力传播在视觉自注意力图上选择关键锚点,并将剩余标记折叠到这些锚点上进行聚合,从而在保持检索性能的同时显著减少存储和计算开销。
Details
Motivation: 现有的多向量视觉语言检索器通过后期交互实现细粒度检索,但存储和计算每个页面的数百个视觉块嵌入会带来巨大开销。现有的无需训练方法(如剪枝或合并)在激进压缩下性能会急剧下降,或在形成代表时未能明确优先考虑重要区域。
Result: 在ViDoRe v1/v2和REAL-MM-RAG基准测试上,使用三种不同的检索骨干网络,AnchorFold在压缩率γ≤0.20时,始终优于所有评估的无需训练基线方法。在ViDoRe v1/v2上,5倍压缩时平均保留了完整索引NDCG@5的98.3%,接近无损压缩;20倍压缩时仍保留92.4%。
Insight: 创新点在于提出了‘聚焦-折叠’框架和递归注意力传播机制,通过图中心性选择结构上重要的锚点标记,并在归一化检索空间中进行相似性驱动的折叠与加权聚合,从而在无需额外训练的情况下,实现了对非锚点贡献的有效保留和对关键区域的容量集中。
Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at $γ\leq 0.20$. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at $5\times$ compression, achieving near-lossless compression, and 92.4% at $20\times$ compression.
[128] UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs cs.CVPDF
Mengxian He, Yunyun sun, Ziyue Gao, Wengkei Lam, Shunyi Zhang
TL;DR: 该论文提出UPolarSQ框架,用于在近视眼底图像中分割视盘和视盘周围萎缩区域并进行生物标志物量化。该方法将笛卡尔坐标系下的感兴趣区域映射到极坐标系,利用改进的U-Net网络进行分割,并直接从极坐标掩码中提取临床生物标志物。
Details
Motivation: 近视引起的后极部重塑常伴随视盘变形和视盘周围萎缩,两者都是临床相关的结构生物标志物。但在笛卡尔坐标的眼底图像中,PPA常呈现为不规则且部分可见的月牙形,导致分割碎片化且量化依赖后处理。
Result: 在内部和外部队列上的实验表明,UPolarSQ改善了OD/PPA的分割效果,并支持可靠的、基于极坐标的近视分析生物标志物估计。
Insight: 创新点在于提出了统一的极坐标域框架,将分割和量化统一在共享的几何表示中。具体包括使用径向-角度解耦模块和边界感知辅助监督来建模各向异性的极坐标特征和径向边界过渡,以及从预测的极坐标掩码中确定性提取生物标志物。
Abstract: Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provide clinically relevant structural biomarkers. In Cartesian fundus images, however, PPA often appears as an irregular and partially visible crescent adjacent to the OD, leading to fragmented segmentation and post-processing-dependent quantification. We propose UPolarSQ, a unified polar-domain framework for OD/PPA segmentation and biomarker quantification in myopic fundus images. UPolarSQ first maps an OD-centered region of interest into polar coordinates, where OD and PPA boundaries can be represented as radial profiles. It then employs UPolarSeg, a U-Net-based segmentation network enhanced with a Radial-Angular-Decoupled Module and boundary-aware auxiliary supervision to model anisotropic polar features and radial boundary transitions. Clinical biomarkers, including disc shape and PPA-width-related measurements, are deterministically extracted from the predicted polar masks, aligning segmentation and quantification within a shared geometric representation. Experiments on internal and external cohorts demonstrate that UPolarSQ improves OD/PPA segmentation and supports reliable polar-native biomarker estimation for myopic analysis.
[129] LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation cs.CVPDF
Jinhong Zhu, Weiqi Yan, Shengchuan Zhang, Liujuan Cao
TL;DR: 本文提出了一种名为LASA(语言和源锚定对齐)的框架,用于解决领域泛化语义分割(DGSS)中的问题。该框架通过三个协同组件——文本和源引导的风格迁移(TSGST)、领域感知查询适配器(DAQA)和领域感知解码器优化器(DADO)——来缓解领域偏移,同时保持特征的完整性和判别性细节。
Details
Motivation: 现有方法(如风格随机化或特征归一化)在缓解领域偏移时,会因粗粒度操作导致特征流形扭曲,或因刚性设计抑制判别性、领域敏感的语义细节,从而损害特征完整性。本文旨在解决这些局限性。
Result: 在具有挑战性的基准测试上进行的广泛实验表明,该方法显著优于最先进(SOTA)的方法。
Insight: 创新点在于利用源特征作为结构锚点和视觉语言模型(VLM)先验作为细粒度指导来引导风格迁移,并通过领域感知的查询重校准和分布对齐来恢复被抑制的判别性细节,从而在泛化性和特征保真度之间取得更好平衡。
Abstract: Domain Generalization Semantic Segmentation (DGSS) focuses on generalizing knowledge from labeled source domains to unseen target domains where data is unavailable during the training phase. While conventional methods utilize style randomization or feature normalization to mitigate domain shifts, they often impair feature integrity. Specifically, style randomization distorts the underlying feature manifold due to its coarse-grained nature, while feature normalization suppresses discriminative, domain-sensitive semantic details owing to its rigid design. To address these limitations, we propose the Language-and-Source-Anchored Alignment (LASA) framework, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO). Concretely, the TSGST module addresses manifold distortion by utilizing source features as structural anchors and vision-language model (VLM) priors as fine-grained guidance. To restore suppressed discriminative and domain-sensitive details, the DAQA module recalibrates object queries via categorical guidance and domain-aware signatures, while the DADO module aligns the resulting query distributions with a shared classifier to ensure consistent categorical responses across domains. Extensive experiments on challenging benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches.
[130] 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents cs.CV | cs.AI | cs.LGPDF
Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
TL;DR: 本文提出了360CityArena,一个用于评估具身智能体在城市环境中探索能力的基准测试。该基准基于日本东京秋叶原地区的真实360度视频重建,包含175个人工设计的任务,涵盖环境理解、路径推理和空间推理三大类别。评估表明,即使是当前最强的LMM智能体(如Gemini 2.5 Flash)的表现也远低于人类水平,揭示了城市规模导航与推理仍面临巨大挑战。
Details
Motivation: 现有户外基准测试要么缺乏足够的真实感,要么复杂度不足,与现实城市环境存在显著差距。因此,需要构建一个高真实感、高复杂度的城市环境基准来推动具身智能体在真实场景中的导航与推理能力发展。
Result: 在360CityArena基准上,最先进的基于LMM的智能体(Gemini 2.5 Flash)的平均任务成功率仅为17.1%,远低于人类水平的77.3%,凸显了当前模型在城市规模具身导航与推理任务上的巨大性能差距。
Insight: 创新点在于利用真实世界的360度视频构建高保真、复杂的城市环境基准,并设计了涵盖定位、地标搜索、路径规划和关系空间推理等核心能力的综合性任务。这为评估和推动具身智能体在真实城市环境中的综合能力提供了一个必要且具有挑战性的测试平台。
Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
[131] LogiShot: Logically Coherent Cross-Shot Video Generation cs.CVPDF
Shuai Guo, Yuhang Yang, Zeyu Zhang, Pengfei Yu, Wei Zhai
TL;DR: 本文提出了LogiShot,一种用于生成逻辑连贯的跨镜头视频的方法。该方法通过联合编码上下文视频和其他条件信号来提供视觉语义线索,并维护视觉记忆以保持跨镜头视觉一致性。作者还构建了一个包含11万个样本的数据集和专用基准来评估跨镜头逻辑连贯性。
Details
Motivation: 当前跨镜头视频生成工作流(如短剧制作)通常依赖孤立的文本脚本或显式参考图像,当用户指令不明确时,生成的片段可能单独看合理但与整体叙事脱节,导致内容不连贯。因此,需要解决跨镜头逻辑连贯性和视觉一致性的问题。
Result: 实验表明,LogiShot在多个镜头的逻辑连贯性方面持续优于现有基线模型。作者构建的专用基准用于评估该性能。
Insight: 创新点在于通过两条互补路径整合信息:联合编码提供密集多模态线索,以及维护视觉记忆以保持一致性。客观来看,构建大规模数据集和专用评估基准也是推动该领域发展的重要贡献。
Abstract: Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.
[132] Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding cs.CVPDF
Donghui Feng, Fengxi Zhang, Changsheng Gao, Wenhan Yang, Qi Wang
TL;DR: 本文提出了一种名为视觉令牌编解码器(VTC)的双路径学习编解码器,用于高效压缩视觉Transformer(ViT)的中间令牌特征。该方法通过分离全局令牌和补丁令牌,并利用补丁令牌在原始网格上的空间相关性进行编码,显著提升了压缩效率。
Details
Motivation: 大型视觉基础模型的分布式部署需要分割ViT主干网络并在计算节点间交换中间令牌特征,在带宽和计算受限的情况下,高效的特征压缩变得至关重要。现有方法将异构的全局和补丁令牌展平为伪图像,忽略了补丁令牌固有的二维网格结构,导致空间冗余未被充分利用。
Result: 在DINOv2和SAM3模型上的实验表明,VTC在分类、分割和检测任务上均优于代表性的ViT特征编码基线方法。在保持90%未压缩特征性能的条件下,VTC在这些任务上将比特率降低了15.7倍至37.4倍。
Insight: 创新点在于提出了一种双路径编解码架构,将全局令牌和补丁令牌分离处理,并针对补丁令牌引入了空间-通道上下文熵模型以利用其二维空间相关性。此外,VTC还集成了后续ViT块的特征匹配监督和单编解码器内的可变速率模块,以支持中间层压缩和实际速率自适应。
Abstract: Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture sequence-axis dependencies while overlooking the native two-dimensional patch-grid structure. In this paper, we show that ViT patch tokens retain strong local spatial correlations on the original grid. To exploit this structural prior, we propose the Visual Token Codec (VTC), a dual-path learned codec that separates global and patch tokens into dedicated coding paths. Global tokens are compressed with a lightweight factorized prior, whereas patch tokens are encoded on the patch-token grid using a spatial-channel context entropy model. To support intermediate-layer compression and practical rate adaptation, VTC further incorporates feature-matching supervision after subsequent ViT blocks and variable-rate modules within a single codec. Experiments on DINOv2 and SAM3 show that VTC consistently outperforms representative ViT feature coding baselines on classification, segmentation, and detection tasks. At 90% of uncompressed-feature performance, VTC reduces bitrate by 15.7x-37.4x across these tasks. We further provide intermediate-layer rate-utility analyses for practical transmission- and storage-oriented deployment scenarios.
[133] SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport cs.CVPDF
Cigdem Beyan, Tonje Knutsen Sordalen, Kim Tallaksen Halvorsen
TL;DR: 本文提出了一种名为SLAP的选择性局部视觉-语言对齐框架,用于鱼类个体重识别任务。该方法通过部分最优传输(POT)建立视觉图像块嵌入与多个身份感知提示嵌入之间的局部对应关系,从而聚焦于最具判别性的身体区域,而非进行全局对齐。实验在Symphodus melops等数据集上验证了其在闭集和开集评估协议下均优于现有基于CLIP的重识别方法。
Details
Motivation: 现有基于CLIP的鱼类重识别方法主要依赖全局图像-文本对齐,导致背景和弱判别性区域对跨模态监督产生干扰。而个体识别的关键线索往往集中在特定局部身体区域,因此需要一种能够选择性关注这些判别性区域的对齐方法。
Result: 在纵向Symphodus melops数据集上的实验表明,该方法在闭集和开集评估协议下均持续优于近期的CLIP-based ReID方法。在其他数据集上的额外评估进一步证明了该方法在不同海洋生物重识别基准上的泛化能力。
Insight: 创新点在于引入了部分最优传输(POT)来实现视觉块与文本提示之间的选择性局部对齐,避免了弱匹配区域的强制对齐,从而学习到更具判别性的视觉表示。从客观角度看,将最优传输理论中的部分匹配思想应用于细粒度跨模态对齐,为解决局部判别性特征学习问题提供了新思路。
Abstract: Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific body regions rather than distributed uniformly across the animal. Nevertheless, recent CLIP-based ReID methods rely predominantly on global image-text alignment, allowing background and weakly discriminative regions to contribute to cross-modal supervision. We propose a selective local vision-language alignment framework that establishes localized correspondences between visual patch embeddings and multiple identity-aware prompt embeddings through Partial Optimal Transport (POT). Rather than enforcing exhaustive correspondence, POT enables selective matching between visual patches and prompt embeddings, allowing the model to emphasize the strongest cross-modal correspondences while avoiding forced alignment of weakly matching regions, thereby yielding more discriminative visual representations for retrieval. The framework is trained end-to-end, while only the adapted visual encoder is retained during inference. Experiments on the longitudinal Symphodus melops dataset demonstrate consistent improvements over recent CLIP-based ReID methods under both closed-set and open-set evaluation protocols. Additional evaluations on other datasets further demonstrate the generalization capability of the proposed method across diverse marine ReID benchmarks.
[134] Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3) cs.CVPDF
Nakul Poudel, Richard Simon, Cristian A. Linte
TL;DR: 本文提出了一种无需掩码标注的手术器械分割方法,通过结合Segment Anything Model 3 (SAM3)和微调后的视觉语言模型Qwen,实现从内窥镜图像中自动进行实例级分割。该方法使用通用文本提示“tool”生成二值掩码,再通过分类扩展为实例分割,在EndoVis数据集上验证了其有效性。
Details
Motivation: 现有手术器械分割方法依赖像素级标注或手动空间提示,限制了可扩展性和自动化;SAM3虽支持基于文本提示的免标注分割,但器械名称作为提示存在领域差距,无法直接使用。
Result: 在EndoVis 2017和2018数据集上评估,该方法虽未达到全监督方法的性能,但显著优于直接使用SAM3进行文本提示的实例级分割。
Insight: 创新点在于提出两阶段框架:利用SAM3的零样本能力生成二值掩码,再结合微调视觉语言模型进行实例分类,为免标注手术器械分割提供了可行方向;客观分析显示,该方法通过通用提示和后续分类缓解了领域差距问题。
Abstract: Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or manual spatial prompts, which limit scalability and automation. The recently introduced Segment Anything Model 3 (SAM3) offers a pathway to annotation-free, automatic segmentation via text-based prompting; however, the instrument name as a text prompt could not be directly used due to a large domain gap. To overcome these limitations, we propose a two-stage framework that achieves instance-level segmentation without requiring ground truth masks or manual interaction. In the first stage, we leverage a natural-language-aligned generic prompt - “tool” - to produce binary masks using SAM3’s zero-shot capability. In the second stage, these masks are extended to instance-level by integrating a vision-language model (Qwen) that is fine-tuned on SAM3-generated masked regions for instrument classification. We evaluate our approach on the EndoVis 2017 and 2018 datasets. Results show that, while our two-stage approach does not reach the performance of current fully supervised methods, it significantly outperforms the direct use of SAM3 for instance-level instrument segmentation with text prompts. Overall, our findings highlight both the limitations and potential of SAM3, suggesting a promising direction toward annotation-free surgical instrument segmentation.
[135] Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline cs.CVPDF
Dipit Saha, Shah Mohammad Abdul Mannan, Mohammad Raihan Rashid, Ruwad Naswan, Ahnaf Tahmid
TL;DR: 本文提出了一种无需训练、从粗到细的两阶段视觉语言模型(VLM)跟踪流程,用于零样本交通事故事件检测。该方法首先通过稀疏采样定位碰撞时刻,然后在精确时间窗口内结合目标检测与跟踪信息,利用冻结的Qwen3-VL-32B-Instruct模型进行细粒度分析,以结构化方式输出事故的时间、位置和类型。
Details
Motivation: 解决大规模交通监控视频中,将原始CCTV录像自动转换为结构化事故记录(包括时间、地点和事故类型)的难题,特别是在无法获得真实世界标注训练数据的严格约束下。
Result: 在ACCIDENT @ CVPR基准的2,027个真实CCTV测试片段上,该方法取得了0.504的三路调和平均分数,比组织者发布的最佳多模型集成基线(0.412)相对提升了22%,达到了新的SOTA水平。
Insight: 创新点在于提出了一种无需训练的两阶段粗到细流程,将冻结的VLM与目标检测(YOLO11x)和跟踪(BoT-SORT)相结合,并通过在第二遍分析中注入稳定的车辆身份和归一化边界框坐标作为视觉覆盖和显式数字描述,增强了模型对动态场景的理解能力。
Abstract: Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, and what type of collision occurred remains unsolved at scale. The ACCIDENT @ CVPR benchmark evaluates exactly this joint prediction under a strict constraint: no labeled real-world training data is available. We introduce a training-free, two-pass coarse-to-fine pipeline that pairs a frozen Qwen3-VL-32B-Instruct vision-language model with YOLO11x object detection and BoT-SORT tracking. A first pass sparsely samples the full clip to anchor the collision moment in time; a second pass re-examines a tight window around that estimate using frames annotated with stable vehicle identities and normalized bounding-box coordinates, which gives the model both a visual overlay and an explicit numeric description of the same scene. On the official 2,027-clip real-CCTV test set, our system achieves a three-way harmonic mean score of 0.504, surpassing all organizer-published baselines including the best multi-model ensemble (0.412) by a 22% relative margin.
[136] Sparse Attention to Emotion: Efficient Facial Emotion Recognition via Token Reduction cs.CV | cs.LGPDF
Aya Manel Zitouni, Aicha Zenakhri, Karim Haroun, Larbi Boubchir
TL;DR: 本文提出了一种名为稀疏情感注意力(SAE)的轻量化面部情感识别模型,通过丢弃对情感上下文无价值的图像令牌,在保持高精度的同时显著降低了计算成本。该模型在RAF-DB数据集上达到了新的最先进水平,并将计算复杂度降低了高达90%。
Details
Motivation: 当前基于视觉Transformer的面部情感识别方法具有二次复杂度,难以在边缘设备部署;作者假设情感识别无需全部面部信息,特定区域(如眼睛、嘴巴)已足够,因此旨在设计高效模型。
Result: 在RAF-DB数据集上,SAE取得了新的最先进结果,即使丢弃90%的图像令牌,仍能以更低成本达到与现有方法相当的竞争性精度。
Insight: 创新点在于利用情感识别的任务特性,通过稀疏注意力机制选择性保留关键面部区域令牌,实现了计算效率与精度的平衡,为轻量化模型设计提供了新思路。
Abstract: Facial Emotion Recognition (FER) is an important task that has significant implications across various fields such as biometrics, health, and human-computer interaction. Current Vision Transformer-based approaches display quadratic complexity $\mathcal{O}(N^2)$, with N being the input sequence length, making them cumbersome to deploy at the edge. In this paper, we hypothesize that the FER task does not necessarily require all facial information to correctly interpret emotional states, as specific regions such as the eyes, the mouth, and parts of the cheeks carry discriminative information that can be sufficient to recognize emotions. Based on this, we propose Sparse Attention to Emotion (SAE), a model that discards image tokens that have no added value to the emotional context, while preserving good accuracy and achieving a significant gain in computational cost. Surprisingly, even after suppressing 90% of the image tokens, our model achieves competitive accuracy to state of the art methods at much lower cost, providing a lightweight Facial Emotion Recognition approach. Experimental results demonstrate that SAE achieves new state of the art results on the RAF-DB dataset while reducing the computational complexity by up to 90%.
[137] City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning cs.CV | cs.CYPDF
Hanan Syed Shabir, Noor Fatima, Safia Baloch, Masroor Hussain
TL;DR: 本文提出了City Sentinel,一个基于人工智能的统一智能监控框架,用于实时多威胁检测。该系统集成了人脸识别、车牌识别、火灾烟雾检测、武器刀具检测、暴力检测和交通事故检测六种功能于一个可扩展平台,采用模块化开源架构,结合了Next.js仪表盘、FastAPI后端、云端PostgreSQL存储以及InsightFace、YOLOv8和EasyOCR等视觉模型。
Details
Motivation: 快速城市化增加了对能够同时监控多种公共安全风险的监控系统的需求,而传统系统通常针对不同功能使用独立解决方案,导致基础设施碎片化和操作界面复杂。
Result: 在配备NVIDIA RTX 3060 GPU的工作站上,系统实现了743毫秒的中位端到端延迟,并在两秒延迟限制内支持四个并发RTSP流;具体性能指标包括91.2%的人脸匹配率、85.7%的车牌读取准确率,以及在火灾、刀具和武器检测模块上mAP@0.5得分介于0.846至0.889之间。
Insight: 创新点在于将多种检测能力集成到一个统一的模块化架构中,实现了广泛的监控覆盖、云端可审计性以及灵活添加新检测功能的能力,同时保持了实用的实时性能;其开源多模型设计为构建可扩展的智能监控系统提供了参考范例。
Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems often use separate solutions for facial recognition, vehicle identification, fire detection, and behavioral analysis, resulting in fragmented infrastructure and multiple interfaces for operators to manage. This paper presents City Sentinel, a unified AI-based surveillance framework that integrates six detection capabilities into one scalable platform: facial recognition, automatic number plate recognition (ANPR), fire and smoke detection, weapon and knife detection, violence detection, and road accident detection. The system combines a Next.js operator dashboard, FastAPI backend, cloud-based PostgreSQL event storage, InsightFace and YOLOv8 vision models, and EasyOCR for plate recognition. Camera streams are processed through dedicated inference workers using RTSP. On a workstation equipped with an NVIDIA RTX 3060 GPU, the system achieves a median end-to-end latency of 743 ms and supports four concurrent RTSP streams within a two-second latency limit. It achieves a 91.2% face-match rate, 85.7% plate-reading accuracy, and mAP@0.5 scores of 0.846 to 0.889 across the fire, knife, and weapon detection modules. In user-acceptance testing, operators could enroll a new identity in under one minute and identify a flagged person from live footage in an average of 12 seconds. The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.
[138] ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision cs.CV | cs.AIPDF
Delin Mao, Chenghao Sun, Jingwei Song, Chishui Chen, Linfeng Zhang
TL;DR: ToolVision提出了一种新的监督学习方法,通过能力对齐的监督来解决多模态模型在视觉工具调用中的监督错位问题。该方法在SFT阶段通过多智能体管道和委员会评分机制筛选有效轨迹,在RL阶段基于工具使用收益进行奖励,从而同时学习何时以及如何使用视觉工具。
Details
Motivation: 现有SFT-then-RL方法在视觉工具调用中存在监督错位:SFT阶段学生模型可能模仿工具调用模式而未真正学会有效使用;RL阶段仅基于结果的奖励会抑制工具使用或鼓励无效操作。
Result: ToolVision-8B在全部七个主要基准测试上超越其基础模型,在三个高分辨率基准测试上超过Thyme-7B、CodeVision-8B和CodeDance-7B,并在V*和HRBench 8K上优于Qwen3-VL-32B-Thinking。
Insight: 创新点包括:SFT阶段的多智能体轨迹探索与委员会评分机制确保能力对齐;RL阶段基于工具使用收益的差异化奖励设计;整个流程无需额外人工标注即可自动构建监督信号。
Abstract: Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how to use tools, but trajectories from stronger teachers may succeed through perceptual capabilities that a smaller student cannot reliably reproduce or exploit, causing the student to imitate tool-call patterns without learning how to make them useful. RL is expected to teach when to use tools, but outcome-only rewards make fallible tool execution a liability and suppress tool use, whereas a blanket bonus for every correct tool-using trajectory encourages valid but ineffective operations. To address these two misalignments, we introduce ToolVision. During SFT, a multi-agent pipeline explores candidate trajectories, and a committee including student-scale models scores stepwise evidence gain to rank and prune the search branches. Only successfully executed trajectories with correct final answers are retained for SFT. Before RL, ToolVision compares the learner’s performance with and without tools, then rewards successful tool use only on questions where tools provide a clear benefit. Both signals are constructed automatically from public task data without additional human annotations of tool use or necessity. ToolVision-8B improves over its base on all seven main benchmarks, surpasses Thyme-7B, CodeVision-8B, and CodeDance-7B on all three high-resolution benchmarks, and outperforms Qwen3-VL-32B-Thinking on V* and HRBench 8K. We will publicly release the datasets and source code.
[139] From Recovery to Drop-off: How Action Post-training Reduces a VLM’s Late-Layer Depth Decodability cs.CV | cs.AI | cs.LGPDF
Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
TL;DR: 本文研究了视觉语言模型(VLM)在通过动作后训练转变为视觉语言动作模型(VLA)后,其空间理解能力(特别是深度感知)的变化。研究发现,VLA在所有解码器层上的深度解码能力均下降(称为’地板’效应),并且在最后几层出现更严重的崩溃(称为’悬崖’效应。通过因果定位,作者发现’悬崖’效应主要由后期MLP层的干扰导致,而注意力机制的影响较小。
Details
Motivation: 动机是探究VLM在进行动作后训练以构建VLA的过程中,其核心的空间几何理解能力(以深度感知为代表)保留了多少,以及这种能力是如何变化的。
Result: 在开源基础模型对Molmo2-ER(VLM)和MolmoAct2-LIBERO(VLA)上的实验结果表明,VLA在所有层的深度解码性能均低于基础VLM(地板效应),且VLA在最后几层的性能出现急剧下降(悬崖效应),而基础VLM的深度解码能力在最后几层是持续改善的。
Insight: 创新点在于揭示了动作后训练对VLM空间理解能力的非均匀损害模式(地板与悬崖效应),并通过因果干预(MLP消融)将性能崩溃定位到后期MLP层的干扰。这为理解VLA训练中知识遗忘的机制提供了新视角,并暗示了针对特定模块进行干预以保留核心能力的可能性。
Abstract: How much of a vision-language model’s (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM’s depth decodability improves through its final layers, the VLA’s collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.
[140] Damage Classification for 3D Point Cloud Data via 3D Data Analysis and Vision Foundation Model-based 2D Projections cs.CVPDF
Evan Perez, Kalelo Dukuray, Erika Ardiles-Cruz, Jie Wei
TL;DR: 本文研究了两种基于3D点云数据的细粒度损伤分类方法:基于3D点云损伤评估(3PDA)算法和基于2D投影损伤评估(2PDA)算法。3PDA利用拓扑数据分析(TDA)从PointNet分割的组件中提取紧凑的几何特征,并结合异常检测算法量化结构退化;2PDA则将3D点云投影为2D视图,利用视觉基础模型(VFM)进行损伤检测。
Details
Motivation: 解决3D点云数据细粒度损伤分类面临的高计算成本和标注数据稀缺的挑战,探索更高效、更通用的损伤评估方法。
Result: 在细粒度损伤分类任务上,3PDA在特定物体几何形状上能达到更高准确率,但计算成本高;2PDA使用VFM的2D投影方法实现了有竞争力的分类性能,其时间复杂度降低了一个数量级,且泛化到更广泛的物体类别。
Insight: 创新点在于将TDA用于3D点云特征压缩以进行异常检测,以及利用2D投影和预训练VFM实现轻量高效的损伤分类。客观来看,2PDA方案为资源受限场景提供了一种有前景的替代方案,平衡了精度与效率。
Abstract: Fine-grained damage classification of 3D point cloud data (PCD) remains a persistent challenge, constrained by high computational demands and limited labeled data. This study examines two methods: 3D PCD-based damage assessment (3PDA) algorithm and 2D projection damage assessment (2PDA) In our 3PDA analysis algorithm, TDA is used to derive compact representations of 3D PCD segmented by pointNet, which are then integrated with anomaly detection algorithms to quantify structural degradation. We show that TDA effectively compresses geometric structure from VFM-segmented components into discriminative feature vectors and that anomaly detection models can reliably distinguish components with varying damage severity using only 3D PCD inputs. In the 2D projection analysis algorithm, we leverage large VFMs for granular damage detection by projecting 3D PCD into 2D views. These projections allow VFM based models to achieve competitive classification performance while requiring only a fraction of the computational cost associated with full 3D data processing. Our results demonstrate that 2D VFM pipelines in 2PDA can perform strongly on fine-grained damage classification tasks, highlighting their viability as lightweight, resource-efficient alternatives to traditional 3PDA architectures. Comparative evaluation shows that the 3PDA attains higher accuracy but only for a narrow subset of object geometries and at substantially higher computational cost due to its reliance on TDA and the scarcity of high-fidelity 3D datasets. In contrast, the 2PDA algorithm yields slightly lower accuracy but offers an order of magnitude reduction in time complexity and generalizes across a far broader range of object categories.
[141] EndoMD-SLAM: Endoscopic Gaussian Splatting SLAM under Optical Degradation with Memory and Static-Transient Decomposition cs.CVPDF
Nuo Chen, Kangqi Ni, Lulin Liu, Joga Ivatury, Ying Ding
TL;DR: 本文提出了EndoMD-SLAM,一种用于内窥镜场景的Gaussian Splatting SLAM框架,旨在解决手术过程中因移动碎屑、冲洗等光学退化现象导致的系统失效问题。该框架通过记忆驱动的门控机制暂停不可靠观测的更新,并利用历史关键帧进行重定位,同时通过自监督的静态-瞬态分解将视觉污染物分离到独立的瞬态场中,从而保护持久解剖地图的几何完整性。
Details
Motivation: 现有基于Gaussian Splatting的SLAM系统依赖严格的多视角光度一致性假设,但在常规内窥镜手术中,间歇性的光学退化(如移动碎屑、冲洗)会严重破坏这一假设,导致系统错误地将相机附着伪影融合到持久3D几何中,造成严重的跟踪漂移和地图损坏。
Result: 在从结肠镜视频构建的专注于退化场景的基准测试上,广泛实验表明,在严重光学退化下,标准基线方法失效,而EndoMD-SLAM保持了几何完整性,将绝对轨迹误差降低了91%,并将渲染保真度提高了9.9 dB PSNR。
Insight: 论文的核心创新在于针对特定退化场景设计了专门的跟踪(记忆驱动门控与漂移感知重定位)与建图(自监督静态-瞬态分解)机制,将瞬态伪影与静态解剖结构显式分离,这为在动态干扰严重的视觉SLAM应用中维持鲁棒性提供了新思路。
Abstract: Dense 3D reconstruction is critical for clinical endoscopic navigation and documentation. While Gaussian Splatting SLAM systems show promise in this domain, they fundamentally rely on strict multi-view photometric consistency. In routine procedures, this assumption is severely violated by intermittent optical degradations like moving debris and water flushing. Standard systems erroneously fuse these cameraattached artifacts into the persistent 3D geometry, causing severe tracking drift and irreversible map corruption. To address this limitation, we propose EndoMD-SLAM, a framework designed to maintain stability under optical degradation through specialized tracking and mapping mechanisms. On the tracking side, a memory-driven gating mechanism detects unreliable observations to suspend map updates and utilizes historical keyframes for drift-aware relocalization. On the mapping side, a self-supervised static-transient decomposition isolates visual contaminants into a dedicated transient field. This explicit separation prevents artifacts from structurally entangling with the persistent anatomical map. We curate a degradationfocused benchmark from colonoscopy videos to systematically evaluate these failure modes. Extensive experiments show that while standard baselines fail under severe optical degradation, EndoMD-SLAM preserves geometric integrity, reducing absolute trajectory error by 91% and improving rendering fidelity by 9.9 dB PSNR.
[142] RMR-Net: Degradation-Evidence-Guided Road-Image Restoration for Defect Detection cs.CVPDF
Amir Ghorbani, Amirali K. Gostar, WeiQin Chuah, Vahid Ghorbani, Aidan Blair
TL;DR: 本文提出RMR-Net,一种用于道路缺陷检测的紧凑型任务感知图像恢复前端网络。它通过估计图像退化证据,融合现有退化参数,并利用有界残差路径恢复高频路面细节,旨在提升车载摄像头在运动模糊、失焦等退化条件下对道路裂缝和坑洞的检测性能。
Details
Motivation: 车载道路摄像头易受运动模糊、失焦、光照不良和噪声等退化因素影响,这些因素会抹去道路缺陷检测器所需的细小裂缝和坑洞边界,因此需要一种针对检测任务优化的图像恢复方法。
Result: 在IVCNZ坑洞数据集和PCM道路损伤数据集上,使用合成退化参数进行实验,RMR-Net在八种保留的退化条件中的七种上取得了最高的mAP50,例如在IVCNZ运动模糊上达到0.140-0.427,在PCM失焦上达到0.060-0.233。
Insight: 创新点在于提出了一种退化证据引导的、任务感知的恢复网络架构,其核心设计包括有界细节残差路径、退化条件化模块以及检测器感知的稳定性约束,这些组件共同作用以有效恢复对检测至关重要的高频细节。
Abstract: Vehicle-mounted road cameras are vulnerable to motion blur, defocus, poor illumination, and noise, which can erase thin cracks and pothole boundaries needed by road defect detectors. This paper presents RMR-Net, a compact task-aware restoration front end that estimates degradation evidence from the image, optionally fuses it with existing corruption context/parameters, conditions lightweight restoration blocks, and returns high-frequency pavement detail through a bounded residual path. The experimental scope is deliberately controlled: the conditioning information used on the Image and Vision Computing New Zealand (IVCNZ) pothole dataset and the Road Damage Dataset: Potholes, Cracks and Manholes (PCM) consists of saved synthetic-generator parameters, not measured vehicle telemetry. A clean-trained, frozen YOLO11s detector evaluates every image source. Across eight held-out degradation conditions, RMR-Net obtains the highest mAP50 in seven, including 0.140-0.427 for IVCNZ motion blur and 0.060-0.233 for PCM defocus. A compact ablation identifies the bounded detail path as the largest local contributor, while degradation conditioning and detector-aware stability terms provide complementary guidance.
[143] Fourier Self-Supervision for Fine-Grained Generalized Category Discovery cs.CV | cs.AI | cs.LGPDF
Sarah Rastegar, Mina Ghadimi Atigh, Pascal Mettes, Yuki M. Asano, Cees G. M. Snoek
TL;DR: 本文提出了一种名为傅里叶自监督的新方法,用于解决细粒度广义类别发现问题。该方法利用图像傅里叶变换,通过双频滤波策略(低通滤波提取抽象类别属性,高通滤波强调边缘纹理等细节)来增强模型对细微差异的辨别能力,从而在未标注数据中识别已知类别并发现新类别。
Details
Motivation: 现有基于自监督和对比学习的方法在细粒度广义类别发现中,往往依赖表面视觉线索而非人类分类所用的内在属性,难以捕捉细粒度差异。本文旨在解决这一问题。
Result: 在多个细粒度数据集上的实验表明,该方法超越了现有最先进方法,即使在类别数量未知的情况下也有效,证明了其在广义类别发现任务中的有效性。
Insight: 核心创新点在于将傅里叶变换与自监督学习结合,并设计了分离的低通与高通滤波双分支架构,分别处理抽象语义和细节特征,通过融合两者的表征来构建更丰富、完整的特征空间,从而提升细粒度识别和新类别发现的性能。
Abstract: Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typically based on self-supervision and contrastive learning, often struggle to capture fine-grained distinctions, relying on superficial visual cues rather than the intrinsic attributes humans use for categorization. We introduce Fourier Self-Supervision, that leverages the Fourier transform of images to enhance the discrimination of subtle differences and support the discovery of new categories. Our method employs a dual frequency filtering strategy: a low-pass filter first extracts broad, abstract attributes that capture high-level category information, while a high-pass filter emphasizes fine details such as edges and textures that are essential for fine-grained recognition. Each operates on a dedicated latent space, and their overlapping representations together yield a richer, more complete feature space. This dual-frequency approach not only refines feature extraction to identify novel categories, but also strengthens the model’s discriminative power in fine-grained category discovery. Experiments on multiple fine-grained datasets show that incorporating Fourier Self-Supervision outperforms state-of-the-art methods, even when the number of classes is unknown, demonstrating its effectiveness for Generalized Category Discovery. Our code is available at: https://github.com/SarahRastegar/FourEx.
[144] Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation cs.CV | cs.AIPDF
Xuanyu Liu, Zheng Fang, Hongyang He, Yundi Hong, Daizong Liu
TL;DR: 本文提出TriNoL框架,用于解决半监督视觉基础模型(VFM)适应中伪标签噪声敏感的问题。该方法将未标记样本根据置信度划分为三个区域,并分配给三个专门的LoRA专家(Positive、Alignment、Negative Expert)进行学习,从而在保持训练成本低的同时提升对噪声监督的鲁棒性。
Details
Motivation: 现有半监督VFM适应方法通常冻结预训练主干并更新轻量模块(如LoRA),但伪标签可靠性混杂,单一LoRA适配器需在同一低秩空间中吸收可靠、模糊和噪声梯度,导致模型对伪标签噪声敏感。
Result: 论文未在摘要中提及具体定量结果或基准测试,但宣称TriNoL通过将不同可靠性的伪标签区域分离到专门的适应路径中,提高了对噪声监督的鲁棒性。
Insight: 创新点在于提出三重专家学习框架,根据伪标签置信度路由样本至不同LoRA专家,实现噪声样本的专门处理;客观分析认为,这种基于置信度的专家分工策略可有效隔离噪声影响,是轻量级适应中提升鲁棒性的可行思路。
Abstract: Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. However, pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambiguous, and noisy gradients in the same low-rank space. This can make VFM adaptation sensitive to pseudo-label noise. We propose \textbf{TriNoL}, a \textbf{Tri}ple-expert learning framework from \textbf{No}isy \textbf{L}abels for semi-supervised VFM adaptation. TriNoL routes unlabeled samples into three confidence regions and assigns them to three LoRA experts: a Positive Expert for high-confidence pseudo-labels, an Alignment Expert for medium-confidence ambiguous samples, and a Negative Expert for low-confidence noisy samples. The VFM backbone remains frozen, and only the LoRA experts and classifier head are updated. By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low.
[145] Learning human joint torques from pixels cs.CVPDF
Chen Chen, Rui Cheng
TL;DR: 该论文提出了VID数据集和基准,用于从单目RGB图像直接预测人体关节扭矩,并设计了VID-Network模型。VID包含超过6.3万帧同步的真实人体图像、运动学标注和生物力学标签,为基于视觉的逆动力学研究提供了首个实用基准。实验表明,VID-Network在整体扭矩估计、关节特定分析和动作特定预测上均优于基线方法。
Details
Motivation: 现有扭矩估计方法依赖表面肌电图、运动捕捉标记或力板等设备,限制了其在普通RGB图像上的应用。本研究旨在将生物力学分析从受控实验室扩展到真实世界场景,直接从视觉观测中估计人体关节扭矩。
Result: 在VID数据集上的实验显示,VID-Network实现了1.7612 N·m/kg的整体平均关节位置误差(mPJE),相比最佳基线提升了39.81%,并在所有评估关节类型和大多数动作类别中取得了最低误差。
Insight: 创新点在于构建了首个基于真实图像的视觉逆动力学数据集(VID)和标准化评估协议,并提出了结合姿态预训练空间概率特征、标记回归和时间扭矩推理的VID-Network模型,为在非受限环境中研究生物力学推断奠定了基础。
Abstract: Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world movement scenarios. Existing torque estimation methods typically depend on surface electromyography, motion-capture markers, force plates, or simulated imitation data, which limits their applicability to ordinary RGB images. In this work, we introduce VID, a vision-based inverse dynamics dataset and benchmark for predicting human joint torques directly from real monocular images. VID contains 63,369 synchronized frames with real human images, kinematic annotations, anthropometric attributes, and OpenSim-derived dynamic labels, providing paired visual and biomechanical supervision for real-image inverse dynamics. We further define a standardized evaluation protocol covering overall torque estimation, joint-specific analysis, and action-specific prediction. To establish a strong reference model, we propose VID-Network, which combines pose-pretrained spatial probabilistic features, marker regression, and temporal torque inference to recover joint torques from image sequences. Experiments on VID show that VID-Network achieves an overall mPJE of 1.7612 N$\cdot$m/kg, improving over the best compared baseline by 39.81%, and obtains the lowest error across all evaluated joint types and most action categories. VID establishes a first practical benchmark for vision-driven human inverse dynamics and provides a foundation for studying biomechanical inference in less constrained environments.
[146] SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision cs.CVPDF
Weixin Ye, Wei Wang, Hongguang Zhu, Xuecheng Nie
TL;DR: 该论文提出了一个名为SI-Edit的框架,用于解决基于草图的图像编辑中实现像素级精度的难题。首先,作者构建了一个高质量数据集SI-Data,该数据集通过多模态大语言模型自动生成包含原始图像、局部几何草图、语义指令和编辑后图像的四元组。在此基础上,SI-Edit框架将语义指令与精确的几何约束相结合,以实现用户意图驱动的局部精细编辑。
Details
Motivation: 当前生成模型在基于草图的图像编辑,尤其是细粒度局部变形方面,难以实现像素级精度。这主要源于缺乏同时提供几何约束和语义指令的高质量公开基准数据集。
Result: 实验结果表明,SI-Edit在基于草图的图像编辑任务上,比基线方法提供了更可靠的结构控制,并实现了与用户意图对齐的精确、像素级局部细化。作者还建立了一套全面的评估指标来衡量结构保真度和语义一致性。
Insight: 主要创新点在于构建了首个用于指令引导局部草图编辑的高质量数据集SI-Data,以及提出了一个整合语义指令与几何约束的协作框架SI-Edit。从客观角度看,利用MLLMs自动化构建数据集的方法和协作空间-语义学习的思路具有借鉴意义。
Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce SI-Data, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called SI-Edit that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the project page.
[147] Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation cs.CV | cs.LGPDF
Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang
TL;DR: 本文提出了一种名为对比掩码保真度(CMF)的无训练、无参考指标,用于直接评估遥感语义分割中候选类别掩码相对于图像证据的可靠性。CMF通过构建保留和擦除的反事实视图,利用冻结的视觉语言模型判断类别证据是否集中在掩码内部且外部缺失,从而揭示人工标注掩码的系统性失真问题。
Details
Motivation: 遥感语义分割模型的训练和评估依赖于人工绘制的掩码,但这些标注往往粗糙、不完整或未对齐,导致高重叠分数可能反映的是与不完美标签的一致性而非对图像的忠实度,从而产生评估悖论。
Result: 在受控掩码损坏验证中,CMF表现良好;在十个遥感基准测试的10,731个图像-类别对上审计显示,人造类别(如建筑、道路、汽车)在62-85%的案例中更支持候选掩码,而模糊土地覆盖则更倾向于人工标注。在盲审的三标注者共识中,CMF与专家判断的匹配率达到81%,优于仅保留评分、模型置信度和训练过的标签质量基线。
Insight: CMF的创新点在于提供了一种无需训练或参考标注的掩码质量评估方法,能够系统性地审计地面真值掩码的失真问题,并通过保守的类别仲裁提升跨域迁移性能,挑战了地面真值不可置疑的假设。
Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.
[148] Visual Distortion Detection in UGC Images Using Large Multimodal Models cs.CV | cs.AIPDF
Ziheng Jia, Yingji Liang, Jiaying Qian, Xiongkuo Min
TL;DR: 本文提出了一种名为VIGIL的方法,利用大型多模态模型(LMM)进行精确的视觉失真检测。通过构建包含超过14万张失真图像的VIGIL-140K训练集,并采用多层级特征同步检测及保留非失真类别中的失真线索等策略,有效解决了合成到真实(S2A)场景的泛化问题。
Details
Motivation: 现有基于LMM的图像质量评估方法主要依赖文本驱动的监督微调,在检测精度上存在局限,且使用合成失真图像作为主要训练数据会导致在真实场景中出现显著的泛化差距,即合成到真实(S2A)问题。
Result: 在领域内合成失真检测和S2A任务上,经过后处理的VIGIL模型均持续优于强基线方法,实现了更优的性能。
Insight: 创新点包括:构建大规模、高质量且覆盖8种主要合成失真类别的训练集VIGIL-140K;将LLM解码器的不同层视为多个检测器,利用多层级特征进行同步失真检测;保留分配给非失真类别的预测中的失真线索,以缓解S2A问题中常见的前景-背景模糊分离问题。
Abstract: The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tuning (SFT). However, this training paradigm exhibits notable limitations in detection accuracy. Moreover, synthetically distorted images, which are often used as the primary training data source, show a significant generalization gap when deployed in real-world scenarios; thus, the \textbf{synthetic-to-authentic (\textit{S2A})} problem represents a critical challenge. Motivated by these issues, we propose \textbf{\textit{VIGIL}}, which leverages the LMM architecture for precise visual distortion detection. From a candidate pool of over 1000K samples, we construct the \textbf{\textit{VIGIL-140K}} training set, which consists of over 140K distorted images. These images are obtained through rigorous quality filtering and carefully crafted distortion injection, covering 8 major synthetic distortion categories. Our model leverages different layers of the large language model (LLM) decoder, treating them as \textit{multiple detectors} that perform synchronous distortion detection using multi-level features. Additionally, we retain distortion cues from predictions assigned to the non-distortion class, which helps mitigate the ambiguous foreground-background (\textit{FG-BG}) separation commonly encountered in the \textit{S2A} problem. After post-processing, our model consistently outperforms strong baselines on both in-domain synthetic distortion detection and \textit{S2A} tasks.
[149] View-Adaptive Renderer for View-Consistent 2D-to-3D Generation cs.CV | cs.GRPDF
U-Chae Jun, Jaeeun Ko, Jiwoo Kang
TL;DR: 本文提出了一种新颖的视点自适应神经渲染框架,用于解决单图像3D生成中因投影模糊性导致的多视图不一致问题。该方法通过独立的视点自适应渲染器校正视图相关误差,并结合自注意力融合模块自适应地整合多视图信息,从而在保证几何一致性的同时,实现了高效且高保真的3D重建。
Details
Motivation: 传统的单目3D生成流程(先合成多视图,再基于NeRF重建)常因固有的投影模糊性导致生成视图间的视觉不连续,进而影响重建模型的准确性。现有解决方案要么计算成本高昂,要么未能充分解决合成视图间的实际不一致问题。
Result: 大量实验表明,该方法能持续提升3D重建的保真度。重要的是,该方法在不依赖基于扩散的SDS监督的情况下,仅使用光度渲染损失和轻量级注意力正则化器,就达到了接近最先进的性能水平。
Insight: 核心创新在于提出了视点自适应神经渲染器,能独立校正视点相关误差,同时共享全局特征主干以保持结构连贯性;以及自注意力融合模块,能自适应整合多视图信息,确保几何一致性,而无需严重依赖间接正则化或计算密集型方法。这实现了精度与效率的良好平衡。
Abstract: Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipelines typically synthesize multiple views from a single input image before applying Neural Radiance Field (NeRF)-based reconstruction. However, inherent projective ambiguities often produce visual discontinuities across generated viewpoints, leading to inaccuracies in reconstructed 3D models. Current solutions either incur significant additional computational burdens or fail to adequately resolve practical inconsistencies between synthesized views. To address these limitations, we propose a novel viewpoint-adaptive neural rendering framework that enables robust 3D reconstruction even when given partially inconsistent multi-view inputs. Our approach introduces view-adaptive neural renderers that independently correct viewpoint-dependent errors while simultaneously sharing a global feature backbone to preserve structural coherence. Furthermore, we propose a self-attention fusion module that adaptively integrates multi-view information, ensuring geometric consistency without relying heavily on indirect regularizations or computationally intensive methods. Through extensive experiments, we demonstrate that our method consistently improves 3D reconstruction fidelity. Importantly, our approach achieves near state-of-the-art performance without diffusion-based SDS supervision, relying primarily on photometric rendering loss with lightweight attention regularizers. This balance between accuracy and efficiency makes the proposed framework highly practical for real-world applications.
[150] Bright-Channel Retinex Enhancement with a Conditional Overdispered-Noise Analysis cs.CV | eess.IVPDF
Jongpil Jeong
TL;DR: 本文提出了一种无需训练的低光照图像增强方法,结合了局部亮通道光照估计、Retinex分解和边缘保持去噪。该方法通过条件负二项式伪计数模型分析Retinex除法放大的异方差噪声,并采用双边滤波进行后处理。在LOL-v1数据集上取得了传统方法中最高的PSNR/SSIM指标,并在Apple M2 Pro CPU上实现了约43 FPS的实时处理速度。
Details
Motivation: 解决低光照图像增强中噪声被Retinex除法放大导致细节损失的问题,同时避免基于学习的方法对训练数据的依赖和计算开销。
Result: 在LOL-v1数据集上获得平均PSNR 17.74dB和SSIM 0.739,在传统方法中达到最高水平;处理400x600图像时在Apple M2 Pro CPU上实现约43 FPS。
Insight: 创新点在于将条件负二项式噪声模型作为诊断工具(而非精确传感器模型)分析Retinex增强中的异方差噪声,并通过亮通道估计与最大似然反射率估计的结合实现无训练增强;其双边滤波的工程化近似在效率与效果间取得平衡。
Abstract: I present a training-free low-light enhancement method that combines local bright-channel illumination estimation, Retinex division, and edge-preserving denoising. For a fixed illumination estimate, a conditional Negative -Binominal psueduo-count method characterises the heteroscedastic noise amplified by division. The unconstrained reflectance ratio is the pixelwise maximum-likelihood estimate, with a boundary solution for zero-valued observations; the implemented estimate additionally applies illumination filtering and range clipping. The NB model is a diagnostic noise analysis rather than a calibrated sensor model, and the final fixed-bandwidth bilateral filter is an empirical approximation rather than the exact Bayesian solution. On the LOL-v1 dataset, the methodobtains mean PSNR/SSIM of 17.74dB/0.739, the highest values among the evaluated with conventional methods. A 400X600 image is processed at approximately 43 FPS on an Apple M2 Pro CPU.
[151] CodecArena: Codec Quality Assessment via Visual Reinforcement Learning cs.CVPDF
Jiaye Fu, Weiqi Li, Qiankun Gao, Yanchen Zhao, Xiandong Meng
TL;DR: 本文提出CodecArena,首个基于视觉-语言框架的视频编码质量评估方法,将编解码器评估建模为参考视频与其重建版本之间的源条件比较推理。通过Facet-GRPO视觉强化学习方案进行优化,结合五个保真度维度(身份、物体、文本、纹理、时间一致性)进行细粒度质量判断。
Details
Motivation: 现有主流指标(如LPIPS和DISTS)主要衡量特征和纹理相似性,而非内容保真度,可能导致重建视频中出现错误人脸或模糊文本等人类明显拒绝的情况仍获得高分。
Result: 在CodecArena-Bench基准测试中,CodecArena在多种编解码器和码率下,与人类对源无关内容的主观评价达到了最先进的一致性,超越了感知指标和先前的视觉-语言评估方法。
Insight: 创新点包括:将编解码器评估转化为源条件比较推理任务;提出Facet-GRPO强化学习方案,利用自动推导的维度方向作为弱锚点,避免单一子分数主导整体偏好,实现可解释的细粒度质量评估;构建了自动偏好数据集CodecArena-1K和人工排名基准CodecArena-Bench以支持训练和评估。
Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.
[152] Multimodal Model Diffing for Feature Discovery and Control cs.CV | cs.AI | cs.CL | cs.LGPDF
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt
TL;DR: 本文提出了MMDiff,一个多模态模型差异分析框架,用于在多模态大语言模型(MLLMs)中发现和控制内部特征。该框架通过训练多模态稀疏自编码器(SAEs),并将其转化为特征级接口,支持特征隔离、任务特定特征检测和特征级控制三种用途。
Details
Motivation: 多模态大语言模型虽然表现出强大的视觉理解能力,但其内部特征难以识别、审计或控制。现有的基于稀疏自编码器的后验解释方法既难以分离出多模态训练所改变的特征,也无法直接用于目标控制。
Result: 在LLaVA-MORE、PaliGemma 2和InternVL3.5三个MLLM家族上进行了评估。移除所发现的稀疏因果特征,在空间任务上平均选择性降低目标行为12%,在OCR任务上降低17%,并将多模态安全攻击成功率降低24%,且不影响VQA性能。对这些特征进行引导,相比标准的单层引导基线,在空间和OCR任务上的准确率平均分别提升了+3.6%和+1.8%。
Insight: 创新点在于提出了一个统一的框架(MMDiff),将多模态稀疏自编码器从单纯的可解释性工具,转变为可用于审计、引导和控制MLLM行为的机制。其核心是通过模型差异分析来隔离由多模态训练引入的特征变化,并结合对比性激活分析来识别因果特征,从而实现对特定行为的精准干预。
Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
[153] UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation cs.CVPDF
Yilei Hua, Beibei Jing, Ce Zheng, Hanyu Zhou, Yawei Luo
TL;DR: 本文提出了UniMoFlow框架,用于基于指令驱动的3D人体运动编辑。该方法通过数据、架构和推理三个层面进行创新:构建了大规模编辑数据集Omni-MoEdit,设计了统一的潜在流匹配模型以共享生成与编辑知识,并引入了源锚定流编辑(SAFE)推理策略。实验表明,该方法在目标文本对齐、编辑有效性和循环一致性方面均有提升,同时保持了良好的源保真度和运动生成质量。
Details
Motivation: 现有3D人体运动指令编辑方法存在局限性:基于生成模型的无训练适配方法控制效果欠佳,而依赖三元组监督的方法则受限于数据规模和语义多样性。本文旨在克服这一瓶颈,将运动编辑直接建立在文本到运动生成的基础上。
Result: 在广泛的实验中,该方法在目标文本对齐、编辑有效性和循环一致性方面表现出改进,同时保持了具有竞争力的源保真度和文本到运动生成质量。评估中引入了语义感知指标以考虑与单一真实参考存在固有偏差的有效编辑。
Insight: 主要创新点包括:1)构建大规模、多样化的编辑数据集Omni-MoEdit的闭环合成与验证流程;2)提出统一的潜在流匹配模型UniMoFlow,在生成和编辑任务间共享语义和运动学知识;3)设计了源锚定流编辑(SAFE)推理方法以实现可控的、锚定于源运动的细化编辑。从客观角度看,将编辑任务深度整合到生成框架中,并系统性解决数据、模型和评估瓶颈,是一个值得借鉴的系统性思路。
Abstract: Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.
[154] Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images cs.CVPDF
Yongsong Huang, Xiaofeng Liu, Tomo Miyazaki, Yaohou Fan, Shinichiro Omachi
TL;DR: 本文针对多模态大语言模型在红外图像上的应用,提出了一个解释感知的评估框架,以区分答案正确性、输出级解释的基于热成像证据的程度以及热成像基础。研究发现,即使答案正确,模型也可能依赖弱证据或可见光证据;移除原始红外图像会削弱热成像基础但几乎不影响准确率;且能力更强的模型在红外图像可用时表现更好。作者进一步提出了无需训练的热成像基础反馈方法,用于诊断解释侧失败并修正解释,同时保持答案不变。
Details
Motivation: 当前通用多模态大语言模型在红外图像上的评估通常仅基于答案准确性,但正确答案并不能保证模型的解释是基于红外热成像证据的,这可能导致模型缺乏可信度。
Result: 通过双LLM共识判断和人工校准检查,发现正确答案可能依赖弱或可见光证据;移除红外图像会显著削弱热成像基础但对准确率影响小;能力更强的模型在红外图像可用时热成像基础更好。提出的热成像基础反馈方法在本地配对输入验证中,在不改变答案的情况下改善了解释侧的基础性。
Insight: 创新点在于提出了一个解释感知的评估框架,强调评估模型解释的热成像基础性,而不仅仅是答案准确性;并提出了无需训练的热成像基础反馈方法,用于自动修正解释错误,提升模型在红外场景理解中的可信度。
Abstract: General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model’s explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.
[155] RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection cs.CVPDF
Zhihao Zhang, Gengwei Zhang, Tianlong Chen, Xiaoming Liu
TL;DR: 本文提出RefineAny3D,一种用于单目3D目标检测的深度优化方法。该方法将深度误差修正视为一个视觉对齐问题,而非数值回归问题,通过扩展视觉语言模型的词汇表并利用大规模思维链数据进行监督,实现了对封闭集检测器、开放词汇检测器和3D自动标注工具的通用后处理优化。
Details
Motivation: 现有深度基础模型虽然具有强大的零样本泛化能力,但缺乏3D检测所需的目标级精度,直接替换检测器预测的深度会降低准确性。因此,作者将目标级深度优化作为一个独立任务来处理。
Result: RefineAny3D作为单一后处理步骤,在封闭集检测器、开放词汇检测器和3D自动标注工具上均带来了稳定的性能提升,并且能够泛化到未见过的类别、场景和相机,无需重新训练。
Insight: 核心创新在于将深度误差修正重新定义为视觉对齐问题,通过引入动作令牌替代数值深度输出,并利用基于显式视觉证据的思维链数据进行监督,从而避免了直接的数值回归,提升了模型的泛化能力和解释性。
Abstract: Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector’s predicted depth degrades accuracy, even falling below the detector’s own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM’s vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.
[156] RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing cs.CVPDF
Hao Li, Ju Dai, Feng Zhou, Mengting Shi, Haofei Wang
TL;DR: 本文提出RAGMesh框架,结合FaME-G2E数据集,通过检索增强机制实现细粒度文本驱动的3D人脸生成与编辑。该方法利用多尺度检索融合模块整合全局与局部几何先验,并引入自适应检索引导监督增强区域对齐,显著提升了局部几何精度与编辑可控性。
Details
Motivation: 现有文本驱动3D人脸方法难以将长文本描述转化为细粒度几何细节(如眉毛张力、脸颊收缩等局部形变),导致几何保真度与编辑精度受限。
Result: 在FaME-G2E数据集上的实验表明,RAGMesh在局部几何精度、文本引导可控性、区域编辑精度和推理效率上均优于现有SOTA方法。
Insight: 创新点包括构建大规模多模态数据集FaME-G2E,以及提出多尺度检索融合模块与自适应检索引导监督机制,通过检索增强策略显式对齐文本语义与面部区域,提升细粒度编辑能力。
Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text–mesh annotations and paired text–blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.
[157] Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression cs.CV | cs.AIPDF
Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu
TL;DR: 本文提出了一种后果敏感的视觉令牌压缩方法,用于视觉-语言模型。该方法根据下游任务中不同错误的潜在成本差异,动态分配视觉计算资源,优先保障高后果问题的准确性,而非对所有令牌进行均匀压缩。
Details
Motivation: 现有视觉令牌压缩方法通常假设所有错误的代价相同,仅追求平均准确率最大化。然而,在实际下游任务中,不同预测错误的后果严重性可能极不对称,因此需要一种能根据错误成本差异来分配计算预算的压缩策略。
Result: 在受控的基准测试中,该方法将高风险问题的错误率从0.300降至0.133,而基于内容的分配器表现与均匀分配相当。在混合工作负载上,该方法将成本加权错误降低了38%,同时延迟比全分辨率推理降低了约21%。
Insight: 创新点在于将错误后果的差异性成本纳入视觉令牌压缩的决策过程,提出了“校准-分配”框架,并证明了当错误成本差异增大时,将令牌预算向高后果问题倾斜的分配原则能显著优化整体任务表现。该方法在多种基准、架构和实现机制上均展现出良好的泛化性。
Abstract: Visual token compression for vision–language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.
[158] NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams cs.CVPDF
Lifang Wu, Yuyang Wu, Yangdong Gao, Fengyu Liu, Ya Jing
TL;DR: 该论文提出了NBA_Streaming,一个用于在线细粒度篮球解说生成的大规模基准数据集。它包含307小时的篮球比赛视频和约3.5万个时间对齐的事件标注,旨在解决现有方法无法处理连续视频流、以及现有数据集在球员身份、细粒度动作和事件链等监督信息有限的问题。论文还提出了一个因果两阶段框架,以提升在线事件定位和基于事实的解说生成能力。
Details
Motivation: 现有篮球解说生成方法主要针对预分割片段或完整视频,不适合连续视频流。现有数据集在球员身份、细粒度动作、事件属性和连贯事件链方面提供的监督有限,限制了生成解说的信息丰富度和事实准确性。
Result: 在提出的NBA_Streaming基准上进行的大量实验表明,现有基线方法在在线时序、事实基础和细粒度描述方面存在困难。论文提出的因果两阶段框架相比现有强基线模型取得了持续改进,但仍有提升空间,凸显了该基准的价值。
Insight: 创新点在于从孤立片段转向连续视频流,构建了一个统一评估事件定位、响应可靠性、事实基础和因果约束下解说质量的基准。提出的框架结合了“完成优先”的事件定位和以球为中心的语义基础,能够从观察到的流中识别完整事件,并组织场景、事件、身份和动作线索来生成解说。
Abstract: Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. However, existing methods are primarily designed for pre-segmented clips or complete videos, making them unsuitable for continuous streams. Existing datasets also provide limited supervision for player identities, fine-grained actions, event attributes, and coherent event chains, restricting the factual richness of generated commentary. To address these limitations, we introduce NBA_Streaming, a large-scale benchmark for online fine-grained basketball commentary generation. It contains 307 hours of basketball broadcasts and approximately 35K temporally aligned events, with annotations of event boundaries, player identities, fine-grained actions, event chains, and natural-language commentary. By moving from isolated clips to continuous streams, NBA_Streaming enables unified evaluation of event localization, response reliability, factual grounding, and commentary quality under causal constraints. We further propose a causal two-stage framework that combines completion-first localization with ball-centric semantic grounding, enabling the system to identify complete events from observed streams and organize scene, event, identity, and action cues for commentary generation. Extensive experiments reveal the difficulty of NBA_Streaming, where existing baselines struggle with online timing, factual grounding, and fine-grained description. Our framework consistently improves over strong alternatives, while the remaining gap highlights NBA_Streaming as a valuable benchmark for streaming sports video understanding and generation.
[159] Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines cs.CV | cs.LGPDF
Yarin Udi, Tom Sharon-Shahak, Roee Masad, Dan Pri-Tal
TL;DR: 本文提出EndoClock框架,用于审计医学世界模型流水线中因多模态数据同步到固定速率网格而可能丢失任务相关证据的问题。通过四类分类法(采样值、网格单元更新模式、原生时序、外部采集通道)定位证据留存位置,并以保守预训练审计方式报告最低证据承载表示。
Details
Motivation: 解决医学多模态数据在固定网格同步预处理过程中,因内生时钟依赖潜在或采集状态而导致任务相关证据被擦除的问题,确保模型能获取完整信息。
Result: 在超声心动图案例中演示了失败场景:B模式视频输出在脉冲波多普勒采集期间停止,而相应测量事件仅记录于外部采集日志,突显了同步可能导致信息丢失的风险。
Insight: 创新点在于提出内生时钟概念和四类证据留存分类法,将理论框架转化为可执行的预训练审计工具,强调在任务设计前需保留原生观测过程以评估同步的信息影响。
Abstract: Medical world models commonly learn from multimodal recordings synchronized onto a fixed-rate grid. This preprocessing resamples each native stream onto a shared time axis. Each stream has an observation clock that governs when observations are emitted or updated. When this clock depends on the latent or acquisition state, it is endogenous. In such settings, synchronization may not be neutral and can erase task-relevant evidence before the model sees the data. We introduce a four-regime taxonomy that characterizes where the evidence needed to distinguish a target event or state survives. The relevant witness may remain in the sampled values, in grid-cell update patterns, in native timing, or only in an external acquisition channel. EndoClock operationalizes this taxonomy as a conservative pretraining audit. It reports the lowest witness-bearing representation supported by the available evidence, or unresolved when no regime can be established. We illustrate this failure in echocardiography, where B-mode video write-outs cease during pulsed-wave Doppler acquisition while the corresponding measurement events remain recorded only in an external acquisition log. This work is a preliminary failure alert and executable audit. Its practical message is to preserve the native observation process long enough to determine whether synchronization has erased information required by the intended task.
[160] PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection cs.CVPDF
Shengbo Qi, Hongyi Fang, Benjia Zhou, Rui Mao
TL;DR: 本文提出了PatchHead,一种用于提升AI生成图像检测器跨数据集泛化能力的轻量级空间聚合头。该方法基于冻结的DINO骨干网络,通过保留并整合DINO补丁令牌的空间二维结构信息,来捕捉图像中空间分布的生成痕迹,从而解决现有方法因过度依赖全局CLS令牌聚合而导致泛化性能不佳的问题。
Details
Motivation: 现有基于视觉基础模型(如DINO)的AI生成图像检测器,通常仅使用全局聚合的CLS令牌进行分类,这可能会掩盖图像中空间分布的生成痕迹,导致模型在训练和测试数据来自不同生成器或数据集时泛化能力差。
Result: 在涵盖人工策划和真实场景的九个跨数据集基准测试中,PatchHead在七个数据集上排名第一,在其余两个上排名第二。它将先前最强方法的平均平衡准确率从91.6%提升至94.6%(+3.0个百分点),并将最差情况准确率从82.4%提升至89.4%(+6.9个百分点),同时仅增加了8.6%的可训练参数和0.08%的额外FLOPs。
Insight: 核心创新在于设计了一个轻量级的空间聚合头(PatchHead),它保留了DINO补丁令牌的二维空间结构,并通过整合相邻区域的证据来学习空间分布的生成痕迹。这从表征层面解释了为何基于空间补丁聚合的方法比基于单一CLS的全局表征具有更强的跨生成器和数据集的迁移可靠性。方法还结合了冻结骨干、LoRA适配器等高效微调策略。
Abstract: AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spatial representations produced by vision foundation models like DINO, existing detectors typically classify images using only the globally aggregated CLS token. We hypothesize that globally aggregating DINO features into a single CLS token obscures spatially distributed generation traces. To test this hypothesis, we introduce PatchHead, a lightweight spatial aggregation head that preserves the two-dimensional organization of DINO patch tokens and integrates evidence across neighboring regions. During training, we freeze the pretrained DINO backbone and optimize only the inserted LoRA adapters, PatchHead, and auxiliary projection head. Across nine cross-dataset benchmarks spanning manually curated and in-the-wild settings, PatchHead ranks first on seven datasets and second on the remaining two. It improves the strongest prior method from 91.6% to 94.6% in average balanced accuracy (+3.0 points) and raises the worst-case accuracy from 82.4% to 89.4% (+6.9 points), while introducing only 8.6% more trainable parameters and 0.08% additional FLOPs. Further qualitative analysis suggests that PatchHead (i) reduces class-conditional domain discrepancy, and (ii) redirects the representation from content-dominated saliency toward spatially distributed authenticity evidence. Together, these observations provide a representation-level account of why spatial patch aggregation transfers more reliably across generators and datasets than a single CLS-based global representation. Our code and models will be made available upon acceptance.
[161] RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation cs.CV | cs.AIPDF
Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou
TL;DR: 本文提出REST(Reward-Enhanced Scored-Trajectory Distillation)框架,一种单阶段的强化学习与蒸馏协同训练方法,用于实现高效少步文本到图像生成。该方法利用强化学习过程中已评分的有限步轨迹作为蒸馏监督源,通过优势调制蒸馏(AMD)机制强化高奖励轨迹的监督并减弱低奖励行为的影响,从而在无需额外图像生成、独立蒸馏数据集或对抗训练的情况下,实现与多步强化学习教师模型性能相当甚至更优的少步无分类器引导推理。
Details
Motivation: 当前高效的文本到图像生成通常需要依次进行基于强化学习的奖励对齐和少步蒸馏,这种顺序执行方式增加了训练成本,且可能在压缩过程中损失奖励收益。本文旨在从强化学习原生视角出发,直接利用扩散强化学习过程中产生的已评分轨迹作为蒸馏监督,避免传统两阶段方法的缺陷。
Result: 在组合生成、视觉文本渲染和人类偏好对齐任务上的实验表明,REST能够实现少步无分类器引导推理,其性能匹配甚至超越其40步的强化学习教师模型,整体额外训练成本低于纯强化学习的25%。在DrawBench PickScore上,REST比RTDMD提升了0.82分,且仅需五分之一的训练迭代次数。
Insight: 核心创新在于将强化学习过程中产生的评分轨迹视为有价值的蒸馏监督源,而非采样副产品,并提出单阶段的RL-蒸馏协同训练框架REST。其优势调制蒸馏(AMD)机制通过基于优势值的符号权重动态调整监督强度,有效引导学生模型模仿高奖励行为并远离低奖励行为,实现了轻量级、即插即用的高效训练。
Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sampling. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher. The student learns segment-wise from the teacher’s evolving rollout trajectories while leaving the original teacher optimization unchanged. To prevent uniform imitation from preserving undesirable low-reward behaviors, we further introduce Advantage-Modulated Distillation (AMD), which transforms rollout advantages into signed weights over a base distillation loss. AMD strengthens supervision from preferred trajectories and mildly repels the student from low-reward ones. The resulting framework is lightweight and plug-and-play, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment show that REST enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost below 25% over pure RL. REST improves DrawBench PickScore over RTDMD by 0.82 while requiring only one-fifth of the training iterations.
[162] GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views cs.CV | cs.AI | cs.IR | cs.MMPDF
Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang
TL;DR: 本文提出GRASP框架,以解决无人机视角下细粒度跨模态理解的双重挑战:宏观层面的跨模态焦点错位(背景干扰导致模型关注全局环境而非具体物体细节)和微观层面的视觉同构(候选对象几何结构相似但细微属性不同)。GRASP通过区域聚焦对齐(RFA)抑制背景干扰、促进以物体为中心的跨模态对齐,并利用语义扰动增强匹配(SPEM)和前景净化的语义原型码书(SPC)构建语义扰动负样本以增强细粒度语义判别能力。
Details
Motivation: 无人机场景固有的广视角和俯视视角给视觉-语言理解带来双重挑战:宏观上,视觉表征中过多的背景杂波导致跨模态焦点错位,模型优先关注全局环境相似性而非特定物体细节;微观上,视觉同构造成歧义,候选对象具有相似几何结构但仅在细微属性上不同。
Result: 在GeoText-1652基准和未见过的ERA数据集上的大量实验表明,GRASP在无人机视角细粒度图像-文本检索任务中取得了有竞争力的性能,验证了其在航空场景跨模态理解中的有效性。
Insight: 创新点在于提出粒度感知的区域对齐和语义原型学习框架,通过区域聚焦对齐(RFA)抑制背景干扰、促进物体中心对齐,以及利用语义原型码书(SPC)构建语义扰动负样本的语义扰动增强匹配(SPEM)策略,共同增强细粒度判别能力,为解决无人机视角下跨模态理解的宏观和微观挑战提供了系统方案。
Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At the macro level, overwhelming background clutter in visual representations leads to Cross-Modal Focus Misalignment, where the model prioritizes global environmental similarities over specific object details. At the micro level, Visual Isomorphism creates ambiguity, where candidates share similar geometric structures yet differ only in subtle attributes. To address these challenges, we propose the Granularity-Aware Region Alignment and Semantic Prototype (GRASP) learning framework, enhancing discriminative capability through two synergistic strategies. Specifically, we introduce Region-Focused Alignment (RFA) to promote object-centric cross-modal alignment while suppressing background interference. Concurrently, to tackle visual isomorphism, we propose Semantic Perturbation Enhanced Matching (SPEM), which leverages a foreground-purified Semantic Prototype Codebook (SPC) to construct semantically perturbed negatives for fine-grained semantic discrimination. Extensive experiments on the GeoText-1652 benchmark and the unseen ERA dataset demonstrate that GRASP achieves competitive performance in drone-view fine-grained image-text retrieval, validating its effectiveness for cross-modal understanding in aerial scenarios. Our code implementation is available at https://github.com/UCAS-JC/GRASP.
[163] UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation cs.CV | cs.AIPDF
Xuewan He, Tong Chu, Zihan Cheng, Yuchen Su, Qianxin Xia
TL;DR: 本文提出了一种名为UniDFKD的统一数据自由知识蒸馏框架,旨在解决现有方法因依赖架构特定统计先验(如批归一化统计量)而无法有效应用于现代架构(如视觉变换器ViT)的问题。该框架通过三个维度(CSC、SSA、SSD)引入与架构无关的语义先验来指导数据合成和知识迁移,从而提升合成数据的语义质量和蒸馏性能。
Details
Motivation: 现有数据自由知识蒸馏方法严重依赖架构特定的统计先验(如批归一化统计量)来指导数据合成,但这些先验在现代架构(如ViT)中常常缺失,导致合成数据语义质量下降和性能严重退化。因此,需要一种不依赖于具体网络架构的统一框架来保证有效的知识蒸馏。
Result: 在CNN和ViT上的大量实验表明,UniDFKD在同类和异类设置下均取得了新的最先进性能,平均绝对性能提升超过20%。
Insight: 创新点在于用显式的、与架构无关的语义先验(通过CSC定义合成内容、SSA锚定空间证据、SSD对齐空间证据)替代了架构特定的统计先验,从而实现了对多种网络架构的统一且高性能的知识蒸馏。从客观角度看,将语言嵌入和空间归因作为先验来源,为数据合成提供了更通用和语义丰富的指导。
Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD methods rely heavily on architecture-specific statistical priors (e.g., Batch Normalization statistics) to guide data synthesis, however, such architecture-dependent priors are often absent in modern architectures such as Vision Transformers (ViTs), resulting in degraded semantic quality of the synthesized data and consequently catastrophic performance degradation. In this paper, we propose \emph{UniDFKD}, a unified data-free knowledge distillation framework that replaces architecture-specific statistics with explicit, architecture-agnostic semantic priors. \emph{UniDFKD} governs the entire synthesis-distillation pipeline along three dimensions: (1) Categorical Semantic Conditioning (CSC) defines \emph{what} to synthesize by persistently modulating the generator with language-derived embeddings to capture semantic diversity; (2) Spatial Semantic Anchoring (SSA) dictates \emph{where} evidence belongs by anchoring the teacher’s spatial attributions to a Gaussian prior; and (3) Spatial Semantic Distillation (SSD) controls \emph{how} knowledge is transferred by explicitly aligning teacher-student spatial evidence alongside predictions. Extensive experiments across CNNs and ViTs demonstrate that UniDFKD establishes a new state-of-the-art, outperforming existing methods by an average absolute margin of over 20% in both homogeneous and heterogeneous settings.
[164] Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation cs.CVPDF
Jun Huang, Meiyi Chen, Zijie Yue, Yuhang Xiao, Fang Li
TL;DR: 该论文提出了一种基于视觉语言模型(VLM)的宫腔镜手术场景分割方法VLM-hyster,旨在解决宫腔镜视频中不同病灶形态高度相似以及存在镜面反射、运动模糊和液体遮挡等伪影带来的分割挑战。该方法利用预训练图像编码器提取鲁棒视觉特征,并结合基于Transformer的解码器进行密集预测,通过设计类别特定的文本提示和引入掩码蒸馏分支来提升分割性能。
Details
Motivation: 解决宫腔镜手术场景分割中因不同病灶形态高度相似以及视频中存在镜面反射、运动模糊和液体遮挡等伪影带来的独特挑战,以更好地理解宫腔镜术中环境并辅助计算机辅助干预。
Result: 在一个包含4020张高分辨率图像的大型多中心宫腔镜手术场景数据集上,VLM-hyster在十五个代表性类别的像素级定位任务中大幅超越了最先进的AI模型,并通过妇科医生的广泛评估以及多中心和前瞻性验证证明了其鲁棒性和泛化能力。
Insight: 创新点在于首次将视觉语言模型应用于宫腔镜手术场景分割,通过设计类别特定的文本提示和引入掩码蒸馏分支来过滤与文本提示相关性低的视觉特征,使模型能更有效地关注特定类别的图像区域,从而提升分割性能。从客观角度看,该方法将VLM的语义理解能力与手术场景的特定需求相结合,为解决医学图像分割中的复杂挑战提供了新思路。
Abstract: Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision-language model (VLM)-based hysteroscopic surgical scene segmentation method, which performs pixel-wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM-hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer-based decoder for dense prediction. Moreover, we design category-specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category-specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high-resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM-hyster substantially outperforms state-of-the-art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM-hyster’s robustness and generalizability. The results suggest that VLM-hyster earns considerable potential in enabling AI-assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at https://github.com/viscom-tongji/VLM-hyster.
[165] Warp-free Cross-view Geo-localization via Feature-space Consensus Mining cs.CVPDF
Zhuo Song, Lian Xu, Runqing Jiang, Yongjian Zhang, Kunhong Li
TL;DR: 本文提出了一种无需几何扭曲的跨视角地理定位方法,通过特征空间共识挖掘来克服街景与卫星图像间的视角差异和外观差异。该方法采用联合视角共识引导的学习框架,在特征空间中动态挖掘并自适应强化语义共识,避免了传统方法因几何变换引入的视觉失真和噪声监督。
Details
Motivation: 现有跨视角地理定位方法依赖几何扭曲来暴露共视线索,但这类变换基于限制性空间假设,在视角依赖的可见性下会引入严重视觉失真,导致噪声监督和脆弱对应关系。本文旨在绕过显式几何扭曲,直接通过特征空间中的语义共识实现鲁棒对齐。
Result: 在四个标准基准测试(如CVUSA、CVACT等)上进行了广泛实验,结果表明该方法达到了最先进的性能水平,验证了挖掘跨视角语义共识对可靠地理定位的重要性。
Insight: 创新点在于完全绕过显式几何扭曲,通过训练时的辅助联合视角路径实现跨视角直接交互,并引入全局模式探针作为语义词典,将不同模态投影到严格对齐的度量空间。共识引导的对比目标使单视角嵌入被显式拉向联合视角锚点,从而将共识挖掘能力蒸馏到单视角编码器中,实现推理时的鲁棒检索。
Abstract: Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.
[166] Revisiting the Current Frame: Physical-Trace-Guided Network Output Correction for Video Restoration cs.CVPDF
Yifeng Lin, Liuxiang Qiu, Guangming Ren, Tiesong Zhao
TL;DR: 本文提出了一种名为ANCHOR的模型无关框架,用于视频恢复任务中的输出校正。该框架将低质量的当前帧作为时间对齐的锚点,通过估计空间信任场来平衡恢复结果与原始观测,从而提升现有视频恢复模型的性能。
Details
Motivation: 现有视频恢复方法利用时序信息时,参考帧可能因物理成像变化、遮挡或时序聚合不完美而引入不一致的退化、内容差异或重建错误,且不同空间位置的输出可靠性未被充分探索。
Result: 在高动态范围视频重建和视频去雨任务上的实验表明,ANCHOR框架能持续提升多种最先进恢复模型的性能,验证了基于可靠性的输出校正的有效性。
Insight: 创新点在于提出了一种模型无关的后处理校正框架,通过物理痕迹证据估计空间信任场,自适应地融合恢复建议与原始观测,而非直接改进恢复网络本身。
Abstract: Video restoration methods exploit temporal information to recover information missing from degraded observations. However, reference frames within the sequence may introduce inconsistent degradation, content discrepancy, or reconstruction errors due to physical image-formation variations, occlusion, and imperfect temporal aggregation. Existing approaches mainly focus on improving restoration networks, while the reliability of the generated outputs at different spatial locations remains largely unexplored. In this work, we propose ANCHOR, a model-agnostic framework that revisits the low-quality current frame as a temporally aligned anchor for video restoration correction. Specifically, ANCHOR estimates a spatial trust field from heterogeneous physical-trace evidence and adaptively balances the restoration proposal with the original observation. Experiments on High Dynamic Range video reconstruction and video deraining demonstrate consistent improvements across various state-of-the-art restoration models, validating the effectiveness of reliability-aware output correction for video restoration.
[167] Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs cs.CVPDF
Ali Cheraghian, Hamidreza Dastmalchi, Hamed Barzamini, Morteza Saberi, Mojtaba Golzan
TL;DR: 本文提出了一种无需训练的幻觉缓解框架,用于动态抑制大型视觉语言模型(LVLM)中的幻觉。该方法通过构建一组解耦的幻觉子空间来隔离不同的幻觉模式,并在推理时根据每个输入实例自适应地计算权重,动态组合投影以选择性抑制最可能的幻觉方向,同时保留图像语义。
Details
Motivation: LVLM在生成文本时经常出现幻觉问题,即生成的文本不准确地描述视觉输入。现有基于模型编辑的方法通常依赖单一的全局子空间来纠正幻觉,无法捕捉不同输入间多样化的幻觉模式,因此需要一种动态的、针对每个实例的幻觉抑制方法。
Result: 在多个视觉语言基准测试和不同LVLM家族上的广泛实验表明,该方法能带来一致的性能提升,证明了其鲁棒性、泛化性和高效性。
Insight: 创新点在于构建了多个解耦的幻觉子空间来表征不同的幻觉模式,并引入了基于每个输入实例的自适应权重机制,实现了动态、针对性的幻觉抑制,避免了全局编辑的局限性,且无需训练。
Abstract: Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language models (LLMs). However, their reliability is frequently undermined by hallucinations, where generated text inaccurately describes the visual input. Although fine-tuning can mitigate this problem, it is computationally expensive and requires large, curated datasets, making training-free alternatives attractive. Among these, model editing is more promising than decoding-based approaches: decoding methods adapt outputs per input but introduce computational overhead and instability, whereas model editing modifies internal representations offline, providing a more efficient and stable solution. However, existing model-editing techniques typically rely on a single global subspace to correct hallucinations, treating all test samples identically and failing to capture diverse hallucination modes across inputs. To address this limitation, we propose a training-free hallucination mitigation framework for dynamic, per-instance suppression at test time. Our method first constructs a set of Disentangled Hallucination Subspaces, each isolating a distinct hallucination mode. During inference, the model adaptively calculates weights reflecting each input’s relationship to these subspaces, guiding a dynamically combined projection that selectively suppresses the most probable hallucination directions while preserving image-grounded semantics. Extensive experiments across multiple vision-language benchmarks and LVLM families demonstrate consistent improvements, highlighting the robustness, generalizability, and efficiency of our approach.
[168] Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation cs.CVPDF
Zhe Li, Honghao Qiao, Zhixin Xu, Qijie Wang, Bo Peng
TL;DR: 本文提出了一种用于游戏风格RGBA视频生成的端到端方法,通过构建GameAlpha-2.4K数据集并训练一个参考条件化的RGBA视频生成器,能够单次前向传播联合生成RGB帧和alpha遮罩。为了提高效率,引入了可见性路由器来早期识别透明token并跳过其后续的DiT更新,同时使用x_0-lock机制引导其沿原始流匹配轨迹收敛。
Details
Motivation: 解决游戏产业中高质量RGBA动画生成的两大挑战:一是缺乏游戏风格的RGBA视频数据集,二是传统先生成后抠图的流程会导致半透明区域模糊且抠图结果不稳定,而现有联合建模RGB与alpha的方法多为文本条件化,在效率和质量上仍有不足。
Result: 模型在FVD指标上优于传统的两阶段流程;所提出的可见性路由器在最后两个DiT去噪步骤中跳过了35%的token计算,相比密集推理实现了1.2倍的主干加速,且质量下降可忽略不计。
Insight: 创新点包括构建了面向游戏风格的RGBA视频数据集GameAlpha-2.4K,以及设计了可见性路由器与x_0-lock机制,实现了在扩散Transformer中高效识别并处理透明区域,从而在保证生成质量的同时显著提升推理速度。
Abstract: RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in gaming industry. However, generating high-quality RGBA animations for games remains challenging for two reasons. First, most existing RGBA video datasets are dominated by photorealistic content, with limited coverage of game assets. Second, the traditional generate-then-matte pipelines estimate alpha only after RGB synthesis, so semi-transparent regions are often blurred by background, resulting in unstable matting outputs. More recently, many methods have begun to model RGB and alpha jointly, but existing approaches are mostly text-conditioned, and still have unresolved issues in efficiency and quality. To address these challenges, we introduce GameAlpha-2.4K, a 2.4K-clip game-style RGBA video dataset built with matte-friendly synthesis, multi-hypothesis alpha recovery, and compositing-based quality gates. Using this dataset, we train a reference-conditioned RGBA video generator that jointly produces RGB frames and alpha mattes in a single pass. To improve efficiency, we propose a visibility router that identifies transparent tokens in an early stage and bypasses their later DiT updates, while x_0-lock guides them along the original flow-matching schedule toward self-predicted endpoints. Our model obtains lower FVD than traditional two-stage pipelines, and the visibility router skips 35% of token evaluations in the final two DiT denoising steps, providing a 1.2x backbone speedup with negligible quality degradation compared to dense inference.
[169] CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device cs.CVPDF
Francisco Guillén, Ricardo Almeida, Bruno Silva, João C. Neves
TL;DR: CableDex是一个基于计算机视觉的系统,通过智能手机拍摄单张照片,自动估算工业线缆卷轴上电缆的长度,解决了人工测量耗时且不准确的问题。该系统集成了相机标定、实例分割、姿态估计和体积计算等技术,适用于五种不同类型的卷轴和多种电缆尺寸。
Details
Motivation: 解决工业场景中,人工测量线缆卷轴长度效率低下、误差大的问题,旨在提供一种快速、准确的自动化测量方案。
Result: 在包含五种卷轴类型的75个卷轴数据集上评估,系统实现了4.90%的平均绝对百分比误差(MAPE),满足工业测量中10%的误差容限;其核心实例分割模型在1000张人工标注图像上训练,达到99.5%的mAP50,单图推理时间为5.66毫秒。
Insight: 创新点在于将移动设备摄影与多阶段计算机视觉流程(标定、分割、姿态、体积计算)结合,实现端到端的便携式工业测量;其高精度分割模型和针对卷轴几何的专用体积估计算法是关键。
Abstract: CableDex is a computer vision system that addresses the time-consuming and inaccurate manual measurement of cable length on industrial reels from a single photograph captured with a mobile phone. The system combines camera calibration, instance segmentation, pose estimation, and volumetric calculation to estimate the cable length across five different reel types and various cable sizes. This system is based on an instance segmentation model trained on 1,000 manually annotated images, achieving 99.5% mAP50 with an inference time of 5.66 ms per image. Evaluated on 75 reels across five reel types, the system achieves a MAPE of 4.90%, within the 10% error tolerance commonly accepted in industrial cable-reel measurement. The demonstration presents the end-to-end pipeline, from reel label scanning and image capture to segmentation and length estimation, through the mobile application.
[170] CoInS-Net: A Continuous Position-Aware Network for Joint Medical Image Interpolation and Segmentation cs.CVPDF
Yujia Sun, Ningfeng Que, Peiting Shi, Rongrong Fu, Yingying Yang
TL;DR: 本文提出了一种名为CoInS-Net的连续位置感知网络,用于联合执行医学图像插值和分割任务。该网络通过一个共享的Swin编码器,结合连续空间坐标查询,实现了插值与分割分支之间的双向交互,旨在解决传统独立处理方法导致的冗余计算和跨切片结构信息利用不足的问题。
Details
Motivation: 解决各向异性医学体数据因稀疏平面采样导致的结构不连续和边界模糊问题,并克服现有方法将插值与分割独立处理所带来的计算冗余和互补信息利用不足的缺陷。
Result: 在四个包含不同模态和解剖区域的公共医学影像数据集上的实验表明,该方法优于传统的单任务方案,实现了插值与分割任务的相互促进。
Insight: 创新点在于提出了一个联合优化框架,通过连续位置插值模块和基于原型的任务交互模块,使两个任务在共享解剖结构的同时保持各自边界级别的特定需求,无需额外标注,为智能临床医学图像分析提供了通用技术方案。
Abstract: Accurate medical image interpolation and anatomical structure segmentation are fundamental for computer-aided diagnosis and treatment planning. Anisotropic medical volumes with sparse through-plane sampling often suffer from structural discontinuity and boundary blur, hindering reliable clinical image analysis. Most existing methods implement interpolation and segmentation independently, which introduces redundant computation and fails to fully exploit complementary cross-slice structural information between sequential slices. To address these issues, we propose a continuous position-aware interaction network, termed CoInS-Net, for joint frame interpolation and lesion segmentation. Unlike conventional cascaded interpolation-then-segmentation paradigms, the framework enables bidirectional interaction under a shared Swin encoder with continuous spatial coordinate queries. A spatially continuous position interpolation module generates target-position features at every scale from the relative coordinate and physical spacing, and a prototype-based task mutual interaction module lets the segmentation and interpolation branches exchange global structure through a small set of shared prototypes rather than dense feature mixing. A multi-scale task-cooperative decoder further separates each scale into shared and task-specific components, so the two tasks reinforce common anatomy while preserving their distinct requirements down to the boundary level, without extra annotations. Experiments on four public medical imaging datasets with diverse modalities and anatomical regions demonstrate that the proposed method outperforms conventional single-task schemes. The joint optimization framework effectively realizes mutual promotion between interpolation and segmentation tasks, providing a reliable and universal technical scheme for intelligent clinical medical image analysis.
[171] Foundation Models are Implicit Deepfake Detectors cs.CVPDF
Stefan Smeu, Dragos-Alexandru Boldisor, Elisabeta Oneata, Dan Oneata
TL;DR: 该论文发现预训练基础模型(如自监督表示)在深度伪造检测中表现出一个一致现象:伪造样本的特征表示幅度系统性地低于真实样本。基于此,论文将深度伪造检测构建为异常检测问题,并证明简单的特征幅度统计方法即可达到与复杂方法相竞争的性能。研究进一步表明,特征幅度降低主要与伪造内容引入的语义偏移有关,而低层次生成指纹作用较小,且该判别信号随基础模型规模增大而增强。
Details
Motivation: 当前深度伪造检测方法严重依赖预训练自监督表示,但其区分真实与伪造媒体的内在机制尚不明确。论文旨在揭示这些表示中的系统性差异,并探索其作为零样本检测器的潜力。
Result: 在多个预训练模型、数据集以及图像和视频领域上,基于特征幅度统计的简单方法在深度伪造检测任务上取得了具有竞争力的性能,与更复杂的检测方法相当。
Insight: 论文的核心创新点在于发现并利用了基础模型中伪造样本特征幅度系统性降低这一鲁棒且可泛化的现象,将其转化为一个简单的异常检测问题。这揭示了大规模预训练模型本身蕴含了强大的零样本深度伪造检测能力,且其能力随模型规模扩展而增强,为利用基础模型进行媒体真实性认证提供了新视角。
Abstract: Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their properties make real and fake media distinguishable. In this work, we uncover a surprisingly consistent phenomenon: across multiple pretrained models, datasets, and both image and video domains, fake samples systematically produce lower-magnitude representations than their real counterparts. Motivated by this finding, we formulate deepfake detection as an anomaly detection problem and show that simple statistics of feature magnitude achieve competitive performance with far more sophisticated deepfake detection methods. We further investigate the origin of this effect and demonstrate that reduced feature magnitude is primarily associated with semantic shifts introduced by fake content, while low-level generative fingerprints play a comparatively smaller role. Finally, we show that this discriminative signal strengthens as the size of the underlying foundation model grows, suggesting that advances in representation learning naturally translate into stronger zero-shot deepfake detectors.
[172] Sekai2: From World Exploration to Interactive World Modeling cs.CVPDF
Kang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang
TL;DR: Sekai2是一个用于视频世界建模的多源真实世界视频数据集,包含来自113个国家或地区的128,892个视频片段,总计2,826小时。每个片段都提供了相机轨迹和分层的时间标注,解耦了主体运动、环境动态、静态场景内容和相机行为。特别地,它引入了982个沿非线性轨迹(包含循环和重访)拍摄的全景序列,为学习持久场景表示和几何一致的世界模型提供了关键监督。
Details
Motivation: 现有数据集(如大规模网络视频或姿态标注数据集)通常缺乏长视频、相机轨迹和时间对齐文本的三者结合,限制了长时程视频生成和相机可控合成的训练。Sekai2旨在填补这一空白,为交互式世界建模提供包含世界探索镜头的综合资源。
Result: 数据集分析表明,Sekai2实现了完整的姿态和字幕覆盖,具有广泛的地理和语义多样性、多样的相机轨迹以及高度非冗余的时间描述。它为长时程视频生成、相机可控合成和交互式世界模型预训练提供了可扩展的资源。
Insight: 创新点在于引入了包含循环和重访的非线性轨迹全景序列,这提供了跨时间和视角对同一位置的重复观察,为学习持久场景表示、长期空间记忆和几何一致的世界模型提供了关键且独特的监督信号。数据集的分层标注结构(解耦运动、动态、静态内容和相机行为)也是一个系统性的贡献。
Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.
[173] RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation cs.CV | cs.AIPDF
Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu
TL;DR: 本文提出了RecoverFly,一个用于无人机视觉语言导航(UAV-VLN)的、具备故障感知能力的强化学习后训练框架。该框架旨在解决现有端到端策略在交互式闭环执行中监督有限的问题,通过令牌级强化学习、故障案例重访以及两阶段长尾场景课程学习等技术,稳定优化策略并提升样本利用效率。
Details
Motivation: 现有端到端无人机视觉语言动作(UAV-VLA)策略的行为克隆目标在交互式闭环执行中提供的纠错监督有限,而直接应用强化学习又面临样本利用效率低、场景分布长尾以及策略分布偏移等挑战。
Result: 在TravelUAV基准测试中,RecoverFly在已见场景、未见地图和未见物体三个划分上均取得了最佳性能。相比AerialVLA初始化策略,在总探索预算仅为训练集大小约30%的条件下,成功率提升了3.12到8.37个百分点。
Insight: 创新点包括:将令牌级强化学习应用于语法约束的自回归无人机动作的稳定优化;通过重访未解决的故障案例来加强纠错学习和样本利用;结合两阶段长尾场景课程与参考策略正则化,以在提升场景适应性的同时保留已习得的能力。这些方法有效提升了策略的鲁棒性和泛化能力。
Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30% of the training-set size, validating its effectiveness, robustness, and generalization capabilities.
[174] FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search cs.CVPDF
Hieu Dinh Trung Pham, Phuong Huu Vu Tran, Thuan Duc Mai, Son Nguyen Minh Le, Khang Le Minh
TL;DR: 本文提出了一种名为FaLCon的检索框架,用于解决基于文本的行人异常搜索任务中的Sim2Real(从合成到真实)挑战。该框架采用由粗到细的检索策略,结合了全局语义匹配和细粒度验证,通过锚定约束、多专家融合和不确定性门控共识来提升检索精度。
Details
Motivation: 基于文本的行人异常搜索任务需要在Sim2Real设置下,根据详细的自然语言描述从真实世界行人图像库中检索目标,这面临两大挑战:一是细粒度差异(如细微动作、物体交互)难以捕捉,二是对整个图库应用大型多模态模型计算成本过高。
Result: 在PAB基准测试上,所提出的软声明感知检索方法达到了86.44%的mAP@10,显著优于单个检索骨干网络。完整的FaLCon框架进一步将性能提升至95.41% mAP@10、94.44% R@1和99.09% R@5,实现了新的SOTA水平。
Insight: 创新点在于提出了一个锚定约束的由粗到细检索框架,通过将昂贵的语义推理限制在一个由全局检索生成的小候选池中,有效平衡了检索精度与计算效率。其核心设计包括:利用完整和拼接描述作为锚点保证召回,使用多个语义方面进行有界校正,以及通过不确定性门控共识模块自适应融合多个专家结果。
Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or appearance attributes, while applying multimodal large language models to the entire gallery is computationally expensive. We propose an anchor-constrained coarse-to-fine retrieval framework that combines global semantic matching with fine-grained verification. First, each query is represented by its original caption, a structured concatenation, and several semantic facets. Heterogeneous vision-language retrievers are then integrated through robust per-query score calibration and soft claim-aware fusion. Full and concatenated captions serve as anchors to preserve candidate recall, whereas appearance, action, and object facets provide bounded corrective evidence. The resulting candidate pool is further refined by a discriminative Qwen3 reranker and two complementary semantic verification modules based on anomaly-aware cloze completion and multi-agent evidence reasoning. Finally, an uncertainty-gated consensus module adaptively reweights the three experts on ambiguous queries. Experiments on the PAB benchmark show that the proposed soft claim-aware retrieval achieves 86.44% mAP@10, substantially outperforming individual retrieval backbones. The complete framework further improves performance to 95.41% mAP@10, 94.44% R@1, and 99.09% R@5. These results demonstrate that preserving strong global retrieval while restricting expensive semantic reasoning to a small candidate pool is effective for fine-grained Sim2Real person anomaly search. Our code will be available on Github.
[175] Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge cs.CVPDF
Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao
TL;DR: 本文提出了一个用于MeViS-Audio竞赛任务的简单分阶段解决方案。该方法首先使用Qwen3-ASR将语音转换为文本,然后利用多个互补的基础模型和分割模型生成视频掩码轨迹,并通过计算掩码一致性来选择最佳轨迹。最后,结合视觉、视听和视频内查询的得分,通过分类器判断目标是否存在。该系统在竞赛中取得了第一名。
Details
Motivation: 解决MeViS-Audio赛道提出的挑战:根据口语运动表达在视频中分割描述的对象,并在目标不存在时返回空掩码。
Result: 提交的系统在MeViS-Audio赛道上获得了0.5952的J&F分数、0.7931的无目标准确率、0.9205的目标准确率,最终得分为0.769589,被竞赛组织者通知排名第一。
Insight: 主要创新点在于提出了一种基于一致性的多模型融合策略,通过计算候选掩码轨迹之间的平均一致性来选择最可靠的预测,而非依赖单一模型。此外,通过显式的方向、计数和复数规则处理复杂查询,并结合多模态特征进行视频级别的目标存在性分类,提升了系统的鲁棒性和准确性。
Abstract: The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J &F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.
[176] GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction cs.CVPDF
Khang Minh Le, Hieu Dinh Trung Pham, Luu Thanh Danh, Nam-Tien Le, Hieu Anh Ngo
TL;DR: 本文提出GeoRoute,一种无需训练的几何感知混合推理框架,用于交通场景的长时未来帧预测。该方法通过多帧时序上下文和视角条件路由,在预训练视频生成模型基础上稳定静态结构并提升时间一致性,适用于前视摄像头和异构交通视角视频。
Details
Motivation: 解决长时未来帧预测中因时序重影、几何漂移和物体运动不一致导致的视觉质量下降问题,特别是在结构化交通场景中直接应用现有潜在视频扩散模型时几何不稳定和时序连贯性退化。
Result: 在AI City Challenge Track 5基准测试中验证,最终系统在顶级团队中取得了有竞争力的性能。
Insight: 创新点在于无需重新训练或微调基础视频模型,通过推理时几何感知细化(如多帧深度分层渲染器)和视角条件混合推理(利用冻结视觉语言模型推断相机组并选择专用运动预测器),提升了静态几何稳定性和低级结构保真度。
Abstract: Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.
[177] Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance cs.CV | cs.AIPDF
Chunxiao Liu, Wei Liu, Anbin Xiong, Erli Meng
TL;DR: 本文提出了一种名为MGN-AIR的像素级多模态引导全合一图像修复框架,旨在解决现有方法对整个图像采用统一修复策略的局限性。该框架通过学习像素级视觉提示,并结合文本和视觉提示提供全局与局部退化线索,从而实现对每个像素进行更细粒度和精确的修复控制。
Details
Motivation: 现有全合一图像修复方法通常对整个图像应用统一的修复策略,忽略了不同区域可能遭受不同类型和严重程度的退化。本文的动机是提出一种像素级的修复方法,以实现更精细和精确的修复过程。
Result: 在涵盖去噪、去雨、去模糊、去雾、去雪和低光增强等多个任务的多个全合一图像修复基准测试上进行了广泛实验。实验结果表明,所提出的方法在性能上持续且显著地超越了现有方法。
Insight: 创新点在于将修复过程细化到像素级别,并利用多模态(文本和视觉)提示提供全局和局部退化指导,实现了对修复过程的更精细控制。从客观角度看,这种像素级多模态引导机制为全合一修复任务提供了一种新的、更灵活的范式。
Abstract: All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various types and levels of corruption using a single model. Recent works have achieved remarkable progress by learning degradation-adaptive prompts or network architectures. However, these methods typically apply a uniform restoration strategy across the entire image, neglecting the fact that different regions may suffer from distinct degradation types and varying degrees of severity. In contrast, we propose to perform restoration at the pixel level, thereby enabling more fine-grained and precise control over the restoration process. Specifically, we present MGN-AIR, a novel pixel-level restoration framework for all-in-one image restoration. Our approach first learns to estimate a pixel-level visual prompt. Then, it leverages both textual and visual prompts to provide global and local degradation cues, guiding the model on where to look and how to restore at each pixel. We conduct extensive experiments on multiple all-in-one image restoration benchmarks, covering a wide range of tasks including denoising, deraining, deblurring, dehazing, desnowing, and low-light enhancement. Experimental results demonstrate that our proposed method consistently and significantly outperforms existing approaches.
[178] XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher cs.CV | cs.LGPDF
Lazar Đoković, Aimee Lin
TL;DR: 本文对轻量级图像特征提取与匹配器XFeat进行了可复现性研究,通过重新实现架构、重新评估原始检查点并进行消融实验,验证了其在标准图像匹配基准上的准确性与效率权衡。研究发现复现模型在MegaDepth-1500和ScanNet-1500上表现接近或优于原始检查点,但原始论文中部分架构设计(如并行关键点分支)的实际效益被高估,且下游任务(如Aachen视觉定位)的结果对未明确的评估细节敏感。
Details
Motivation: 针对XFeat论文、补充材料和公开代码在实现细节上的不一致,本研究旨在通过可复现性分析验证其作为轻量级图像匹配器的有效性,并深入探究其架构设计选择的合理性。
Result: 在MegaDepth-1500和ScanNet-1500基准测试中,复现模型与原始检查点性能接近,部分情况下甚至更优,支持了XFeat在准确性与效率间取得良好权衡的结论;但在Aachen视觉定位任务中,结果低于报告值,表明评估细节敏感。
Insight: 研究强调了可复现性分析中重新实现与重新评估的区别,揭示了原始论文中并行关键点分支对半稠密匹配的重要性被高估,且跨模态匹配(如视网膜、热可见光图像)在严重模态偏移下性能急剧下降,为轻量级特征设计提供了更细致的见解。
Abstract: We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors’ released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work. This distinction between re-evaluation and reproduction is important, as the paper, supplement, and public code differ in several implementation details, including the backbone layout, fusion block, and training losses. Empirically, our reproduced models closely match and, in some cases, outperform the re-evaluated original checkpoint on MegaDepth-1500 and ScanNet-1500, supporting the main claim that XFeat provides a strong accuracy-efficiency trade-off for standard image-matching benchmarks. Our ablations provide a more nuanced view of two architectural arguments from the original paper. In particular, the parallel keypoint branch is important for semi-dense matching, but its benefit is less pronounced than originally claimed, while the evidence for the specific placement of the single skip-connection remains inconclusive. Finally, we reproduce the original downstream evaluations and find close agreement for homography estimation, while Aachen visual localization remains below the reported results, even for the released checkpoint, suggesting sensitivity to underspecified evaluation details. We then extend the analysis to zero-shot out-of-distribution and cross-modal matching across retinal, thermal-visible, and multimodal remote-sensing imagery, where XFeat remains effective in some settings but degrades sharply under severe modality shifts.
[179] Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework cs.CVPDF
Dongxu Ge, Shansong Liu, Cheng Gong, Xiao-Lei Zhang, Chi Zhang
TL;DR: 本文针对音频到图像(A2I)生成任务,提出了一个高质量的三模态数据集A2I-Set和一个名为AudioCanvas的生成模型。该研究旨在解决现有数据集在图像保真度和跨模态对齐精度上的不足,从而提升A2I生成的质量。
Details
Motivation: 现有A2I生成方法的性能受限于传统数据集,这些数据集往往缺乏高保真图像和精确的跨模态对齐,导致即使微调强大的文本到图像(T2I)模型也难以实现高质量生成,这限制了该领域的实际应用。
Result: 实验表明,基于A2I-Set微调的AudioCanvas模型,在生成结果的视觉表现力和跨模态对齐方面,普遍优于现有方法。
Insight: 主要创新点在于构建了一个统一、高质量、包含音频、图像和详细文本描述的三模态数据集A2I-Set,并通过人工监督创建了一个混合来源的测试集,为音频视觉研究提供了关键的数据基础。从客观角度看,高质量、对齐良好的多模态数据集是提升特定跨模态生成任务性能的关键因素。
Abstract: As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted increasing research attention in recent years. Nevertheless, despite the remarkable visual quality of modern text-to-image (T2I) models, the performance of A2I remains fundamentally limited by traditional datasets, which often lack both high-fidelity images and precise cross-modal alignment. As a result, existing methods still struggle to achieve high-quality audio-to-image generation through finetuning strong T2I models, thereby constraining practical applications in this area. Motivated by this gap, we introduce A2I-Set, a unified, high-quality tri-modal dataset consisting of 323K paired audio, images, and detailed text captions, specifically designed for audio-visual research, including audio-conditioned image generation. Besides, we developed a new mixed-source test set for the A2I task through human supervision. We further propose an A2I model, AudioCanvas, fine-tuned on our A2I-Set. Experiments show that AudioCanvas achieves more visually expressive as well as cross-modal alignment results that generally outperforming existing approaches. Our dataset and source code are available at https://github.com/gdx012/A2I-Generation.
[180] PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images cs.CVPDF
Changhai Ma, Ziyu Wu, Yunkang Zhang, Fangting Xie, Mengting Niu
TL;DR: 本文提出了一种名为MDP-Net的端到端网络,用于从多个设备的时序压力图像中直接估计人体网格。为了解决现有单设备方法监测范围受限的问题,作者构建了高质量的多设备时序压力数据集MDP,并引入了一种受混合专家框架启发的多模态融合机制,以实现跨设备压力信息的有效互补与增强。
Details
Motivation: 基于压力的人体姿态监测因其保护隐私的特性,已成为非侵入式感知的主要方法,但现有方法通常局限于单一设备,限制了有效监测范围。
Result: 在自建的MDP数据集上,MDP-Net实现了12.6厘米的关节位置误差,证明了融合多设备压力信息对于日常人体姿态监测是一种有效且有前景的新方案。
Insight: 主要创新点在于提出了首个从多设备压力数据估计3D人体网格的端到端网络MDP-Net,并构建了相应的多设备时序压力数据集;其受MoE启发的多模态融合机制,为跨设备异构信息的有效融合提供了新思路。
Abstract: Human pose monitoring is crucial in fields such as rehabilitation assessment and human-computer interaction. Due to its privacy-preserving nature, pressure-based human pose monitoring has become a primary approach for unobtrusive sensing. However, existing methods are generally limited to a single device, which restricts the effective monitoring range. To address this limitation, we propose MDP-Net, an end-to-end network capable of directly estimating human meshes from temporal pressure data across multiple devices. We introduce a multimodal fusion mechanism inspired by the Mixture of Experts (MoE) framework to achieve effective complementarity and enhancement of cross-device pressure information. To support the training and evaluation of MDP-Net, we constructed MDP, a high-quality multi-device temporal pressure dataset that includes various pose labels such as 2D/3D joints and human meshes. Experimental results demonstrate that MDP-Net achieves a joint position error of 12.6 cm on the MDP dataset. These results prove that fusing multi-device pressure information is an effective and promising new solution for daily human pose monitoring.
[181] DocPure: Prompt-Free Unified Document Restoration via Degradation-Aware Structure-Guided Wavelet Modulation cs.CVPDF
Lingming Su, Wanglong Lu, Tao Wang, Kaihao Zhang, Nan Zhang
TL;DR: DocPure是一种无需提示的统一文档修复框架,通过退化感知结构引导的小波调制,能够处理多种退化类型(如模糊、噪声、压缩伪影和阴影)的文档图像恢复。该方法利用退化感知结构自编码器预测干净的结构先验,并采用结构引导的小波交互机制在频域和空间语义之间建立桥梁,实现跨频率自适应调制以提升恢复质量。
Details
Motivation: 现有统一文档修复方法通常需要训练多个退化特定模型、依赖手动任务提示或跨任务数据配对,这限制了其效率和泛化能力。DocPure旨在解决这些限制,通过无需提示的框架实现退化感知的文档修复,简化训练和推理过程。
Result: 在去模糊、去噪、压缩伪影减少和去阴影等多个任务上的广泛实验表明,DocPure相比最先进方法(SOTA)表现出强大的性能,实现了高质量的文档图像恢复。
Insight: 创新点包括退化感知结构自编码器(通过退化信息路由正则化预测结构先验,仅在训练中使用退化标签作为辅助监督)和结构引导的小波交互机制(利用低频子带调制高频恢复,确保结构一致性),这些设计提升了模型对多种退化的适应性和恢复效果。
Abstract: High-quality document images are pivotal for information archiving and downstream automatic processing. However, they are frequently compromised by diverse degradations during uncontrolled acquisition and transmission. While unified document restoration techniques have been proposed to restore images from multiple degradations, they often struggle with training multiple degradation-specific models, reliance on manual task-specific prompts, or cross-task data pairing. To address these limitations, we propose DocPure, a prompt-free unified framework that achieves degradation-aware document restoration. We design a degradation-aware structure auto-encoder with degradation-informed routing regularization to predict clean structural priors from degraded inputs. The model is prompt-free at inference, and degradation labels are only used as auxiliary supervision for the routing regularization during training. Furthermore, we introduce a structure-guided wavelet interaction mechanism to bridge frequency-domain features and spatial semantics. Within the structure-guided wavelet interaction mechanism, a cross-frequency adaptive modulation utilizes low-frequency sub-bands to modulate high-frequency recovery, ensuring structural consistency. Extensive experiments demonstrate that DocPure achieves strong performance compared with state-of-the-art methods across various tasks, including deblurring, denoising, compression artifact reduction, and deshadowing.
[182] From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation cs.CV | cs.AIPDF
Zeyuan Ma, Jiaxin Chen, Di Huang
TL;DR: 本文提出了一种统一的语义到决策框架,用于解决无人机视觉语言导航(UAV-VLN)中的三个耦合问题:视觉观察中指令相关地标的基础语义较弱、长时程历史信息利用不足以及局部陷阱或重复探索下的决策不稳定。该框架包含指令基础语义增强模块、相关性感知动态时间聚合策略和拓扑感知决策方法。
Details
Motivation: 当前UAV-VLN方法存在三个相互关联的问题:视觉观察中对指令相关地标的语义基础较弱、长时程历史信息利用不足,以及在局部陷阱或重复探索下决策不稳定。
Result: 在广泛使用的AerialVLN和OpenFly基准测试上的实验表明,该方法达到了最先进的性能水平。
Insight: 创新点在于提出了一个统一的语义到决策框架,通过指令基础语义增强、相关性感知动态时间聚合和拓扑感知决策优化,将语义基础与决策优化紧密结合,以提升长时程导航的鲁棒性和准确性。
Abstract: UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. First, we present an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state. Subsequently, we develop a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frames into structured landmark prompts for the decoder. Finally, we devise a topology-aware decision method that combines local-optimum cognition with group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. Experiments on the widely used AerialVLN and OpenFly benchmarks clearly demonstrate that our method achieves state-of-the-art performance.
[183] VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation cs.CVPDF
Jiajun Xu, Yanghao Zhou, Jingyun Liao, Yu Bai, Jinxing Zhou
TL;DR: 本文提出了VideoVIBE,一个基于视频的细粒度诊断基准,用于评估一次性生成的交互式网页应用的质量。该基准包含约1.7K个视频问答实例,覆盖了语义逻辑、视觉动态、结构时序和功能故障。同时,作者提出了一个无需训练的多智能体系统V2Lens,通过视觉和源代码验证来提升诊断性能。
Details
Motivation: 现有评估方法通常只对孤立的产物或最终任务结果进行评分,无法揭示故障的具体类型和原因,因此需要一种更细粒度、基于行为观察的诊断性评估基准。
Result: 在13个闭源和开源视频多模态大模型中,Gemini-2.5-Flash表现最佳,得分为64.54;而提出的V2Lens系统达到了71.72分,提升了7.18分,显著优于现有模型。
Insight: 创新点在于将网页操作录屏转化为细粒度的诊断任务,并构建了一个大规模、多类别的故障诊断数据集。同时,提出的V2Lens系统采用无需训练的多智能体协作框架,通过证据驱动的验证和选择性修正来提升诊断的准确性和可靠性。
Abstract: Natural-language-driven “vibe coding” enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.
[184] Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation cs.CV | cs.AIPDF
Yifei Xue, Yuanchen Fei, Hao Zhang, Chenzhi Nie, Tie ji
TL;DR: 本文提出了一种名为GeoCon-Bench的新基准,用于评估AI生成内容(AIGC)视频的质量,其核心是通过定量测量生成视频序列中跨帧的几何一致性,来评估视频内容是否符合物理定律(如运动规律)。该基准通过估计全局运动、拟合单应性或基础矩阵模型,并报告内点比率和几何误差等互补指标来实现。
Details
Motivation: 现有的视频质量评估(VQA)方法主要关注视觉和谐度、视频-文本一致性和特定领域对齐,但缺乏衡量视频内容对物理定律(如几何一致性)遵循程度的定量指标。本文旨在填补这一空白,为AIGC视频提供一个基于物理原理一致性的质量评估基准。
Result: 在多个最先进的AIGC模型上的实验表明,GeoCon-Bench作为一种视频质量评估指标是可靠的。该基准还附带了一个包含6种运动类别、20个场景的数据集。
Insight: 创新点在于首次提出了一个专门用于量化评估AIGC视频几何一致性(即对物理定律的遵循程度)的基准和指标集(如内点比率、几何误差),将视频质量评估从传统的感知或语义层面扩展到了物理合理性层面。
Abstract: Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality assessment (VQA) metrics to evaluate AI-generated content (AIGC) videos and guide model optimization. Existing studies assess video quality through visual harmony, video-text consistency, and domain-specific alignment, yet lack quantitative metrics for measuring fidelity to physical laws. To address this limitation, we present a novel benchmark that evaluates the quality of AIGC videos based on their compliance with physical principles by quantitatively measuring geometric consistency across frames extracted from generated sequences. This serves as a proxy for estimating the extent to which generated videos conform to real-world physical rules. Specifically, GeoCon-Bench captures global motion through translation estimation, fits homography or fundamental matrix models using background correspondences, and reports complementary metrics, including inlier ratio and geometric error. We also release a dataset containing 20 scenes across six motion categories. Experiments on state-of-the-art AIGC models demonstrate the reliability of GeoCon-Bench as a video quality assessment metric.
[185] Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection cs.CV | cs.AI | cs.LG | eess.IV | eess.SPPDF
Weize Cai, Yongqi Dong, Zhida Shao, Yichen Liu, Zixin Fu
TL;DR: 本文提出了一种结构增强与质量感知的框架,用于提升车道线检测的鲁棒性。该框架包含一个门控水平-垂直令牌模块来增强主干网络特征的结构连续性,以及一个线质量感知动态锚框评分模块来校准分类置信度与定位质量之间的关联。
Details
Motivation: 解决基于锚框的车道线检测器中存在的两个耦合问题:主干网络特征在部分可见车道线上失去结构连续性,以及分类置信度与线级定位质量解耦,导致不准确的锚框在非极大值抑制前持续存在。
Result: 在VIL-100数据集上,该方法将ADNet-R34的F1@50分数从89.97提升至91.28,减少了假阳性和假阴性。在CULane和TuSimple数据集上的额外实验、广泛的消融研究、分数分布诊断和运行时分析证实了该方法在结构性和排序方面的互补性改进,且计算开销最小。
Insight: 创新点在于通过轻量级的方向性令牌交互增强特征的结构连续性,以及通过质量监督、硬负样本抑制和成对排序来动态校准锚框评分,而无需增加推理分支,从而在保持高效推理流程的同时提升了检测性能。
Abstract: Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based detectors provide efficient candidate generation, their performance is limited by two coupled issues: backbone features often lose structural continuity along partially visible lanes, and classification confidence may decouple from line-level localization quality, allowing inaccurate anchors to persist before non-maximum suppression (NMS). We propose a structure-enhanced and quality-aware framework that improves lane representation and dynamic-anchor scoring while preserving the inference pipeline of the Anchor Decomposition Network (ADNet). Specifically, a Gated Horizontal-Vertical Token (GHVT) module enhances mid- and high-level backbone features via lightweight directional token interactions with a learnable residual gate. In parallel, Line-Quality-Aware Dynamic Anchor Scoring (LQAS) calibrates existing classification logits using quality supervision, hard-negative suppression, and pairwise ranking without adding inference branches. On the VIL-100 dataset, our method improves ADNet-R34 from 89.97 to 91.28 in F1 score at the 0.5 intersection-over-union threshold (F1@50), reducing both false positives and false negatives. Additional experiments on CULane and TuSimple datasets, extensive ablations, score-distribution diagnostics, and runtime analysis confirm complementary structural and ranking improvements with minimal computational overhead.
[186] LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection cs.CV | cs.LGPDF
Peter Lorenz, Anjith George, Marcel Sébastien
TL;DR: 本文系统评估了32种基础模型在面部呈现攻击检测(PAD)任务中的表现,发现零样本提示性能接近随机猜测,而使用LoRA微调虽能在单个数据集内实现低于2%的ACER,但跨数据集泛化性能显著下降。研究表明,预训练表示和适应数据集对跨数据集泛化的影响大于轻量级适应策略本身。
Details
Motivation: 解决现有PAD方法在跨数据集评估中性能下降的问题,探索基础模型在PAD任务中的潜力,并系统评估不同架构和训练过程的基础模型,以弥补现有研究主要关注CLIP类模型的不足。
Result: 在MCIO基准(包括MSU-MFSD、CASIA-FASD、Replay-Attack和OULU-NPU)上,零样本提示性能接近随机水平;LoRA微调(可训练权重少于1%)在大多数情况下实现低于2%的单个数据集内ACER,但跨数据集ACER显著更高。
Insight: 创新点在于首次系统评估多种基础模型在PAD任务中的表现,并揭示LoRA微调主要优化单个数据集内的决策边界,而跨数据集泛化更依赖于预训练表示和适应数据集的质量,而非轻量级适应策略本身。
Abstract: Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variations in sensors or lighting conditions can reduce the effectiveness of detectors from near-perfect to nearly random. Foundation models (FMs) have emerged as a promising alternative because typical PAD datasets, such as the MCIO benchmarks (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU), are small relative to the scale used for web-based pretraining. However, existing PAD systems primarily focus on CLIP-based foundation models, while overlooking other FMs with different architectures and training procedures. This study addresses this question by systematically evaluating 32 FMs. Zero-shot prompting achieves performance near chance across model families and scales. The vision encoders, when low-rankadapted (LoRA) with fewer than 1% trainable weights, achieve below 2% intra-dataset ACER in most cases, while cross-dataset ACER is substantially higher. LoRA primarily refines the decision boundary within a dataset, suggesting that pretrained representations and the adaptation dataset play a larger role in cross-dataset generalization than the evaluated lightweight adaptation strategy.
[187] NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation cs.CV | cs.AIPDF
Haiyang Yan, Jinyue Guo, Yanchao Zhang, Bingqing Wang, Zhenchen Li
TL;DR: 本文提出NeuroRefiner,一种用于3D荧光显微镜神经元分割的多智能体精炼系统。它通过模拟人类专家迭代全局观察和局部编辑的工作流程,利用三个协作智能体诊断拓扑错误、生成校正指令和验证精炼质量,并结合专门的TopoRefineNet工具进行体素级编辑,以生成拓扑更准确、可解释性更强的分割结果。
Details
Motivation: 现有方法难以同时保留神经元的局部细节和全局拓扑结构,导致分割结果碎片化,无法准确处理神经元稀疏且细长的形态。
Result: 在BigNeuron、CWMBS和ZBFWB数据集上的实验表明,NeuroRefiner优于现有最先进方法,在具有挑战性的ZBFWB数据集上F1分数显著提升了3.02%。
Insight: 创新点在于将分割精炼过程形式化为一个多智能体协作系统,模拟人类专家工作流,并引入了结合跨模态特征融合的专用工具TopoRefineNet,通过多轮推理和编辑迭代提升拓扑准确性。
Abstract: Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses significant challenges to existing segmentation methods. These methods struggle to preserve both local details and global topology, leading to fragmented results. To address this, we propose NeuroRefiner, a multi-agent system that formalizes the human expert workflow involving iterative global observation and local editing. Specifically, NeuroRefiner comprises three collaborative agents dedicated to diagnosing topological errors, generating correction instructions, and validating refinement quality. To facilitate agent instruction-guided segmentation refinement, we propose TopoRefineNet, a dedicated 3D U-Net-based tool that leverages cross-modality feature fusion to generate refined masks. Through multi-round agent reasoning and voxel-level editing, NeuroRefiner produces topologically more accurate segmentations with enhanced interpretability. Experiments on the BigNeuron, CWMBS, and ZBFWB datasets demonstrate that NeuroRefiner outperforms state-of-the-art methods, notably achieving a 3.02% improvement in F1 score on the challenging ZBFWB dataset.
[188] DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation cs.CV | cs.AIPDF
Zian Li, Litong Gong, Borui Liao, Pengfei Liu, Xinyu Wang
TL;DR: 本文提出DUET和DUET+方法,通过噪声级专家分工(sCM专家处理高噪声步以生成多样结构,DMD专家处理低噪声步以优化外观细节)来调和视频生成中两步蒸馏的质量与多样性权衡,避免了损失级组合的优化困难。
Details
Motivation: 针对扩散模型视频生成中迭代采样成本高的问题,现有两步蒸馏方法存在质量与多样性权衡:轨迹级蒸馏(如sCM)偏向多样性,而分布级蒸馏(如DMD)偏向质量,DUET旨在极端两步生成中协调这两类范式。
Result: 在Wan2.1-T2V-1.3B骨干网络上,DUET将sCM的两步生成质量提升至接近DMD水平,同时保持其结构多样性(约为DMD的两倍);DUET+进一步提升了整体质量并保留了多样性优势。
Insight: 创新点在于噪声级专家分工的简单有效范式,通过独立训练专家避免优化冲突,并结合RL引导的专家适应(DUET+)缓解中继接口和高噪声阶段的瓶颈,为两步视频生成提供了兼顾质量与多样性的解决方案。
Abstract: Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality–diversity trade-off between its two dominant paradigms: trajectory-level distillation (e.g., sCM) favors diversity, whereas distribution-level distillation (e.g., DMD) favors quality. Targeting extreme two-step video generation, we introduce DUET, which reconciles the two paradigms through a noise-level duet of experts: an sCM expert takes the high-noise step to lay out diverse structure, and a DMD expert takes the low-noise step to refine appearance detail. Since the two experts are trained independently with their native objectives, DUET sidesteps the optimization difficulties of loss-level combinations and delivers quality and diversity jointly rather than trading one for the other. We further identify the relay interface and the high-noise stage as the remaining bottlenecks, and address them with RL-guided expert adaptation, yielding DUET+. With the Wan2.1-T2V-1.3B backbone, DUET lifts the two-step quality of sCM close to the level of DMD while retaining nearly all of its structural diversity—about twice that of DMD—and DUET+ further improves overall quality while preserving this diversity advantage. Together, these results establish noise-level expert specialization as a simple, effective paradigm for reconciling diversity and quality in two-step video generation.
[189] EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization cs.CVPDF
Yifei Cao, Guolong Wang, Mingliang Hou, Xiya Bu, Daming Liu
TL;DR: 本文提出EgoHieraLoc,一个受人类视觉分层处理机制启发的统一框架,用于解决第一人称视角视频中的视觉查询定位问题。该框架通过分割引导的判别性解析、查询感知的鲁棒定位以及区域自适应模块来精化边界,并引入几何-语义联合置信度机制将2D定位扩展至3D。
Details
Motivation: 现有视觉查询定位方法在物体边界模糊、全局上下文难以有效指导细粒度定位时面临挑战。本文动机源于人类视觉处理此类模糊性的分层能力,旨在模仿其快速筛选、选择性注意、上下文反馈及多视角可信度整合的过程。
Result: 在VQL-2D和VQL-3D基准测试上进行了广泛实验,结果表明所提方法取得了最先进的性能。
Insight: 创新点包括:1)将人类视觉分层处理机制(前景筛选、选择性注意、上下文反馈)系统化为一个统一的计算框架;2)提出几何-语义联合置信度,通过乘性耦合分割置信度、局部深度一致性、多视角反投影一致性和三角测量基线质量,实现仅当视角在语义和几何上均可信时才贡献于3D估计,从而鲁棒地集成多视角证据。
Abstract: Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly screens foreground candidates, selectively attends to the target despite distractors, refines perception via feedback between global context and local detail, and, when a single view is unreliable, integrates evidence across viewpoints according to its credibility. Inspired by these competencies, we propose \textbf{EgoHieraLoc}, a unified framework for VQL-2D and VQL-3D. A Discriminative Parsing Module first extracts foreground-aware query representations using segmentation priors; a Query-Aware Module then performs robust target localization through discriminative correlation filtering with deformable modeling; and a Regional Adaptation Module feeds multi-scale context back into local regions to recover precise object boundaries. To extend this perceptual hierarchy to 3D localization, we introduce Geometric-Semantic Joint Confidence (GSJC), which multiplicatively couples segmentation confidence with local depth consistency, multi-view back-projection consistency, and triangulation-baseline quality, so that a viewpoint contributes to the 3D estimate only when it is credible both semantically and geometrically. Extensive experiments demonstrate state-of-the-art performance on both VQL-2D and -3D benchmarks.
[190] Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cs.CVPDF
Jiahao Shao, Yuanbo Yang, Yiyi Liao, Yujun Shen, Ceyuan Yang
TL;DR: 本文提出TextCall方法,通过仅保留工具调用时的结构化文本(如工具名称、坐标、目标描述等)而跳过返回的图像像素,验证了在视觉推理任务中,工具调用的文本支架而非返回的像素承载了主要信息。实验表明,该方法在保持或提升性能的同时,显著降低了延迟和API调用开销。
Details
Motivation: 针对工具增强的视觉语言模型中,返回的图像像素对性能增益贡献有限的现象,研究旨在探究工具调用过程中结构化文本支架是否才是关键信号,以理解模型’用图像思考’的本质。
Result: 在LoRA、全微调和强化学习等多种设置下,TextCall在多个基准测试中匹配或超越了完整’用图像思考’方法的性能,同时将延迟降低了29-46%,并避免了因看到返回图像而停止调用工具的直接回答失败模式。
Insight: 创新点在于提出了工具调用文本支架假说,并通过TextCall实验设计验证了在现有基准中,结构化文本是负载信号,而返回图像是冗余的;这为设计更高效的视觉推理系统提供了新思路,即可能无需依赖高成本的图像返回与处理。
Abstract: Tool-augmented vision-language models increasingly “think with images”: they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not carry the gain, what does? We hypothesize that the load-bearing signal is the structured text emitted before any returned pixel arrives: tool name, coordinates, target description, and intent. This textual scaffold encodes where to look and what to find. We introduce TextCall (call-but-no-return) to test this: it keeps the scaffold but replaces returned images with the text placeholder [Image output skipped]. Three studies support the hypothesis. (i) Non-necessity of returned pixels: across LoRA, full fine-tuning, and RL, TextCall matches or exceeds full thinking-with-images; under RL it preserves tool use at the reported checkpoint, avoiding the failure mode where, under matched settings, seeing the returned image causes the model to stop calling tools and answer directly. (ii) Sufficiency of the scaffold: on matched training queries, scaffold-only input yields equivalent accuracy to returned-image input. (iii) Component specificity: decomposing the scaffold into reasoning text and spatial code shows both components contribute, with the dominant one varying by task. Together these results support the Tool-Call Scaffold Hypothesis: in current thinking-with-images distributions, the active signal is the structured text emitted at tool-call time; the returned image is a redundant carrier. TextCall preserves accuracy while reducing latency by 29-46% and eliminating tool-execution API calls. Our claims hold for current thinking-with-images benchmarks; constructing tasks where pixels are genuinely load-bearing remains an open direction.
[191] CIFA: Contextual-Intersectional Fairness Auditing for Hidden Subgroup Discovery in Face Analysis cs.CVPDF
Nazia Aslam, Khalid Adnan Alsayed, Thomas B. Moeslund, Kamal Nasrollahi
TL;DR: 本文提出了CIFA(Contextual-Intersectional Fairness Auditing)框架,用于系统性地审计人脸分析模型中由人口统计属性(如性别、种族)与上下文因素(如光照、图像质量)交互作用产生的隐藏子群性能下降问题。该框架通过人口统计、上下文及上下文交叉审计,结合最差组发现来识别和排序最脆弱的属性组合。
Details
Motivation: 现有计算机视觉的公平性评估主要依赖总体准确率和人口统计子群分析,但模型性能也受上下文因素影响,这些因素可能与人口特征相互作用,导致在总体准确率和人口公平性看似可接受的情况下,某些隐藏子群的性能严重下降。
Result: 在FairFace、CelebA和UTKFace数据集上,使用ResNet-50和ViT-B/16模型进行性别分类评估,结果表明总体准确率和仅人口统计评估会掩盖显著的上下文交叉差异。通过审计-缓解-再审计协议评估几种现有缓解策略,发现虽然某些最差组差异有所减少,但没有单一策略能在所有数据集和架构上一致地消除它们。
Insight: 创新点在于提出了一个结构化的上下文交叉公平性审计框架,将上下文因素正式纳入公平性评估,并系统性地发现和优先处理隐藏子群风险。客观来看,该研究强调了在公平性评估中考虑属性交互作用的重要性,并提供了一个可复现的框架来发现和重新评估风险,这对于构建更鲁棒和公平的视觉系统具有借鉴意义。
Abstract: Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitive to contextual factors such as illumination, blur, image quality, facial accessories, and appearance attributes. These factors may interact with demographic characteristics, producing hidden subgroups in which performance degrades substantially despite strong aggregate accuracy and apparently acceptable demographic fairness. To address this, we propose the Contextual-Intersectional Fairness Auditing Framework (CIFA), a structured framework for identifying subgroup vulnerabilities arising from interactions between demographic and contextual attributes. CIFA performs demographic, contextual, and contextual-intersectional auditing, followed by worst-group discovery to identify and rank the most vulnerable attribute combinations. We evaluate CIFA on gender classification using ResNet-50 \cite{he2016deep} and ViT-B/16 \cite{dosovitskiy2020image} across FairFace \cite{Karkkainen2021}, CelebA \cite{Liu2015}, and UTKFace \cite{Zhang2017}. Our results show that aggregate accuracy and demographic-only evaluation can mask substantial contextual-intersectional disparities. We further assess several established mitigation strategies through an audit–mitigate–reaudit protocol and find that, although some worst-group disparities are reduced, no single strategy consistently eliminates them across datasets and architectures. These findings establish contextual-intersectional auditing as an important component of fairness evaluation and provide a reproducible framework for discovering, prioritizing, and reassessing hidden subgroup risks in face analysis systems.
[192] LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection cs.CVPDF
Renshan Zhang, Haoyang Meng, Yixiao He, Rui Shao, April Hua Liu
TL;DR: 本文提出LookAgain,一种基于视觉反思的闭环GUI定位模型,将界面元素定位重新定义为多轮预测-观察-修正过程,通过“定位”和“确认”两个基本操作迭代优化坐标假设,显著提升了小目标、密集控件和分布外界面的定位精度。
Details
Motivation: 现有GUI定位模型在单次预测后缺乏反思机制,导致对小目标、密集控件和分布外界面性能下降,本文旨在通过闭环视觉反思解决这一范式局限。
Result: 在包含拒绝感知和通用GUI定位的基准测试中,LookAgain均取得一致性能提升,并达到最先进水平(SOTA)。
Insight: 创新点在于将GUI定位重构为多轮闭环反思过程,通过视觉标记锚定空间先验进行迭代修正,并采用监督微调与基于终端正确性的强化学习(GRPO)联合训练策略。
Abstract: Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, densely packed controls and out-of-distribution interfaces. We attribute this gap to a paradigmatic limitation shared by existing approaches: none of them treats a produced coordinate as a hypothesis to be reflected upon and revised under new visual evidence. This manifests as three coupled issues: 1) Lack of post-hoc reflection. The prediction is frozen at the moment of emission, leaving no internal mechanism to challenge or refine it. 2) Visual evidence decoupled from the prediction. The auxiliary visual evidence is gathered to support the upcoming coordinate rather than to scrutinise the one already committed to. 3) Refinement over views, not over predictions. The iterative zoom-in refines the inspected region instead of inheriting a previous coordinate as a spatial prior to be corrected. In this paper, we propose LookAgain, a closed-loop GUI grounder driven by post-prediction visual reflection. LookAgain reformulates grounding as a multi-turn predict-look-again-refine process with two primitives: “locate” posts a coordinate hypothesis, renders a marker on the image and appends a local patch of the predicted region. It anchors the next reasoning step to the previous prediction as a spatial prior; “confirm” accepts or reject the hypothesis and terminates the procedure. We train the LookAgain grounder with SFT on constructed reflective trajectories as a cold start, followed by GRPO with terminal grounding correctness as the sole reward. Extensive experiments show that LookAgain consistently improves performance on both refusal-aware and general GUI grounding benchmarks, achieving state-of-the-art results. Comprehensive ablations further verify the effectiveness of the proposed framework.
[193] Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes cs.CVPDF
Mario Malizia, Marnix Enting, Rob Haelterman, Ken Hasselmann
TL;DR: 本文提出了一种从少量未标记的现场照片中合成带标签训练图像的方法,以解决植被中隐藏的小物体检测数据稀缺和跨站点泛化差的问题。该方法利用视觉语言模型构建粗略的3D植被场景,并放置3D物体网格来生成标注,再通过轻量适配器和扩散模型进行纹理重绘,同时通过分级掩码锁定控制对物体本身的修改程度。
Details
Motivation: 解决在植被中隐藏的小物体(如人道主义排雷中的地雷)标注图像稀缺,以及基于其他站点标注训练的检测器在新部署站点上泛化性能差的问题。
Result: 在一个人道主义排雷基准测试上,使用该方法合成的图像训练的标准检测器,其性能匹配或超过了使用来自不同站点的更大规模真实标注数据集训练的检测器,实现了无监督的站点自适应。消融实验表明,该方法对照片和裁剪预算不敏感,且域内精度不能预测跨站点性能。
Insight: 核心创新在于利用VLM和3D场景几何自动生成标注,并通过扩散模型进行可控的纹理重绘以适应目标域。关键发现是分级掩码锁定(即控制对物体像素的扩散程度)是影响性能的最重要因素,轻微扩散物体本身比完全保护其像素更能提高少数类召回率。
Abstract: Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing labels collected at another site, we synthesize labeled training images from a handful of unlabeled photographs of the deployment site itself. A vision–language model generates a coarse 3D vegetation scene from one photograph; placing 3D object meshes in the scene yields bounding boxes, segmentation masks, and per-instance occlusion directly from the scene geometry, without manual annotation. A lightweight adapter fine-tuned on the photographs conditions a diffusion pass that re-textures the renders, and a graded mask-lock sets how much diffusion may touch the object itself. In our runs this grade was the most influential curation choice: lightly diffusing the object improves minority-class recall over fully protecting its pixels, while unrestricted diffusion dissolves it. Trained on these images, a standard detector matched or exceeded its counterpart trained on a larger labeled dataset of real images from a different site, consistently across seeds on a humanitarian-demining benchmark; the comparison is thus unsupervised site adaptation from a handful of photographs against conventional cross-site label reuse. In our ablations the gains were largely insensitive to the photograph and crop budgets, and in-domain accuracy did not predict cross-site performance.
[194] ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection cs.CVPDF
Jingtai He, Shiyuan Meng, Wenchao Meng, Qinmin Yang
TL;DR: 本文提出ADOPD框架,一种基于参考特权策略蒸馏的方法,用于提升多模态大语言模型在工业异常检测任务中的性能。该方法通过让教师模型在匹配与不匹配参考条件下评估学生模型生成的序列,定义令牌级学习方向和序列级权重,从而将参考比较的优势内化到模型参数中,实现无需额外检索的零样本推理。
Details
Motivation: 工业异常检测需要识别与正常视觉模式的细微偏差。现有基于MLLM的方法在推理时通过比较查询图像与参考图像可提高准确率,但依赖额外的检索和处理开销。本文旨在探索是否能够将参考比较的优势内化到模型参数中,使模型在仅输入查询图像时也能达到类似效果。
Result: 在MMAD基准测试上,ADOPD在零样本推理设置下达到了77.31%的平均准确率,相比其骨干模型Qwen3-VL-4B提升了6.14个百分点,并且比该模型在单样本设置下的性能还高出2.64个百分点,实现了新的SOTA水平。
Insight: 核心创新在于提出了参考特权策略蒸馏框架,通过教师模型在匹配/不匹配参考视图下的似然差异来校准学习信号,使学生模型能够从参考比较中学习到细粒度的异常检测策略,而无需在推理时访问参考图像。
Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) can improve recognition accuracy by comparing query images with references at inference time, but these benefits rely on additional retrieval and processing. We investigate whether the benefits of reference comparison can instead be internalized in the model parameters. Access to references during training allows a reference-aware teacher to supervise a query-only student. However, the teacher may favor plausible responses based on query cues or language priors rather than valid visual information. We propose ADOPD, a reference-privileged on-policy distillation framework. The teacher evaluates student-generated rollouts under matched and mismatched references. The matched-reference teacher-to-student log-ratio defines the token-level learning direction, specifying what the student should learn. The likelihood gap between the two reference views estimates reference-specific support and calibrates the sequence-level weight. ADOPD achieves 77.31% average accuracy on the MMAD benchmark under zero-shot inference, improving the Qwen3-VL-4B backbone by 6.14 points and outperforming its one-shot setting by 2.64 points. Experiments show that ADOPD learns a fine-grained anomaly inspection strategy from reference comparison. The project will be available at https://github.com/withTai/ADOPD.
[195] World Tokens: Enhancing Embodied Policies with Training-Time World Modeling cs.CVPDF
Qu Tang, Benhui Zhuang, Bo Yuan, Xue Yu, Longteng Guo
TL;DR: 本文提出了一种名为World Tokens的具身策略架构,其核心是一个World Adapter模块,旨在桥接视觉语言理解、世界动态建模和动作生成。该方法在训练时利用世界建模来增强动作策略,但在部署时移除世界模型分支,从而保持高效的推理速度。
Details
Motivation: 现有的视觉-语言-动作(VLA)模型擅长闭环控制,但未显式建模物理场景的动态演变;而新兴的世界-动作模型(WAMs)虽能捕捉时空演化,但推理成本高昂。本文旨在设计一种既能利用世界建模优势,又能保持高效部署的具身策略。
Result: 该方法在LIBERO基准上极具竞争力,在SIMPLER基准上取得了最佳报告平均性能,在真实世界R1 Pro任务上显著优于仅动作的基线模型,并且能以VLA级别的延迟生成每个动作块。
Insight: 创新点在于通过World Adapter将VLM特征转换为固定数量的世界令牌,这些令牌同时作为未来视频去噪器和动作专家的条件,使得来自视频去噪的梯度能直接塑造用于动作预测的表征,而部署时无需在线视频模型推理,实现了训练时世界建模与部署时高效性的解耦。
Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert’s sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.
[196] MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation cs.CV | cs.AIPDF
Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng
TL;DR: 本文提出了MedPixel,一个统一的医学像素-语言模型,通过共享的语言-掩码接口连接临床语言与视觉推理。模型使用MedPLG-440K数据集进行训练,结合多任务监督微调和像素级偏好优化,支持显式定位、隐式推理、空间交互、基于掩码的解释和医学VQA等多种任务。
Details
Motivation: 现有医学视觉语言模型缺乏精确定位能力,而医学分割器通常依赖明确的类别或精确的空间提示,两者之间存在监督不匹配的问题。本文旨在弥合这一鸿沟,实现语言与像素级定位的统一。
Result: MedPixel在像素级预测和响应生成方面均表现出色,在多种任务上实现了强大性能,并在外部定位基准测试中展现出有效的零样本迁移能力以及对不完善空间提示的鲁棒性。
Insight: 创新点在于提出了统一的像素-语言模型架构和共享的语言-掩码接口,并引入了通过临床动机合成构建的大规模像素-语言数据集MedPLG-440K,以及使用真实掩码作为离线验证器进行像素级偏好优化的训练方法。
Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language–mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.
[197] Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization cs.CV | cs.AIPDF
Dinh Tan Nguyen, Quang-Hien Kha, Le-Hoang Nguyen, Minh-Toan Dinh, Xuan-Huy Nguyen
TL;DR: 本研究探讨了在现代骨干网络支持下,多任务DETR框架在乳腺X光摄影中联合进行影像级恶性预测和病灶定位的性能。通过在OPTIMAM和SGM1k数据集上的评估,发现ConvNeXtV2和DINOv3等现代骨干网络显著优于传统ResNet风格特征,其中ConvNeXtV2在OPTIMAM上表现最佳,DINOv3在SGM1k上表现最强。
Details
Motivation: 旨在通过联合影像级预测和候选区域定位,提升AI在乳腺X光摄影中的实用性,研究共享表征如何同时支持恶性预测和病灶定位任务。
Result: 在OPTIMAM数据集上,ConvNeXtV2达到97.96% AUC、99.89%灵敏度、25.08% mAP@.5和74.38% recall@.25;在SGM1k数据集上,DINOv3达到90.97% AUC、86.28%灵敏度、82.00%特异度、27.04% mAP@.5和77.32% recall@.25,均展现了SOTA或强竞争力性能。
Insight: 论文宣称的创新点在于系统评估了现代骨干网络(如ConvNeXtV2、DINOv3)在多任务DETR框架下的有效性,客观分析表明骨干网络质量是影响多任务乳腺X光摄影分析性能的关键因素,ConvNeXtV2被证明是该框架下特别匹配且强大的CNN骨干。
Abstract: Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet-style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best overall performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% mAP@.5, and 74.38% recall@.25. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% mAP@.5, and 77.32% recall@.25. These findings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework.
[198] Financial Numerical Prediction and Allocation as Token Generation cs.CV | cs.LGPDF
Xu Ouyang, Moontae Lee
TL;DR: FinATOM提出了一种无需特定任务头的统一方法,通过受限的token生成直接进行金融数值预测和资产配置。该方法将股票收益预测建模为自回归生成波动率标准化收益token,并将动态ETF配置建模为生成归一化的多头权重。
Details
Motivation: 传统金融预测方法依赖特定任务头(如回归、排序或策略头),将语言模型与最终评估的数值对象分离。本文研究因果语言模型是否可以通过受限的token生成直接表示预测和决策。
Result: 在2023-2025年ETF测试中,配置策略将合并总夏普比率从1.428提升至1.529,在5个基点交易成本模型下净夏普比率从1.394提升至1.494。多模态配置输入实现了最高的三期平均夏普比率1.540。在FinTexTS上,监督微调和策略策略分别实现了73.52%/2.68和73.72%/2.69的累计收益/夏普比率。
Insight: 创新点在于提出了一个无头的统一接口,将金融预测和决策直接转化为语言模型的token生成任务,并通过结合序数/排序监督与token级策略阶段(如DAPO增强的GRPO)进行训练。这为使用语言模型直接处理金融数值任务提供了可行性验证。
Abstract: Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ultimately evaluated. We investigate whether a causal language model can instead represent forecasts and decisions directly through constrained token generation. FinATOM introduces a unified, head-free interface for three-step stock-return forecasting and dynamic five-ETF allocation. The forecasting model autoregressively emits volatility-standardized return tokens and is trained with ordinal and ranking supervision followed by a one-epoch token-level policy stage. The allocation model generates normalized long-only weights; supervised fine-tuning imitates a causal mean–variance anchor, and DAPO-augmented GRPO optimizes realized 21-day Sharpe subject to anchor consistency. In 2023–2025 ETF tests, the allocation policy improves pooled gross Sharpe from 1.428 to 1.529 and net Sharpe under a 5-bp transaction-cost model from 1.394 to 1.494. The multimodal allocation input attains the highest three-period mean Sharpe of 1.540, with its clearest advantage in 2025. On FinTexTS, the SFT and policy strategies achieve 73.52%/2.68 and 73.72%/2.69 cumulative-return/Sharpe, respectively. These results support the feasibility of direct language-model token generation for financial numerical prediction and decision-making, while motivating broader tests across assets, regimes, and random seeds.
[199] Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football cs.CV | cs.CY | cs.LGPDF
Seongjin Choi
TL;DR: 该论文提出了一种新的足球比赛分析框架,用于评估无球阶段的控球质量,区分创造空间的控球与无效的控球。它包含两个层面:一个基于事件的‘垃圾控球指数’来识别低威胁的控球序列,以及一个基于视频分析的‘空间创造指数’来量化控球是否实际创造了空间。
Details
Motivation: 现有基于事件的控球价值框架(如预期威胁、VAEP)只评估有球动作,忽略了无球阶段控球的核心问题:控球是创造了空间,还是无效的循环传递?论文旨在解决这一盲点。
Result: 在2026年世界杯的103场比赛数据上,‘垃圾控球指数’与球队积分(r=-0.37)和预期进球差(r=-0.51)呈负相关,且在控制有球动作价值后仍能提供额外信息。在35个被标记的控球时段中,视频分析显示74%为未创造空间,仅6%实际创造了空间。
Insight: 创新点在于将事件数据与视频分析结合,构建了一个两层的评估体系,首次系统性地量化了无球控球在空间创造方面的质量,弥补了现有模型仅关注有球动作的不足,为战术分析提供了新维度。
Abstract: Ball possession is the most-cited and most-misleading number in football: 60% recycled in one’s own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP, on-ball value) price on-ball actions but ignore the off-ball question a sterile possession poses: did holding the ball create space, or was the circulation dead? We answer this in two layers. First, an event-side junk-possession index prices each possession sequence by its peak threat gain under an expected-threat grid and – after reconstructing the live scoreline to exclude lead-protecting circulation – flags low-threat sequences in tied-or-losing states. On the 2026 FIFA World Cup (103 matches, 206 team-matches) the flag correlates negatively with points (r=-0.37) and xG difference (r=-0.51, partly index-coupled). It is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the junk flag stays strongly negatively associated with points (p<0.0001, also match-clustered) while VAEP is not significant – in this same-match (descriptive) regression it adds information beyond this on-ball action-value model. Second, for a flagged window we resolve whether it was spatially dead or space-creating by projecting broadcast video to pitch coordinates and measuring a Space-Creation Index (SCI): a net pitch-control change capturing whether the possession seized space or pushed the opponent’s block back. Across 31 of 35 flagged windows from nine World Cup matches (a purposive sample), 74% are spatially non-space-creating, 19% weak progression, and 6% space-creating windows the event flag alone would score as failure – including a side with 73% of the ball that exited on penalties (two non-creating windows). The two layers separate space-creating-but-unconverted from sterile possession, a distinction event-only on-ball value cannot make.
[200] From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing cs.CVPDF
Jutao Xiao, Yuan Qu, Dongsheng Ma, Fan Wu, Tianyao He
TL;DR: 本文针对现有文档解析器在复杂真实世界表格解析上的性能不足问题,提出了TableParseMap诊断基准和DEC(Decompose–Enhance–Correct)框架。TableParseMap包含916个真实表格,揭示了当前最优解析器TEDS分数仅为85.03,与现有基准分数存在显著差距。DEC框架利用通用视觉语言模型作为控制器,通过分解、增强和纠正三个步骤,在不重新训练的情况下提升冻结解析器的性能,在TableParseMap基准上实现了显著改进。
Details
Motivation: 现有文档解析器在标准基准(如OmniDocBench v1.6)上TEDS分数超过93,但社区反馈和作者审计发现,它们在复杂真实世界表格上仍存在持续失败。为了量化这一差距并解决解析器在大型表格、弱视觉线索和视觉不一致性方面的局限性,研究旨在开发一个诊断性基准和改进框架。
Result: 在提出的TableParseMap基准上,评估的最强解析器仅达到85.03 TEDS。DEC框架在三个冻结解析器上平均提升TEDS 1.57点;在TableParseMap上,整体提升1.89点,在结构错误上提升2.62点,在大型表格上提升5.66点。
Insight: 创新点包括:1) 引入TableParseMap诊断基准,系统组织真实世界表格的挑战场景和失败类型,揭示聚合分数掩盖的弱点;2) 提出DEC代理框架,利用通用VLM作为控制器,通过分解、增强和纠正步骤,以视觉一致性为指导改进冻结解析器,无需重新训练;3) 设计Visual Consistency Gate和Ranker,在推理时无需真实HTML的情况下选择性干预和验证更新,支持回滚。从客观角度看,该研究强调了评估基准需更贴近真实复杂场景,并展示了如何通过智能后处理框架弥补现有模型的固有缺陷。
Abstract: Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challenging scenarios and nine failure types. The strongest evaluated parser achieves only 85.03 TEDS, showing that aggregate benchmark scores conceal substantial weaknesses. Our analysis attributes these failures to three complementary limitations: large tables exceed the reliable processing scale of a single pass, weak or ambiguous visual cues hinder structure perception, and the reconstructed table may remain visually inconsistent with the image. We therefore propose DEC (Decompose–Enhance–Correct), a visual-consistency-guided agentic framework that improves frozen table parsers without retraining. DEC uses a general VLM as the controller: Decompose partitions large tables along structure-aware boundaries, Enhance exposes weak visual evidence and reparses transformed views, and Correct diagnoses and repairs residual errors. A Visual Consistency Gate (VC-Gate) selectively triggers intervention, while a Visual Consistency Ranker (VC-Ranker) verifies candidate updates and supports rollback without ground-truth HTML at inference time. We further derive a 1,977-table Consensus-Hard Set from 4,556 candidates through offline metrics and cross-model consensus. Across three frozen parsers, DEC improves TEDS by 1.57 points on average; on TableParseMap, gains reach 1.89 points overall, 2.62 on structural errors, and 5.66 on large tables.
[201] Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains cs.CV | cs.AIPDF
Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
TL;DR: 本文提出了Sci-VBench,一个用于评估科学领域中知识密集型和推理密集型视频生成的综合基准。该基准包含1,253个专家标注的示例,涵盖自然科学、医疗保健、人文社会科学和工程学四个核心学科的60个主题,要求模型生成具有科学推理和知识基础合成的时序丰富视频。研究还建立了一个基于量规的评估协议,并发现非专业人类评估员和MLLM-as-Judge系统都能与专家判断达成较高一致性。通过对16个前沿专有和开源模型的评测,研究发现尽管自动感知质量得分相近,但在提示基础和科学因果正确性方面存在显著差异,揭示了视觉真实性的进步尚未转化为对科学和因果动态的可靠建模。
Details
Motivation: 当前视频生成模型主要关注表面视觉逼真度,缺乏对科学领域知识密集和推理密集型任务的评估。为了解决这一问题,需要建立一个专门的基准来评估模型在科学视频生成中的知识基础和推理能力。
Result: 在Sci-VBench基准上评估了16个前沿专有和开源模型。结果显示,自动感知质量得分(如FVD)在不同模型间差异不大,但在Prompt Grounding和Scientific and Causal Correctness指标上表现差异显著,且专有模型与开源模型之间存在明显差距。评估协议验证了非专业人类评估员和MLLM-as-Judge系统能与专家判断达成较高一致性。
Insight: 论文的创新点在于构建了首个专注于科学领域知识推理的视频生成基准,并提出了一个可扩展的、基于量规的评估协议。从客观角度看,该工作揭示了当前视频生成模型在科学正确性和因果推理方面的不足,强调了超越视觉质量、评估模型深层理解能力的重要性,为未来研究指明了方向。
Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
[202] Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning cs.CVPDF
Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji
TL;DR: 本文提出了一种名为潜在动态推理(LDR)的视频世界模型,旨在从像素中学习并准确模拟世界的物理动力学规律。该方法将潜在状态转换建模为显式的运动学积分,在结构化潜在空间中进行计算,从而显著提升了模型在分布外场景下的动态外推能力。
Details
Motivation: 当前主流的视频扩散模型主要拟合像素外观,而没有显式建模像素随时间变化的动态规律,导致生成的视频帧可能在视觉上合理但不符合物理定律。本文旨在从像素中直接学习并准确捕获这些底层动力学。
Result: 在PhyWorld基准测试的五个任务(匀速运动、抛物线、碰撞、反弹、迫近)上,LDR在分布外场景下的性能远超基线视频扩散模型,其分布内与分布外误差差距比基线小20倍以上,同时参数量减少26倍,运行速度快143倍。
Insight: 核心创新在于将潜在状态转换显式地建模为运动学积分过程,并在结构化潜在空间而非密集卷积特征上进行计算,这使得模型能够更好地外推学习到的动力学规律,首次实现了视频世界模型在训练分布之外的动态外推。
Abstract: The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20$\times$ smaller than the video diffusion baseline’s, under both single- and joint-task training at 256$^2$ resolution, while using 26$\times$ fewer parameters and running 143$\times$ faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first video world model that extrapolates learned dynamics beyond its training distribution. Project page: https://lat-dyn-reason.github.io/
[203] DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning cs.CVPDF
Mainak Singha, Niccolò Biondi, Elisa Ricci, Subhankar Roy
TL;DR: 本文提出DistMoE,一种用于分布式视觉指令调优的混合专家模型方法。它在语言解码器的每一层中,用一个客户端特定的私有前馈网络专家来增强公共前馈网络,旨在获取领域特定知识。通过引入一个以公共专家为锚点的专家组合阶段,仅更新路由器和轻量级私有投影适配器,减少了客户端特定的偏移,实现了无需跨客户端排练的组合。在推理时,DistMoE在公共和私有专家上进行模块化路由,实现了无需显式领域标签的token级领域组合。
Details
Motivation: 多模态大语言模型展现出强大的多模态指令跟随能力,但适应不同视觉-语言领域通常假设集中式数据访问和昂贵的联合训练。当数据分布在私有、领域特定或权限受限的客户端时,这种假设具有限制性。因此,需要一种方法在保护数据隐私和权限的同时,实现有效的领域适应。
Result: 在多个视觉-语言基准测试上的实验表明,DistMoE能够实现灵活的专家重用、有效的领域适应,并在保持对客户端特定知识的模块化控制的同时,取得了有竞争力的性能。
Insight: 主要创新点在于提出了一个无需排练的、以公共专家为锚点的专家组合训练阶段,通过各向同性正则化损失来对齐表示尺度,从而解决了独立专家训练导致的表示漂移问题。这使得模型能够在无需跨客户端数据共享(排练)的情况下,实现模块化的、token级的领域知识组合与路由,为分布式、隐私敏感场景下的模型适应提供了新思路。
Abstract: Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it cross-client rehearsal-free composition. During inference, DistMoE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge. Codes are available at https://github.com/mainaksingha01/DistMoE.
[204] Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection cs.CVPDF
Wenti Yin, Xiang Wang, Huaxin Zhang, Hanqing Wang, Hongbo Shao
TL;DR: 本文提出了一种无需训练的对比事件裁决方法(CEAVAD)用于视频异常检测。该方法通过构建危险-正常事件对比对,将推理单元从孤立的异常概念转变为可证伪的事件假设,并在竞争性解释与视频证据的交互中建立推理时的解释边界,从而支持时序定位的异常检测和基于证据的解释。
Details
Motivation: 现有无需训练的方法虽然利用了预训练模型的丰富语义知识和推理能力来解释视觉内容,但这些能力并不能直接定义异常决策标准:更丰富的异常描述能更好地捕捉危险相似性,但无法解决异常性判定问题。
Result: 在三个广泛使用的视频异常检测基准测试上的实验表明,CEAVAD在无需训练范式下实现了最先进的性能。
Insight: 核心创新在于将推理单元从概念层面提升到事件假设层面,并引入可证伪的对比裁决机制。该方法通过构建机制特定的危险-良性对比对,并让它们与视频证据竞争,从而在推理时动态形成解释边界,这为基于大模型的零样本或无需训练异常检测提供了新的范式。
Abstract: Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality. To this end, we propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing explanations and video evidence. Specifically, CEAVAD first uses public-safety knowledge to construct hazard-benign event contrasts, pairing each hazard mechanism with a generic normal account and a mechanism-specific benign counterpart. It then determines whether the target interval better supports a hazard explanation or its benign competitor, yielding a revisable contrastive boundary proposal for the target. Finally, CEAVAD adjudicates between the competing explanations to determine whether the hazard hypothesis survives the video evidence, supporting both temporally localized anomaly detection and evidence-grounded explanations. Experiments on three widely used VAD benchmarks demonstrate that CEAVAD achieves state-of-the-art performance under the training-free paradigm.
[205] Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots cs.CVPDF
Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar, Abdelrahman Shaker, Rao Muhammad Anwer
TL;DR: 本文提出了CVPD(对比反事实视觉过程蒸馏),这是首个完全自包含的、用于多模态大语言模型(MLLMs)的密集、同策略、令牌级视觉自蒸馏框架。该方法通过识别模型自身的视觉盲点(即放大特定区域会改变答案分布,而移除该区域则不影响全图行为),并利用反事实准则将其转化为对比监督信号进行自我蒸馏,从而提升模型性能。
Details
Motivation: 现有MLLMs的自改进方法通常依赖基于奖励的粗粒度反馈,而视觉领域的蒸馏方法通常需要依赖外部标注、工具或更强模型构建的上下文。本文旨在提出一种完全自包含的密集视觉自蒸馏框架,不依赖外部资源。
Result: 在Qwen3-VL-8B-Instruct模型上,CVPD在十二个基准测试中超越了六个自演化基线方法(包括依赖外部GPT-4o监督的方法),且没有性能回退。具体在OCRBench上提升+3.60,MMStar细粒度感知上提升+3.38,MMStar逻辑推理上提升+3.08,同时在更广泛的多模态基准上保持或提升了性能。
Insight: 核心创新在于提出了一个完全自包含的视觉自蒸馏框架,其关键是通过模型自身响应识别“反事实视觉盲点”作为监督信号,这避免了对外部资源的依赖。从客观角度看,将模型内部不一致的感知区域(盲点)转化为对比学习目标,是一种新颖且高效的自我监督信号挖掘机制。
Abstract: Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model’s answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model’s own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of $+3.60$ on OCRBench, $+3.38$ on MMStar Fine-Grained Perception, and $+3.08$ on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.
cs.HC [Back]
[206] The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction cs.HC | cs.CL | stat.APPDF
Tyler Dooskin, Squoosh Technical Staff
TL;DR: 这篇论文研究了使用多模态大语言模型(LLM)仅通过网页截图来预测真实A/B测试结果的能力。经过六周的预注册实验,主要结论是:模型在大多数情况下无法可靠预测,但模型自信的预测(通过投票边际筛选)能识别出一个子集,在该子集上预测与统计显著的真实结果有中等一致性。研究发现,无论是模型还是人类专家,其内部共识很高,但与真实结果的关联性却很弱,表明共识本身并非有效证据。
Details
Motivation: 动机是探究多模态LLM是否能够仅凭网页截图,可靠地预测真实世界A/B测试(例如网页版本转化率)的胜出方,以替代或辅助成本高昂的真实测试。
Result: 在330个真实A/B测试上,Gemini 3 Flash模型的预测与“地面真值”的总体一致性较低(Cohen’s kappa = 0.14)。在统计显著的测试子集上,证据仍不充分(kappa = 0.11,置信区间包含零)。然而,通过筛选模型自信的预测(投票边际门控,覆盖49%的测试),在显著标签上达到kappa = 0.31。所有标准改进方法(如使用更昂贵的模型、优化提示)均未通过预注册检验。人类专家组内部一致性为kappa=0.53,但对真实结果的预测也处于随机水平。
Insight: 核心创新点在于系统性地揭示了LLM(及人类)在A/B测试预测任务中的局限性:模型/专家内部的高共识(可复现性、说服力强)并不等同于预测准确性,这挑战了仅依赖共识作为证据的常见做法。方法论上的创新包括严格的预注册实验、对“地面真值”数据质量的批判性分析(指出44%的标签来自非显著测试)、以及引入“投票边际门控”机制来识别模型可能可靠的预测子集,并量化了模型投票的有效独立性。论文强调了校准性弃权(即知道何时不知道)的重要性,并提供了完整的负结果和证据层级记录。
Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no – and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen’s kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the “ground-truth” labels in a leading CRO agency’s catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones – a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge’s confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly – judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only ~0.2, so a 16-vote panel carries about 2 effective independent votes – and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa ~ 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.
[207] How People Evaluate AI-, Expert-, and Peer-Style Financial Advice cs.HC | cs.AI | cs.CL | cs.CY | cs.SIPDF
Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
TL;DR: 本研究通过一项预注册的随机对照实验(N=285),探究了人们在评估金融建议时如何受到建议来源(AI金融助手、认证理财规划师、在线社区论坛)和来源标签(正确标注、未标注、错误标注)的影响。研究发现,在金融建议内容完全一致的情况下,专家建议在多数评估指标上优于AI建议,且这一优势在无标签条件下依然存在;错误标注AI建议为专家建议能提升其情境适配性和整体质量评分,同时削弱专家建议的优势。
Details
Motivation: 随着生成式AI日益成为日常决策(包括金融选择)的常见信息来源,理解人们如何评估AI生成的金融建议变得至关重要。
Result: 在10项评估结果中,专家建议有9项优于AI建议(效应量|d|=0.20–0.47);即使在无来源标签条件下,专家建议仍在8项结果上优于AI建议(最大d=0.60)。错误标注使AI建议的情境适配性和整体质量评分显著提升(d=0.42),并削弱了专家建议在情境适配性上的优势(d=-0.36)。
Insight: 研究揭示了金融建议评估同时受到显示来源标签和信息层面沟通线索的联合影响。来源披露并非中立的透明度机制,而是一个解释性框架,其准确性以及与信息线索的交互会影响信任和依赖。AI建议对显示标签最为敏感,而建议风格的差异在AI标签下最为明显。
Abstract: As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content—including facts, numerical values, recommendation direction, and core reasoning—was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20–0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
[208] A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data cs.HC | cs.CVPDF
Joseph Bingham
TL;DR: 本文提出了一种动态语义框架,用于将人类指代表达式与视觉感知数据关联起来。该框架通过将协议状态外化为三个显式的指称-对象绑定集合,并结合基于SIFT单应性和通用质量指数的轻量级感知对齐流程,解决了现有视觉语言模型在词汇顺应性方面的不足。在斯坦福重复指称游戏语料库上的评估显示,该框架在单次导演话语中能将正确目标纳入其前5假设集的准确率达到83.56%。
Details
Motivation: 现有视觉语言模型在词汇顺应性方面存在缺陷,无法像人类一样通过重复交互收敛于共享的指代表达式。本文旨在通过显式建模协议状态和动态语义更新规则,填补这一能力差距。
Result: 在斯坦福重复指称游戏语料库(包含超过15,000条抽象七巧板刺激的导演-匹配者话语)上,框架从单次导演话语中前5假设集准确率为83.56%,接近人类匹配者约77-80%的top-1准确率。消融实验量化了SIFT对齐、UQI、查询预处理和图像增强等组件的贡献。
Insight: 创新点在于结合了透明的符号层(动态语义绑定更新)与可解释的感知通道(基于SIFT和UQI),实现了词汇顺应性的结构化恢复。框架强调可审计性,但未实现交互闭环,且感知检索存在泄漏效应需量化处理。
Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents that they do not shorten references, reuse successful expressions, or maintain stable pact state across turns. We present a framework that addresses the gap by externalizing pact state into three explicit, inspectable sets of referent-object bindings ($Γ, Ξ, Ω$), updated by a dynamic-semantics context-change rule. The symbolic layer sits on top of a lightweight perceptual-alignment pipeline that grounds noisy human referring expressions in crowd-sourced imagery via SIFT homographies and the Universal Quality Index. Evaluated on the Stanford Repeated Reference Game corpus (over 15{,}000 director-matcher utterances on abstract tangram stimuli), the framework places the correct target in its top-5 hypothesis set 83.56% of the time from a single director utterance. Human matcher top-1 accuracy on the same corpus is approximately 77-80%. We also report results on a held-out condition in which obvious tangram-adjacent images are excluded from the retrieved set, which provides a more conservative measurement of the grounding signal. Ablations isolate the contribution of each component: SIFT alignment, UQI, query preprocessing, and image augmentation. The central contribution is the combination: a transparent, auditable symbolic layer that recovers the structure of lexical entrainment turn by turn, paired with a perceptual channel whose behavior can be examined ablation by ablation. We also discuss in detail what the framework does not do. It is not interactive, it does not close the loop with the director, and its retrieval-driven perceptual channel is vulnerable to a class of leakage effects that we quantify and bound rather than wave away.
cs.AI [Back]
[209] The Knowing-Saying Gap: When Probes See Errors that Confidence Misses cs.AI | cs.CL | cs.LGPDF
Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
TL;DR: 该论文探讨了语言模型中线性探针检测上下文错误的能力与模型实际错误预测之间的脱节现象。研究发现,探针能近乎完美地检测上下文损坏,但无法可靠预测最终答案的正确性,且基于探针的干预措施效果高度依赖模型和错误类型。
Details
Motivation: 解决语言模型部署监控中,探针检测能力与模型置信度表达之间的不一致问题,即模型’知道但不说出’错误,影响可靠性评估。
Result: 在多跳算术链任务中,探针检测与答案正确性无关;结构化置信度格式崩溃为两个值且错误率无法区分;探针持久性假设被证伪。干预措施如分支选择在Llama-3.1-8B上表现最佳(救援4例,破坏0例),而其他方法救援与破坏率相近。
Insight: 创新点在于揭示了探针监控与语言化置信度的互补必要性,但无单一干预措施普适;实际部署需基于模型感知和错误类型感知的路由策略,强调监控系统的情境依赖性。
Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered “persistence beats peak” hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.
[210] IntelliAudit: Using Large Language Models to Evaluate Audit Controls cs.AI | cs.CL | cs.CR | cs.HC | cs.MAPDF
Allison Wilson, Sina Moradi Sabet, Diar Shakimov, Panteha Shahrivar, Mohammad Reza Bagheri
TL;DR: 本文提出了IntelliAudit,一个基于检索的多智能体系统,用于自动化评估IT审计中的安全与合规控制。该系统能够从异构证据中检索相关信息,生成基于证据的评估,挑战不利发现,裁决分歧,并为审计员提供包含引用证据、理由、缺失证据分析和补救建议的推荐。
Details
Motivation: IT审计中,审计员需要判断分散在不同策略、记录、电子表格和操作工件中的异构证据是否满足语义上的安全与合规控制要求,这一过程难以自动化,因为审计结论依赖于证据的充分性而非简单的关键词匹配。
Result: 该系统在ISO/IEC 27001标准上进行了实例化,并通过专家审计员评审和审计准备用户反馈在多个模拟组织中进行评估。评估表明,IntelliAudit能够支持控制解释、基于证据的推理和审计准备工作流,但也揭示了人工监督对于校准充分性判断和纠正过于宽松建议的重要性。
Insight: 论文的创新点在于将基于检索的多智能体系统架构应用于复杂的、语义驱动的IT审计证据评估任务,实现了从证据检索到生成可解释建议的端到端流程。客观来看,其核心洞察是这类系统应作为决策支持工具而非自主认证系统,强调了人机协同在专业判断任务中的必要性。
Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational artifacts, and because audit conclusions depend on evidentiary sufficiency rather than keyword matching. We present IntelliAudit, a retrieval-grounded multi-agent system for IT audit evidence evaluation. Given a control and an evidence corpus, IntelliAudit retrieves relevant artifacts, generates an evidence-grounded assessment, challenges adverse findings, adjudicates disagreements, and produces an auditor-facing recommendation with cited evidence, rationale, missing-evidence analysis, and remediation guidance. We instantiate IntelliAudit on ISO/IEC 27001 and evaluate it across multiple simulated organizations using expert auditor review and audit-readiness user feedback. The evaluation shows that IntelliAudit can support control interpretation, evidence-grounded reasoning, and audit-preparation workflows, while also revealing the importance of human oversight for calibrating sufficiency judgments and correcting overly permissive recommendations. These results suggest that retrieval-grounded multi-agent systems can assist audit evidence review, but should remain decision-support tools rather than autonomous certification systems.
[211] GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering cs.AI | cs.CL | cs.IT | cs.LGPDF
Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang
TL;DR: 本文提出了GRACE框架,用于解决混合表格数据(包含连续数值和离散类别)的聚类问题。该框架通过多视角LLM查询策略,在属性-值级别获取外部语义知识,将异构值映射为知识驱动的描述,从而构建统一的语义度量空间,并一次性提取通用语义表示,避免了将LLM嵌入迭代优化循环带来的巨大计算开销。
Details
Motivation: 传统混合数据聚类算法完全依赖数据集内部统计来估计类别关系,这限制了学习到的度量仅基于经验共现,而忽略了概念上明显但统计上未观察到的关联。虽然LLM提供了外部世界知识,但将其以文本为中心的推理应用于高度抽象的表格概念存在显著挑战,且将LLM嵌入迭代度量学习循环会导致难以承受的计算开销,迫使在语义丰富性和可扩展性之间做出妥协。
Result: GRACE在可扩展性上与传统统计驱动基线方法相当,同时在11个竞争方法中实现了更优的聚类准确性和概念可解释性。
Insight: 核心创新点在于将语义获取从数据集级别转移到属性-值级别,通过一次性LLM查询提取通用语义表示,从而将昂贵的LLM调用与迭代优化解耦。此外,框架通过将外部语义与数据集内部统计证据进行交叉验证,确保其与特定数据集的聚类结构保持一致,兼顾了外部知识引入和内部数据适配。
Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines the learned metric to empirical co-occurrences and ignores conceptually obvious yet statistically unobserved affinities. Although LLMs offer external world knowledge, applying their text-centric reasoning to highly abstract tabular concepts presents significant challenges. Bridging this modality gap to construct a semantically complete metric typically requires embedding LLMs into iterative metric learning loops to dynamically optimize cross-modality representations. This incurs intractable computational overhead, forcing a compromise between semantic enrichment and scalability. Therefore, we propose GRACE, an LLM-grounded framework for scalable mixed-data clustering. GRACE shifts semantic acquisition to the attribute-value level via a multi-perspective LLM querying strategy, mapping heterogeneous values into knowledge-informed descriptions. Crucially, this one-shot grounding extracts general-purpose semantic representations that embed heterogeneous attributes into a unified space, decoupling expensive LLM invocation from iterative optimization. Furthermore, GRACE cross-validates these external semantics against dataset-internal statistical evidence to ensure alignment with the dataset-specific cluster structure. Ultimately, GRACE matches the scalability of conventional statistics-driven baselines while achieving superior clustering accuracy and conceptual interpretability over 11 competing methods. The source code is available at https://github.com/develop-yang/GRACE-GRACE-A
[212] Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE cs.AI | cs.CLPDF
Ramakrishna P. Kompella, Aadit Mahajan
TL;DR: 本文研究了多语言混合专家模型中的安全对齐不一致问题,即模型在英语中会拒绝有害请求,但在低资源语言中却可能遵从。通过机制性地追踪印度语多语言模型sarvam,作者发现安全检测本身是有效的,但拒绝行为的实际生成(写入)是由一个特定的、可定位的混合专家写入电路完成的,该电路受一个注意力反对者的抑制。作者量化了干预该电路的不同方式(如抑制反对者、放大写入者)的成本和效果。
Details
Motivation: 解决多语言模型安全对齐的不一致性问题,即模型在低资源语言中比在高资源语言中更容易遵从有害请求,并探究其背后的机制性原因。
Result: 在sarvam模型中发现,有害检测方向在模型中层(L11)近乎语言不变(英语与印度语余弦相似度≈0.9),但实际写入拒绝的电路位于下游。干预该电路的实验表明:抑制注意力反对者成本低且有效,放大写入者成本高昂,而对外部头进行手术式编辑无效。该电路的组织结构和揭示它的梯度方法在另一个无关的MoE模型中也复现了。
Insight: 创新点在于将多语言安全失败机制性地定位并定价为一个特定的、可干预的写入电路,而非检测失败。这为多语言安全修复提供了成本量化的干预地图,并揭示了安全决策(上游检测)与安全执行(下游写入)在模型中的分离。
Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reasoning model, and find it is not a failure to detect harm. Harm is encoded as an internal direction that is nearly language-invariant in mid-network (English-vs-Indic cosine ${\approx}0.9$ at $L11$), and steering that direction upstream causally controls refusal. But the detection direction is orthogonal to the change that actually writes the refusal, which is late and assembled over the course of generation rather than read off in a single forward pass. We attribute the write to a specific, localizable circuit, a mixture-of-experts writer held in check by an attention opposer and price every way of intervening on it: damping the opposer is cheap and effective, amplifying the writer is a cost wall, and surgical edits to the responsible heads do nothing. The circuit’s organization, and the gradient method that exposes it, recur in a second, unrelated MoE model, while the lever’s strength is architecture-specific. The result is a cost-measured map of where a multilingual safety repair can land, and what it costs
[213] HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails cs.AI | cs.CL | cs.LGPDF
Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum
TL;DR: HoloAegis提出了一种极简参数化的拓扑推理框架,用于实现零样本LLM安全护栏。该方法将文本映射到单位球面后,通过纯几何推理(基于预计算的系统拓扑锚点库和吉布斯-玻尔兹曼自由能计算)进行安全评估,并引入双时间尺度指数移动平均来检测多轮语义漂移。
Details
Motivation: 解决当前LLM安全护栏面临的核心矛盾:微调会扭曲预训练表示,而生成式判别器则带来过高的推理成本。论文旨在探索是否可以通过对冻结的语义表示进行纯几何推理来实现安全防护。
Result: 在8个基准测试上评估,HoloAegis达到了最先进的准确率(在AuthenHallu上AUC为1.0000,在HarmBench上为0.9802),具有亚毫秒级延迟、零冷启动数据,并展示了跨语言迁移能力(在中文CHIFRAUD上AUC为0.9758)。
Insight: 核心创新在于将安全评估形式化为基于预计算拓扑锚点库的几何推理问题,实现了表示与推理的解耦。其宣称的’拓扑边界稳定性猜想’表明,稀疏的锚点质心比全向量空间方法能更稳定地抵御高频词汇扰动,这为构建高效、鲁棒的安全护栏提供了新思路。
Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge the prevailing paradigm by asking: can safety be achieved through pure geometric reasoning over frozen semantic representations? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning. We term our approach minimally parametric because the only free parameters are the anchor count K and the temperature tau, both fixed after construction and requiring no gradient-based training. An un-fine-tuned encoder maps text to a unit sphere, after which all decisions are purely geometric. We formalize safety evaluation as a Gibbs-Boltzmann Free Energy computation over a pre-computed System Topology Anchor Bank, and we introduce Dual Time-Scale Exponential Moving Averages to detect progressive multi-turn semantic drift. Our key theoretical insight is a Topological Boundary Stability Conjecture: we provide theoretical motivation and strong empirical evidence that sparse anchor centroids stabilize the decision boundary against high-frequency lexical perturbations far better than full vector space methods. Evaluated across 8 benchmarks, HoloAegis achieves state-of-the-art accuracy (1.0000 AUC on AuthenHallu, 0.9802 on HarmBench) with sub-millisecond latency, zero cold-start data, and cross-lingual transfer (0.9758 AUC on Chinese CHIFRAUD).
[214] MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models cs.AI | cs.CLPDF
Rahma Simin Ali, Jawad Hossain
TL;DR: 本文研究了在低资源语言孟加拉语中,对于数学推理任务,使用教师生成的思维链监督是否比仅使用答案监督的微调方法更有优势。作者构建了孟加拉语数学推理数据集MathShikkha,并在四个4B-7B参数规模的小语言模型上进行了对比实验。研究发现,思维链监督的价值取决于骨干模型的能力和数据分布变化,其主要优势在于提升孟加拉语忠实度、提供可审计的推理过程以及增强领域外鲁棒性,而非显著提升领域内的推理有效性。
Details
Motivation: 解决低资源语言(如孟加拉语)中数学推理的挑战,探究教师生成的思维链监督是否比传统的仅答案监督微调能带来额外收益。
Result: 在自建的MathShikkha数据集上,对于三个较强的骨干模型,思维链监督相比仅答案监督没有显著提升;但对于较弱的4B模型,则显著提升了18.56分。在更大的、经过污染审计的BanglaMATH基准测试中,思维链监督对所有四个模型都显著优于仅答案监督,提升幅度在20.1到28.1分之间。
Insight: 思维链监督的主要价值并非直接提升推理内容的正确性,而是增强了目标语言(孟加拉语)的忠实度、产生了可检查的推理过程,并提高了模型在领域外数据上的鲁棒性。其效果高度依赖于骨干模型的能力和数据分布,为低资源语言场景下的监督策略选择提供了重要洞见。
Abstract: Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B–7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15–52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p < 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1–28.1 points (all $p < 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen’s $κ= 0.76$–$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision’s value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.
[215] Can Open-Weight Models Compete on Financial Text Comprehension? cs.AI | cs.CL | cs.IR | q-fin.GNPDF
Jan Spörer
TL;DR: 本文评估了开源权重模型在金融文本理解任务上的表现,更新了包含2,967个问题的Financial Touchstone基准,并测试了20个模型。研究发现,开源模型如Kimi K2.6在准确性上排名第三,挑战了推理架构或专有模型权重是金融理解必要条件的假设。
Details
Motivation: 尽管开源权重模型在通用基准上已接近前沿专有模型,但其在真实世界金融任务中的可靠性尚未得到充分验证,本文旨在填补这一空白。
Result: 在Financial Touchstone基准上,Anthropic的Claude Opus 4.6准确率最高(88.4%),Google的Gemini 2.5 Pro幻觉率最低(0.08%)。开源模型Kimi K2.6准确率排名第三,GLM 5和Mistral 3分别第四和第五。信息检索是主要瓶颈,占所有失败的48.9%。
Insight: 开源权重模型在金融理解任务上可以表现出色,无需依赖推理架构或专有权重。研究还发现,中国模型的地缘政治内容过滤器会拒绝合法的金融问题(0.08%尝试),且拒绝行为与访问路径和模型本身相关,这揭示了模型部署中的潜在偏差问题。
Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba’s proprietary flagship Qwen3-Max. Anthropic’s Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google’s Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
[216] ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons cs.AI | cs.CLPDF
Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang
TL;DR: 本文提出了ComboShoppingBench,这是一个用于评估LLM智能体在模拟电商和外卖环境中进行预算约束下、使用优惠券的组合购物任务的基准。该基准通过任务合成生成包含优惠券、预算约束和用户查询的测试用例,并采用LLM评判和确定性验证相结合的方式进行评估。实验表明,即使是强大的LLM智能体在该基准上也表现不佳。
Details
Motivation: 解决现实世界中组合购物任务(如设备配置、备餐、活动策划)的评估挑战,这类任务需要联合推理商品兼容性、可用性、店铺要求、配送费、优惠券和预算,且存在多个可行方案,使得精确匹配指标不适用,而纯语义评估又无法检测不可行订单或错误支付。
Result: 在ComboShoppingBench上对多种LLM智能体进行的实验表明,即使是强大的智能体也表现挣扎,突显了在可靠、约束感知的组合购物方面仍有巨大的改进空间。
Insight: 创新点在于提出了一个开放但可验证的组合购物基准,其核心是通过探索智能体生成可行且语义连贯的购物篮作为“见证”,来指导生成具有对齐评估标准的测试用例,并采用LLM评判(语义)与确定性验证(逻辑约束)相结合的混合评估框架。
Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.
[217] Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics cs.AI | cs.CL | cs.CY | cs.GTPDF
Yen-Shan Chen, Yu Chian Duan, Chih-En Kuo, Jian-Bin Wu, Yun-Nung Chen
TL;DR: 该论文提出了Avalon-ToM-Bench,一个基于《抵抗组织:阿瓦隆》游戏非对称信息机制的细粒度心智理论评估基准。它将心智理论分解为认知与动机推理、推断与行动的2x2分类,并通过人工构建的视角受限查询进行评估。对28个大型语言模型的测试揭示了其在心智理论推理上的关键局限性,特别是在社会推理表达而非知识缺失方面。
Details
Motivation: 现有心智理论评估要么过于简化静态场景,要么交互式设置诊断能力有限,无法精细评估智能体的心智状态推理能力。
Result: 在Avalon-ToM-Bench上评估28个LLM,结果显示模型游戏规则理解强但心智理论能力弱;线性探测隐藏状态准确率达77-82%,远高于模型自身思维链的62-70%;专门推理训练带来+11.0分的显著提升,而测试时思维链仅带来+1.1分的边际收益。
Insight: 创新点在于通过非对称游戏机制构建细粒度、可诊断的心智理论分类评估框架。核心发现是心智理论失败主要源于社会推理表达障碍而非知识缺失,且稳健的心智理论依赖于习得的推理策略而非推理时的额外思考。
Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy – epistemic versus motivational reasoning crossed with inference versus action – using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequently represent correct mental-state inferences in their hidden states but fail to express them during generation – linear probes recover 77-82% accuracy versus 62-70% from the models’ own chain-of-thought. 3) Policy, not deliberation. Dedicated reasoning training yields substantial improvements whereas test-time chain-of-thought provides only marginal gains (+11.0 versus +1.1 points on average), suggesting that robust ToM depends on a learned reasoning policy rather than increased inference-time deliberation.
[218] Mismatch Matters: On-Policy Distillation Beyond Token Agreement cs.AI | cs.CLPDF
Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang
TL;DR: 本文揭示了在线策略蒸馏(OPD)中存在的一种失效模式——退化一致性,即学生模型通过重复循环实现与教师模型的近乎完美的令牌一致性,但全局响应却存在缺陷。为此,作者将关注点从一致性转向师生不匹配,并将不匹配令牌主要分为两类:学生超额令牌和学生赤字令牌。为应对这两类不匹配,论文提出了TIDE方法,通过有界Hellinger塑形抑制最严重的超额采样,并通过解析式教师Top-K注入恢复赤字概率质量。在多个数学推理基准测试中,TIDE均优于标准OPD及近期基线方法。
Details
Motivation: 动机在于发现现代大语言模型后训练流程中的核心组件——在线策略蒸馏存在退化一致性的失效模式,即学生模型可能通过局部重复来虚假地匹配教师令牌,导致全局响应质量低下。因此,需要超越简单的令牌一致性,深入分析并解决师生模型输出之间的不匹配问题。
Result: 在多个数学推理基准测试(使用多个Qwen3师生模型对)上,TIDE方法一致地超越了标准OPD以及近期的令牌选择和奖励塑形基线方法。特别是在师生不匹配严重的情况下,TIDE将Avg@8指标从6.9%提升至20.3%,将平均响应长度减少了3.6倍,并大幅减少了格式错误。
Insight: 创新点在于将师生不匹配系统地分类为学生超额令牌和学生赤字令牌,并针对性地提出了TIDE方法,结合有界损失塑形和解析式概率注入,分别处理这两类问题。这为改进在线策略蒸馏提供了一种更精细、更稳定的训练视角,超越了单纯追求令牌一致性的传统做法。
Abstract: On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher-student mismatch, and find that mismatch tokens can be mainly categorized into two types: student-excess tokens and student-deficit tokens. Student-excess tokens are generated by the student but assigned near-zero probability by the teacher; their log-ratio corrections grow unbounded and destabilize the update. Student-deficit tokens, in contrast, are preferred by the teacher but rarely sampled by the student; their absence blocks the transfer of the teacher’s reasoning patterns. To tackle these mismatch directions, we propose TIDE (Token-level Independent Deficit-Excess correction), which applies bounded Hellinger shaping to suppress the most severe sampled excesses and an analytic teacher top-$K$ injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher-student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. Moreover, the gains of TIDE are more pronounced under strong teacher-student mismatch, where it improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures. Code is available at https://github.com/yzc-666/TIDE
[219] Towards Expert-level Medical AI for Real-time Video Consultations cs.AI | cs.CL | cs.CVPDF
Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother
TL;DR: 本文介绍了AMIE (Video),一个基于Gemini的多智能体系统,首次在实时临床视频会诊中实现了专家级AI性能。该系统整合了低延迟对话、临床推理和实时视听感知,在随机对照的OSCE研究中,AMIE (Video)在病史采集、诊断、管理和体格观察等方面表现与初级保健医生相当或更优。
Details
Motivation: 解决现有基于文本的医疗AI丢弃了关键的感知维度(如非语言线索)且无法服务无法书面描述症状的患者的问题,旨在将医疗AI扩展到实时视听交互,以匹配临床实践的自然沟通复杂性。
Result: 在包含30名初级保健医生、15名患者演员和100个临床场景的随机OSCE研究中,AMIE (Video)在病史采集、诊断、管理和体格观察方面的表现被临床评估者评为与初级保健医生相当或更好;患者演员在评估和解释病情方面更偏好AMIE,但在建立融洽关系和伙伴关系方面更偏好医生。
Insight: 创新点在于构建了一个整合实时视听感知的低延迟多智能体系统,并为此建立了远程医疗临床视听线索的分类学和自动化评估方法;客观来看,这是首个在结构化临床考试中展示出与医生相当综合能力的实时视频会诊AI,标志着向处理临床实践感官复杂性的增强护理AI系统迈出了重要一步。
Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE’s approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)’s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
[220] CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models cs.AI | cs.CV | cs.ROPDF
Hsu-kuang Chiu, Stephen F. Smith
TL;DR: 本文提出了CMU-Drive基准测试和V2V-VLA模型,旨在解决多智能体协同自动驾驶问题。CMU-Drive是一个用于评估多辆联网自动驾驶车辆在安全关键场景下协同性能的闭环端到端基准。V2V-VLA是一个协同视觉-语言-动作模型,能在单次前向传播中联合生成驾驶动作、未来路径点、语言推理和通信策略。
Details
Motivation: 现有视觉-语言-动作模型主要针对单个自动驾驶智能体,缺乏对协同感知、推理和规划的支持,本文旨在填补多智能体协同端到端自动驾驶评估与建模的空白。
Result: 在提出的CMU-Drive基准上进行了实验,为协同VLA驾驶建立了首个基准和基线,为未来多智能体、闭环、端到端协同自动驾驶研究奠定了基础。
Insight: 创新点在于将协同驾驶整合到一个统一的VLA模型框架中,实现了动作、路径、推理和通信的联合生成,并构建了首个专门用于评估多车协同驾驶的闭环端到端基准。
Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.
[221] CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits cs.AI | cs.CVPDF
Xinqi Yang, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du
TL;DR: 该论文提出了CircuitReason-1k基准测试,包含1000个真实的教科书电路问题,用于全面评估从视觉电路图到符号推理的长程、多步骤推理能力。基准测试包含配对的问题、图表、答案和参考解决方案,并采用证据优先的构建流程和面向推理的分类法。
Details
Motivation: 解决现有评估方法在电路分析任务上的不足,该任务不仅需要识别图像中的组件,还需要完成符号接地、拓扑恢复、物理模型选择、方程构建、中间量传播以及保持单位、符号、方向和相位约定等一系列复杂的长程推理步骤。
Result: 在三个商业聊天机器人系统和六个开源多模态大语言模型上进行评估,最高准确率达到84.8%。然而,在长程推理问题上性能持续下降,定性分析揭示了在拓扑-目标绑定、物理约定和后期输出传播方面存在持续失败。
Insight: 创新点在于构建了一个专注于长程视觉到符号推理的、证据优先的基准测试,并引入了结合保守类型评分和身份盲多模型语义共识的评估方法,为衡量多模态模型将技术视觉证据转化为持续、物理有效的符号推理能力提供了聚焦的测试平台。
Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning process. Each problem pairs one or more circuit diagrams with a self-contained question, a typed or semantically specified answer, and a reference worked solution. An evidence-first construction pipeline aligns questions, figures, and solutions, while a reasoning-oriented taxonomy organizes problems by circuit type and dependency depth. Evaluation combines conservative typed scoring with identity-blinded multi-model semantic consensus, retaining every problem in the denominator. Across three commercial chatbot systems and six open-source multimodal large language models, the highest-scoring system reaches 84.8% accuracy. However, performance consistently deteriorates on long-horizon problems, and qualitative analysis exposes persistent failures in topology-to-target binding, physical conventions, and late-stage output propagation. \benchmark{} provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning. Code are available at GitHub - CircuitReason/CircuitReason1K.
[222] GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models cs.AI | cs.CVPDF
Zhihang Liu, Mei-Po Kwan, Jinlin Wu, Hao Li
TL;DR: 本文提出GeoPhysAdapter方法,用于提升滑坡测绘的跨域迁移能力。该方法基于冻结的视觉基础模型,通过地形、物质和降雨触发等多尺度地球物理信息进行约束,并在像素和候选滑坡体两个决策单元上进行有界自适应,有效减少了跨域误报。
Details
Motivation: 滑坡测绘在应急响应和区域风险评估中至关重要,但新触发的滑坡缺乏即时标注,跨域可迁移性成为关键。现有视觉基础模型在未见区域、事件和数据源上仍会产生高置信度误报,而地形、物质和降雨触发等地球物理信息因尺度不匹配(如10米网格重采样)与分割决策单元错位,加剧了不确定地理上下文问题(UGCoP)。
Result: 在包含四个公共数据源、55个全球滑坡事件和7,890个测试样本的事件隔离PILD数据集上,像素级自适应移除了507,817个错误像素,错误减少7.76%;而将决策单元提升至候选滑坡体,在相同样本、锚点和基线条件下,错误减少提升至23.99%(约为像素级效果的3.1倍),IoU提高0.031(相对提升14.2%),并实现每损害一个像素纠正9.92个像素的效果。
Insight: 创新点在于提出多尺度地球物理适应框架,将地形、物质和降雨触发分别约束为密集空间引导、区域调制和事件时序强制,并在像素和候选滑坡体两个决策单元进行有界自适应,当支持不足时精确回退到视觉预测。这解决了地球物理信息与分割决策单元尺度不匹配的问题,显著提升了跨域滑坡测绘的准确性。
Abstract: Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still generate high-confidence false alarms. Terrain, material, and rainfall triggering can constrain such errors, but their supports are local, regional, and event-scale, so that resampling onto a 10~m grid misaligns them with the segmentation decision unit and compounds the uncertain geographic context problem (UGCoP). We propose GeoPhysAdapter, which anchors on a frozen vision foundation model, restricts terrain, material, and triggering to dense spatial guidance, regional modulation, and event-timing forcing, and applies bounded adaptation at two decision units, the pixel and the candidate landslide body, reverting exactly to the visual prediction where support is insufficient. On an event-isolated PILD dataset of four public sources, 55 global landslide events, and 7,890 test samples, 70.3% of cross-domain false-positive mass lies in near-pure spurious bodies of median equivalent diameter 207m, matching coarse-prior support rather than the pixel. Pixel-level adaptation removes a net 507,817 erroneous pixels and reduces error by 7.76%, whereas raising the decision unit to the candidate body, under identical samples, anchor, and baseline, increases error reduction to 23.99%, approximately 3.1 times the pixel-level effect, improves IoU by 0.031 (14.2% relative), and corrects 9.92 pixels per pixel harmed. The data and code are publicly available at: https://github.com/Liu-Zhihang/geophysadapter.
[223] CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation cs.AI | cs.CV | cs.LG | cs.ROPDF
Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera
TL;DR: 本文提出了CADEngBench,一个用于评估计算机辅助设计(CAD)模型工程能力的双轨基准测试。该基准包含CADEngBench-P(针对300个参数化零件,测试B-Rep有效性、工程与可制造性检查、参数扰动、功能编辑及线性静态有限元分析匹配)和CADEngBench-A(针对150个零件对,测试关节检索、面边定位、关节框架预测及运动学验证)。通过对八个多模态、代码生成模型进行测试,发现编辑现有CAD模型比生成新模型更容易,而复杂编辑和有限元分析匹配仍具挑战性,且装配预测常无法准确恢复关节或配合实体。
Details
Motivation: 当前CAD模型评估往往仅关注外观正确性,但工程级CAD模型还需满足设计要求、参数可预测性、可控编辑、结构响应匹配及有效关节连接等工程行为。本文旨在填补这一空白,建立一个全面评估CAD模型工程能力的基准。
Result: 在八个多模态、代码生成模型上的测试结果表明,编辑提供的CAD模型比从零生成更容易;复杂编辑和匹配的线性静态有限元分析(FEA)仍然困难;装配预测通常能定位相关区域,但难以准确恢复记录的关节或配合实体。该基准揭示了现有模型在工程行为测试上的不足。
Insight: 论文的创新点在于首次提出了一个系统评估CAD模型工程行为(而非仅外观)的双轨基准,强调了参数化设计、装配推理和物理仿真的综合测试。从客观角度看,该基准将CAD评估从几何层面提升到功能与工程层面,为未来CAD生成与理解模型的发展提供了关键的评估框架和方向指引。
Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification. Across eight multimodal, code-capable models, editing supplied CAD is substantially easier than generating it, while complex edits and matched FEA remain difficult. Assembly predictions often locate the relevant region but fail to recover the recorded joint or mating entities. These results show that CAD evaluation must test engineering behavior rather than appearance alone.
cs.NE [Back]
[224] An evolutionary model of animats with VLM-based subjective evaluation cs.NE | cs.AI | cs.CL | cs.HC | cs.MAPDF
Shota Miyazaki, Takaya Arita, Reiji Suzuki
TL;DR: 本研究提出了一种将视觉语言模型(VLM)的主观评价融入遗传算法适应度评估与选择过程的框架。该框架利用VLM对虚拟软体机器人的运动序列图像进行基于主观评价词(如“可爱地”、“奇怪地”)的成对比较,并将比较结果作为选择压力,从而同时演化形态与运动。实验表明,VLM的主观选择比随机选择加速了种群收敛,并产生了与各评价词相对应的独特形态与运动。
Details
Motivation: 解决如何将人类主观、定性的语言评价(如“可爱”、“奇怪”)整合到进化计算的选择过程中,以引导虚拟生命体(animats)演化出符合特定主观描述的形态与行为,并探索VLM在实现这种主观评价中的作用。
Result: 在虚拟软体机器人演化实验中,VLM的主观选择相比随机选择加速了种群收敛,并演化出与各主观评价词(如“可爱地”、“奇怪地”)相对应的独特形态与运动模式。辅助人类实验表明,尽管单次选择与VLM不完全一致,但演化出的形态与运动趋势在定性上相似,且人类评估易产生疲劳。
Insight: 创新点在于将VLM作为主观评价器引入进化计算,实现了基于自然语言指令的演化引导。一个关键发现是VLM可能并非字面应用评价词,而是将其分解为多个内部评价标准进行判断,这为分析VLM主观判断结构提供了基础框架,并有望推动基于主观评价的进化计算与人工生命研究。
Abstract: In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness evaluation and selection processes of a genetic algorithm. As the target of evolution, we employ virtual soft robots with flexible morphologies and locomotion and present the VLM with sequence images representing the locomotion of two individuals. Selection is performed via pairwise comparisons based on subjective evaluation terms such as adorably and weirdly. The outcomes of these comparisons are used as selection pressure within the genetic algorithm, enabling the simultaneous evolution of morphology and locomotion. Experimental results demonstrate that subjective selection by the VLM accelerates population convergence compared to random selection, while also giving rise to distinctive morphologies and motions corresponding to each evaluation term. An auxiliary experiment with human participants further showed that, although individual pairwise choices only partly agreed with the VLM selections, the resulting morphological and locomotion tendencies were qualitatively similar and repeated human evaluations imposed noticeable fatigue. Moreover, the observation that similar evolutionary outcomes emerged across different evaluation terms suggests that the VLM does not apply these terms in a purely literal manner but instead decomposes them into multiple internal evaluation criteria when making judgments. This work visualizes the evolutionary process through which subjective linguistic expressions are mapped onto embodied phenotypes and provides a foundational framework for analyzing the structure of subjective judgment in VLMs. The proposed approach is expected to contribute to new developments in evolutionary computation and artificial life research based on subjective evaluation.
cs.RO [Back]
[225] SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments cs.RO | cs.CVPDF
Xuan Yao, Yuze Zhu, Junyu Gao, Zongmeng Wang, Changsheng Xu
TL;DR: 本文提出了一种名为SC²-WM的自校正世界模型框架,用于解决连续环境下的视觉语言导航问题。该框架通过引入内部反馈机制实现闭环决策,利用世界模型的前瞻信息进行状态级规划修正,并在测试时通过条件性世界感知适应进行模型级校正,以提升导航的鲁棒性和泛化能力。
Details
Motivation: 现有VLN-CE方法大多依赖开环执行,缺乏在推理过程中检测和校正内部状态漂移的机制,导致在部分可观测环境下导航决策不够鲁棒。
Result: 在标准VLN-CE基准测试上的实验表明,该方法提高了导航的鲁棒性和泛化性能。
Insight: 创新点在于将闭环反馈引入世界模型,通过状态级规划修正和条件性模型更新相结合,实现了从内部状态到模型参数的多层次自校正机制,为部分可观测环境下的序列决策提供了新思路。
Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise-ikun/SC2_WM.
[226] Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning cs.RO | cs.CVPDF
Shilin Shan, Chuhao Zhou, Ruize Wang, Xinyan Chen, Xiangyu Chen
TL;DR: 这篇综述论文提出了一个名为TF-ART的统一分类法,用于系统性地回顾和分类结合触觉与力感知的机器人学习方法。它从多模态感知和多阶段系统设计的联合视角出发,对现有方法进行梳理,并探讨了物理交互的任务设置与基础设施需求。
Details
Motivation: 现有综述尚未从统一视角,即同时涵盖多模态感知(如力、触觉、视觉、语言)和多阶段系统设计(如高层策略、动作细化模块、底层控制器),来系统回顾力与触觉感知的机器人学习。本文旨在填补这一空白。
Result: 论文提出了TF-ART分类法,将现有方法映射到一个统一的分层结构中,用于表征其如何组织观测模态、编码融合异构感官输入、跨多阶段生成并细化动作,以及将学习策略连接到反应式机器人末端控制。
Insight: 主要创新点是提出了一个统一的多模态与多阶段框架分类法(TF-ART),为力与触觉感知的机器人学习领域提供了系统化的方法论视图,并整合了算法与实践的双重视角。
Abstract: Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on force regulation and adaptive control. In this context, recent robot learning methods have made substantial progress by integrating force, tactile, vision, language, and proprioceptive sensing into learned manipulation policies. In parallel, many systems adopt multi-phase architectures that combine high-level policies, action-refinement modules, and low-level controllers to bridge semantic task understanding with reactive physical execution. Despite these advances, existing surveys have not explicitly reviewed force- and tactile-aware robot learning from a unified perspective that jointly captures multimodal sensing and multi-phase system design. This survey addresses this gap by proposing TF-ART, a Tactile/Force-Aware Robot learning Taxonomy for multimodal and multi-phase frameworks, which maps individual methods into a unified hierarchical structure. The framework characterizes how recent works organize observation modalities, encode and fuse heterogeneous sensory inputs, generate and refine actions across multiple phases, and connect learned policies to reactive robot-end control. Building on this methodological view, we further examine the task settings and infrastructure requirements of physical interaction, thereby integrating both algorithmic and practical perspectives on force- and tactile-aware robot learning.
[227] AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization cs.RO | cs.AI | cs.CVPDF
Peng Xu, Chengcheng Wang, Shaohua Wan
TL;DR: 本文提出AeroDPO框架,旨在解决轻量级无人机视觉语言导航(UAV-VLN)中感知质量与模型规模权衡以及行为克隆鲁棒性不足的问题。通过结合高保真视觉输入和基于确定性物理模拟回滚的自动化直接偏好优化(DPO)流程,该方法在仅使用20亿参数模型的情况下,在未映射场景中实现了49.16%的成功率并显著降低碰撞率,达到了新的最先进水平。
Details
Motivation: 当前无人机视觉语言导航的端到端范式通常依赖数十亿参数的大语言模型,导致边缘部署延迟过高;同时,轻量级模型虽能匹配性能,但基于纯行为克隆的策略缺乏显式负反馈,在分布外场景中表现出严重的碰撞风险。
Result: 在未映射场景的基准测试中,AeroDPO将成功率提升至49.16%,同时大幅抑制碰撞率,超越了70亿参数基线的整体成功率,为自主空中智能体建立了新的SOTA。
Insight: 创新点在于揭示了感知质量比语言推理能力更关键,并设计了零成本的自动化DPO流程,通过物理模拟回滚自主生成偏好数据以替代人工标注,有效增强了轻量级模型的鲁棒性和安全性。
Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
[228] LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding cs.RO | cs.CVPDF
Zhewei Zhang, Puyue Wang, Guanren Qiao, Yijie Weng, Jiawei Hu
TL;DR: 本文提出了LIRA,一种用于视觉-语言-动作(VLA)模型的局部跨层信息路由机制。它将预训练视觉语言模型(VLM)的中间特征路由到动作解码器的过程建模为深度感知的信息路由,通过并行融合块聚合相邻层的查询特征,并与任务令牌及本体感觉输入整合,从而预测机器人动作。
Details
Motivation: 现有VLA模型在将VLM中间特征路由到动作解码器的接口设计上探索不足,要么仅暴露表征层次的一小部分,要么将解码器块与VLM层僵硬地一一对应,限制了跨深度获取互补任务证据的能力。
Result: 在LIBERO、LIBERO-Plus、CALVIN ABC→D和真实世界操作任务上,LIRA在相同0.5B参数配置下,相比VLA-Adapter基线提升了主要聚合指标。在LIBERO-Plus的零样本迁移中,平均成功率从59.1%提升至78.0%,提高了18.9个百分点,表明其在受控分布偏移下具有更好的鲁棒性。
Insight: 创新点在于将VLM到动作的条件化过程形式化为局部、深度感知的信息路由,通过为每个并行融合块分配一个以对应VLM层为中心的局部窗口来聚合相邻层信息,同时保持了主干架构、动作解码器和监督训练方案不变,实现了灵活且有效的特征利用。
Abstract: Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Existing designs either expose only a narrow part of the representation hierarchy or rigidly match each decoder block to one VLM layer, restricting access to complementary task evidence across depths. We introduce LIRA, a local cross-layer action-conditioning mechanism that formulates VLM-to-action conditioning as depth-aware information routing. LIRA operates on task-token features and LIRA Query features derived from intermediate VLM states, then assigns each Parallel Fusion Block a depth-aligned local window centered on its corresponding VLM layer. Parallel Fusion Blocks aggregate neighboring LIRA Query features and integrate them with task-token features and proprioceptive inputs before action prediction. This routing interface leaves the backbone architecture, action decoder, and supervised training recipe unchanged. Across LIBERO, LIBERO-Plus, CALVIN ABC$\rightarrow$D, and real-world manipulation, LIRA improves the principal aggregate metrics over the VLA-Adapter baseline under the same 0.5B-parameter configuration. In zero-shot transfer to LIBERO-Plus, LIRA increases average success from 59.1% to 78.0%, an 18.9-point gain indicating improved robustness under controlled distribution shifts.
[229] PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets cs.RO | cs.CVPDF
Jie Huang, Xiaohe Li, Jiahao Li, Fangli Mou, Chen Qian
TL;DR: PhysX-CoT提出了一种从单张图像生成仿真就绪3D资产的新方法,将生成过程重新定义为显式的结构化物理推理链(Chain of Thought),通过分解部件级状态(如分解、2D/3D定位、关系、几何和表面线索)来监督和优化,从而提升几何、尺度和物理属性的生成质量。
Details
Motivation: 现有方法通常将单图像到3D资产生成视为输出中心的视觉语言模型任务,导致部件放置与局部形状在全局坐标流中纠缠,且中间物理状态缺乏监督或验证,限制了生成资产的仿真可用性。
Result: 在统一协议下(使用相同骨干网络、数据和冻结解码器重新训练所有基线),PhysX-CoT在几何、尺度和物理属性指标上均优于最接近的全任务基线;在Unreal Engine 5中生成的资产解析、碰撞和关节运动具有高有效性。
Insight: 创新点包括将生成过程重构为显式结构化物理推理链,实现部件状态分解与独立监督;通过几何因子化(3D框处理放置、局部编码处理形状)和CoT对齐的GRPO优化,提升物理一致性;中间状态可作为奖励目标,增强模型可解释性与可控性。
Abstract: Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset for a decoder to turn into geometry and physical fields, leaving the image-to-3D reasoning implicit. We argue the limiting factor is this output-centric view: part placement and local shape are entangled in one global-coordinate token stream, and the intermediate physical states are never exposed for supervision, conditioning, or verification. PhysX-CoT instead casts single-image asset generation as an explicit structured physical reasoning process, an ordered and machine-parseable trajectory of part-level states covering decomposition, 2D and 3D grounding, relations, coarse geometry, and surface cues that we separately supervise, use to condition geometry, and treat as reward targets. Geometry is factorized so that 3D boxes carry placement and local codes carry shape, and CoT-aligned GRPO optimizes parse validity, grounding, geometry, placement, and physical consistency. Under a unified protocol that retrains all learned baselines on the same backbone, data, and frozen decoder, PhysX-CoT outperforms the closest full-task baseline across geometry, scale, and physical-attribute metrics. Oracle, token-matched, and state-order controls show the explicit states are functional rather than cosmetic, and in Unreal Engine~5 the generated assets parse, collide, and articulate at high validity.
[230] Action- and Language-Conditioned Video Assessment for Embodied Control cs.RO | cs.CVPDF
Hwanhee Kim, Jaehyun Jang, Seungmin Cha, Hyeonseo Yun, Donghoon Lee
TL;DR: 本文提出了ALVA(动作与语言条件视频评估),一种用于评估具身智能体执行多步自然语言指令任务进度的轨迹评估器。该方法利用预训练视觉语言模型(VLM),通过两个阶段(基于动作总结视觉转换,再根据指令评估总结)来生成离散的轨迹级进度分数。在模拟3D家庭环境中,ALVA展现出保守的评估模式,误报率接近零,并能作为终端反馈有效提升闭环策略优化的性能。
Details
Motivation: 解决基于视觉的具身智能体在执行多步指令时,现有基于最终帧匹配或连续嵌入相似性的反馈方法可能忽略对判断任务完成至关重要的中间状态转换的问题。
Result: 在模拟3D家庭环境任务中,ALVA的误报率接近零。当用作闭环策略优化的终端反馈时,其性能优于静态图像和基于嵌入的视觉基线方法,并缩小了与真实情况(ground-truth oracle)的性能差距。
Insight: 创新点在于提出了一个结合视觉观察、执行动作序列和自然语言指令的轨迹级评估框架,其两阶段(总结与评估)设计利用了预训练VLM,为具身控制任务提供了一种可解释的反馈机制。
Abstract: Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete trajectories. Conventional approaches based on final-frame matching or continuous embedding similarity may overlook intermediate transitions that are necessary for determining whether an instruction has been completed. We propose ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction. The method uses a pre-trained vision-language model (VLM) in two stages: it first summarizes frame-to-frame visual transitions conditioned on the executed actions and then assesses the generated summary with respect to the instruction to produce a discrete trajectory-level progress score. In simulated 3D household environments, ALVA exhibits a conservative assessment pattern with near-zero false-positive rates. When used as terminal feedback for closed-loop policy optimization, it provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle. These results support action- and language-conditioned video assessment as an interpretable feedback mechanism for the evaluated simulated embodied-control tasks.
[231] SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models cs.RO | cs.CVPDF
Junjie He, Junfeng Li, Zhide Zhong, Haodong Yan, Ruixin Li
TL;DR: 本文提出SG-WAM,一种用于世界-动作模型(WAMs)的语义引导方法,旨在解决现有WAMs因文本编码器独立于视觉观察而导致的预测视频与语言指令语义错位问题。该方法利用视觉语言模型(VLM)作为语义规划器,预测文本接地和空间感知的语义前瞻,并将其作为高层语义引导注入WAM,以确保未来视频生成和动作预测忠实遵循语言指令。
Details
Motivation: 现有世界-动作模型主要依赖视觉线索生成未来视频和动作,由于现成的文本编码器独立于视觉观察嵌入指令,导致预测的视频常与语言指令语义错位,从而降低了预测动作的准确性。
Result: 在仿真和真实世界的大量实验表明,该方法具有优越性,展现了精确的操控能力和强大的指令跟随能力。
Insight: 创新点在于利用VLM作为语义规划器,生成文本接地(识别正确目标物体)和空间感知(提供场景几何信息)的语义前瞻,并将其作为高层引导注入WAM,从而增强了模型的指令接地能力,确保了语义对齐。从客观角度看,这是一种将高级语义规划与低级动作生成有效结合的框架创新。
Abstract: World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that leverages a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of world-action models. Specifically, we train a VLM-based planner to predict text-grounded and spatial-aware semantic foresight. The text-grounded semantic foresight grounds the instruction by identifying the correct target objects, and the spatial-aware semantic foresight provides the scene geometry for precise manipulation. We then inject this foresight into the world-action model as high-level semantic guidance, ensuring that both future-video generation and action prediction faithfully follow the language instruction. Extensive experiments in simulation and the real world demonstrate the superiority of our semantic guidance method, showcasing precise manipulation and strong instruction-following capabilities.
[232] VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction cs.RO | cs.CVPDF
Hongjin Ji, Guoyang Xia, Luoyang Sun, Fangxiang Feng, Lei Ren
TL;DR: 本文提出了一种可靠的测试时训练框架VANE,用于视觉-语言-动作模型在闭环操作中的自适应。该方法通过预测未来视觉表征来隔离和评估候选更新,仅在获得未来证据支持时才提交更新,从而实现选择性和可逆的适应。
Details
Motivation: 现有测试时训练方法在闭环操作中难以可靠使用,因为共享适应空间可能混合不兼容的任务修正,且在线更新可能在未知后果的情况下改变后续动作。
Result: 在SimplerEnv WidowX上,VANE比相应的TTT基线平均成功率提高了3.2个百分点;在Google Robot上的结果进一步表明部署时的收益仍依赖于具体任务和机器人本体。
Insight: 创新点在于将适应过程条件化于当前视觉-语言上下文,并通过学习执行动作的未来视觉后果来隔离和评估更新,提供了一种基于证据的、受约束的VLA策略适应方法。
Abstract: Test-time training (TTT) offers a lightweight way to adapt vision–language–action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision–language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.
[233] FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving cs.RO | cs.CVPDF
Guolei Huang, Tengfei She, Yuxuan Lu, Yao Huang, Yuqi Ye
TL;DR: FactorDrive是一个端到端自动驾驶框架,通过规划关键因素驱动的自适应多步推理来提升规划性能。它首先进行大规模驾驶领域指令微调以建立基础驾驶知识,然后构建PCF-CoT数据集,将规划推理基于轨迹相关的空间物理证据并围绕场景特定的规划关键因素组织推理路径。此外,该方法引入了质量搜索引导的组相对策略优化,通过蒙特卡洛树搜索发现高质量推理路径并优化策略。
Details
Motivation: 现有基于视觉语言模型的端到端自动驾驶方法未能充分将空间物理证据整合到规划推理中,且推理适应能力粗糙,无法满足场景特定的规划需求;同时,在自动驾驶后训练中,为提升规划质量而进行的推理路径优化尚未得到充分探索。
Result: 在开环(nuScenes)和闭环导向(NAVSIM)基准测试上的大量实验表明,FactorDrive实现了最先进的规划性能。
Insight: 创新点在于提出了规划关键因素驱动的自适应多步推理机制,以及结合蒙特卡洛树搜索与组相对策略优化的后训练方法,以发现和优化高质量推理路径,从而提升轨迹规划性能。
Abstract: Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.
[234] Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells cs.RO | cs.CV | cs.HC | cs.NIPDF
Emma Takács, Mátyás Hajós, Ádám Juniki, Ádám Fischer, Zoltán Komáromi
TL;DR: 本文提出了一种基于5G的无线可重构工作单元系统,旨在消除人机协作中的基础设施障碍。该系统集成了电池供电的多传感器平台和计算机视觉模块,用于物体检测、姿态估计和手部识别,并通过5G将计算密集型任务卸载到边缘,以解决带宽-延迟权衡问题。系统在匈牙利和挪威的多种5G基础设施中进行了部署和评估,展示了其可移植性和适应性。
Details
Motivation: 传统人机协作工作单元受限于电源和数据线缆,限制了模块化和重新配置的灵活性,且适用于实时感知和安全协作的商用无线设备选择有限。本文旨在通过无线和5G技术消除这些基础设施障碍,以支持动态、高混合、低产量的工业场景(如再制造)。
Result: 系统集成的计算机视觉模型在合成和真实数据上训练,能在不同光照和背景下可靠检测定向抓取姿态和人类手部(mAP@50-95为97.74 ± 0.10%,平均推理时间为12.5毫秒)。网络实验在兼容的网络-设备配对中实现了低至12毫秒的往返响应时间,适合安全、自适应的人机协作,但也揭示了当前5G部署中互操作性的实际限制。
Insight: 创新点包括提出了一种高度灵活、基于5G的无线系统原型,集成了电池供电的多传感器平台和增强的视觉模块(用于手部识别),并通过边缘计算卸载解决了带宽-延迟权衡。从客观角度看,该系统在真实5G网络环境中的可移植性验证和互操作性挑战的识别,为未来无线人机协作系统的实际部署提供了重要见解。
Abstract: Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements. Traditional setups are constrained by power and data cabling, restricting modularity and reconfigurations, while the selection of commercial wireless devices suitable for real-time perception and safe collaboration are limited in availability. This paper presents a highly flexible, wireless, 5G-based system that serves as a versatile experimental testbed for applications including remanufacturing, operator training, and user studies. To eliminate infrastructure barriers, the workcell integrates a novel battery-powered, multi-sensor platform prototype. Additionally, to support operator safety and system adaptability across environmental shifts, the system integrates a computer vision module for object detection and pose estimation, further augmented for robust hand recognition. Trained on synthetic and real data, the model reliably detects oriented grasping poses and human hands across varying lighting and background conditions (with an mAP@50-95 of 97.74 +- 0.10% and a mean inference time of 12.5 ms). Offloading these computationally intensive tasks to the edge via 5G, the proposed architecture contributes to resolving the bandwidth-latency trade-off. To demonstrate portability, the system was implemented in both Hungary and Norway, and was evaluated across a combination of public and private, Standalone and Non-Standalone 5G infrastructures. The performed network experiments produced results in round-trip response times down to 12 ms in case of compatible network-device pairings, suitable for safe, adaptive HRC. However, these measurements also revealed practical limitations related to interoperability in current 5G deployments that should be addressed in future works.
[235] RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance cs.RO | cs.CV | cs.LGPDF
Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su
TL;DR: 本文提出了RynnValue,一种基于时间距离的机器人操作价值基础模型,通过利用时间戳直接生成监督信号,无需偏好或进度标注,在超过7000小时、约300万指令条件视频片段上进行了训练。该模型在RBM-EVAL-OOD基准上超越了完全基于偏好监督的SOTA方法,并能零样本泛化到未见过的任务、机器人和视角,通过基于势能的塑形转换为密集奖励后,显著提升了真实世界策略的成功率。
Details
Motivation: 通用奖励模型已成为扩展机器人学习的瓶颈,现有方法依赖于任务内部锚点(如偏好或归一化进度),这些锚点难以在不同机器人和数据源之间迁移,因此需要一种可扩展的监督目标来学习价值相关能力。
Result: 在RBM-EVAL-OOD基准上,RynnValue的平均Kendall’s tau_a达到0.675,超越了完全偏好监督的SOTA(0.655),是仅基于进度方法(0.292)的两倍多;转换为密集奖励后,在线策略成功率从52.5%提升至72.5%,离线策略从63.8%提升至82.5%。
Insight: 创新点在于使用时间距离作为可扩展的监督目标,结合随机时间采样、时间顺序洗牌和价值隔离注意力机制来抑制捷径学习,从而构建了一个无需偏好标注、能跨任务和机器人泛化的价值基础模型,为通用机器人策略提供了实用的奖励接口。
Abstract: General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall’s tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
eess.SP [Back]
[236] Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations eess.SP | cs.CV | cs.LGPDF
Hongxiang Gao, He-yang Xu, Yuwen Li, Minghui Zhao, Zhipeng Cai
TL;DR: LuminaECG是一个临床结构化的心电图推理框架,它将心电图解释重新定义为基于测量的视觉阅读。该框架通过将心电图信号渲染在标准心电图纸上,明确描绘波形边界,并使用颜色编码分割将波形分解为离散的视觉测量基元,然后训练一个通用的2B视觉语言骨干网络,在不修改架构的情况下将这些基元与诊断推理关联起来。
Details
Motivation: 研究动机是探索心脏病专家解读心电图的过程(包括定位波形成分、测量节律和间期模式,并将这些结构化观察转化为诊断证据)是否能作为心电图智能体的有效先验知识。
Result: 在开放、专有和心电图专家零样本基线测试中,LuminaECG改善了波形测量和诊断恢复。它在CODE-test基准测试中达到了具有临床意义的读者层级,能够在不重新训练的情况下跨不同地理区域的心电图数据集进行迁移,并且其生成的报告结构包含了一个新兴的预后信号。
Insight: 论文宣称的创新点在于提出了一个将临床医生解读过程结构化为视觉测量基元的框架,并证明了这种基于可测量波形证据与临床知识对齐的监督对于构建有效的心电图智能体至关重要,而不仅仅是依赖更大的模型。从客观角度看,其将专业领域知识(医生先验)系统地嵌入到视觉语言模型训练中的方法具有借鉴意义。
Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence. Whether this expert reading process can serve as an effective prior for ECG agents remains unclear. To address this question, we introduce LuminaECG, a clinically structured ECG reasoning framework that reformulates ECG interpretation as measurement-grounded visual reading. ECG signals are rendered on standard electrocardiographic grid paper to preserve the spatial and scale cues used in clinical reading. P-wave, QRS-complex, and T-wave boundaries are explicitly delineated, and color-coded segmentation decomposes the waveform into discrete visual measurement primitives. A general 2B vision-language backbone is then trained with low-rank supervised fine-tuning to associate these primitives with diagnostic reasoning, without architectural modification. Across open, proprietary, and ECG-specialist zero-shot baselines, LuminaECG improves both waveform measurement and diagnostic recovery. It reaches a clinically meaningful reader tier on the CODE-test benchmark, transfers across geographically diverse ECG datasets without retraining, and generates reports whose structure contains an emergent prognostic signal. These findings suggest that effective ECG agents require not only larger models, but supervision that preserves the alignment between measurable waveform evidence and clinical knowledge.
cs.IR [Back]
[237] Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence cs.IR | cs.CV | cs.MMPDF
Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi
TL;DR: 本文提出了一种面向持续视觉世界的搜索基础设施模型,区别于传统视频检索系统。该模型基于分析器定义的场景、持久化理解产物、共享时间轴上的视觉记忆以及能力声明索引,并通过VideoDB数据格式实现。在语义检索实验中,使用通用组件构建的流水线在多个召回率指标上优于商业视频原生引擎。
Details
Motivation: 解决智能体在摄像头、屏幕、流媒体和档案等持续产生观察数据的视觉世界中,进行搜索时面临的系统性问题,包括连续数据流处理、多时间粒度理解、无需回放完整记录即可选择上下文,以及结果必须与可检查的源证据保持关联。
Result: 在涵盖四个公共数据集、超过9800个查询的语义检索对比中,基于通用组件构建的流水线在Recall@1/@3/@10的宏平均指标上(73.09/83.39/91.20)优于商业视频原生引擎(65.75/77.13/89.10),后者仅在Recall@50上略高(96.42 vs 96.07)。
Insight: 核心创新在于将视觉世界搜索视为基础设施问题,提出了一个模型无关的架构,明确区分记忆(保留的一切)、上下文(为任务选择的内容)和证据(支撑的源片段)。该系统设计将分割、采样、模型选择、嵌入和排序作为系统决策,并将实时流视为一等数据源,其性能表明系统设计对检索质量的影响目前超过了视频特定预训练的作用。
Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model-agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first-class sources, with video-native foundation models offered as fixed APIs. In a semantic-retrieval comparison against a commercial video-native engine spanning 9,800+ queries over four public datasets, a pipeline of general-purpose components achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video-specific pretraining, and visual-memory infrastructure can deliver it while keeping playable, source-grounded evidence first-class.
cs.LG [Back]
[238] TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing cs.LG | cs.CLPDF
Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic
TL;DR: 本文提出了TEMPER方法,一种基于张量网络的高效流形约束参数化方案,用于增强残差路由的表达能力。该方法通过将残差路由中的生成器表示为多模态张量并利用张量网络进行参数化,在保持动态路由能力的同时显著减少了参数数量。
Details
Motivation: 现有超连接方法通过多残差流和动态信息流提升了残差路由的表达能力,但其密集、非结构化的生成器导致参数数量随流数快速增长,形成了生成器层面的瓶颈。
Result: 在语言建模和常识推理任务上的综合实验表明,TEMPER在匹配或超越现有方法性能的同时,所需额外参数显著减少。在8个残差流时,TEMPER取得了最佳CORE分数,且比mHC方法减少了约84%的额外参数,展现了更优的性能-参数效率权衡。
Insight: 核心创新在于将路由生成器结构化为低秩张量网络,这既控制了学习路由子空间的维度(张量秩),又将生成器近似误差与路由逻辑及块输出差异关联起来,从而在保证表达力的同时提升了参数效率和模型可解释性。
Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing. However, a generator-level bottleneck remains in existing methods: they use dense, unstructured generators for pre-branch aggregation, residual mixing, and post-branch redistribution, which results in parameter count growing rapidly with the number of streams. To address this issue, we propose \underline{\textbf{T}}ensorized \underline{\textbf{E}}fficient \underline{\textbf{M}}anifold-constrained \underline{\textbf{P}}arameterization for \underline{\textbf{E}}xpressive Residual \underline{\textbf{R}}outing (\textbf{TEMPER}), which represents these generators as multi-way tensors over the input-stream, feature, and output-stream modes, and parameterizes them using tensor networks. Such a structured low-rank formulation is shown to preserve token-dependent manifold-constrained routing interface while substantially reducing parameter growth. It also promotes interpretability and intuition, as: i) tensor ranks control the dimensionality of the learned routing subspace, with full ranks recovering dense routing; while ii) the generator approximation errors bound differences in routing logits and, consequently, in the routed-block outputs. Comprehensive experiments show that TEMPER matches or outperforms existing methods across language modeling and commonsense reasoning tasks, while requiring substantially fewer additional parameters. At eight residual streams, TEMPER achieves the best CORE score while using about $84%$ fewer additional parameters than mHC, thus showing a stronger performance-parameter efficiency trade-off.
[239] Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning cs.LG | cs.CLPDF
Yifu Huo, Shunjie Xing, Chenglong Wang, Peinan Feng, Qiaozhi He
TL;DR: 本文提出了一种基于环境反馈的多时间尺度信用分配方法(EFCA),用于解决智能体强化学习中因奖励延迟和稀疏性导致的信用分配难题。该方法通过整合短期反馈信号和中长期状态历史信号,为中间决策提供更细粒度的监督,从而提升长期任务的成功率和质量。
Details
Motivation: 智能体强化学习在现实环境中常面临延迟和稀疏奖励的挑战,现有信用分配方法忽略了环境交互过程中产生的丰富过程信息(如交互历史),而这些信息可为识别单个动作的贡献提供有价值的监督。
Result: 在ALFWorld和WebShop基准测试中,EFCA在任务成功率和任务质量上均持续优于强基线方法,突显了基于环境反馈的多时间尺度信用分配在长期智能体强化学习中的有效性。
Insight: 创新点在于利用环境反馈直接提取短期和中长期过程信号,并通过回报重加权机制整合这些信号,为信用分配提供了更细粒度和环境基础的监督,这为处理长期稀疏奖励问题提供了新思路。
Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.
[240] Parameter Exploration for RLVR via Variational Learning cs.LG | cs.AI | cs.CLPDF
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
TL;DR: 本文提出了一种名为Perturbed Parameter Policy Optimization (3PO)的方法,通过参数空间探索来增强大型语言模型(LLM)的强化学习。该方法从后验分布中采样不同的策略来生成轨迹,从而提供了一种与动作空间探索互补的控制机制。实验表明,在数学推理和代码生成任务上,3PO方法能以相近的计算成本,持续提升下游性能,并减少训练中的无效轨迹。
Details
Motivation: 现有LLM强化学习方法(如温度缩放)主要在动作空间控制探索,这只能影响输出分布的方差而无法重新排序token,限制了探索能力并可能导致训练发散或停滞。因此,研究旨在探索参数空间作为补充的探索机制。
Result: 在OLMo-3-1025-7B和Qwen2.5-Math-7B模型上,针对数学推理和代码生成任务的实验表明,3PO方法在FLOPs成本几乎相同的情况下,持续优于标准的GRPO方法,提升了平均下游性能,并显著减少了零优势组以及格式错误或不正确的轨迹。
Insight: 核心创新点在于将探索从传统的动作空间转移到参数空间,通过从后验采样不同策略来生成多样化的轨迹。这提供了一种新的、互补的探索控制维度,能更有效地促进LLM在强化学习中的探索,从而提升学习稳定性和最终性能。
Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.
[241] Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition cs.LG | cs.CVPDF
Yani Guan, Dengpan Dong, Zi Wei, Shuang Luo, Dan Hannah
TL;DR: 该论文研究了光学化学结构识别(OCSR)任务中合成图像与真实文档之间的性能差距问题。通过在合成渲染结构和真实专利、期刊图表及手绘图像混合数据上微调21个识别器,发现真实标注训练数据是提升性能的关键因素,而视觉塔LoRA等适应策略的效果则因基础模型而异。
Details
Motivation: 解决OCSR任务中模型在合成图像上表现优异(>91%准确率)但在真实文档基准(如ACS、CLEF-IP、USPTO)上性能骤降(<16%)的“合成到真实”差距问题。
Result: 在Qwen2.5-VL模型上,仅使用9.5%真实数据即可将ACS精确匹配率从0.15提升至0.37,50.2%时达0.46;最佳配置在合成渲染图像上达到0.96精确匹配,在ACS、CLEF-IP、UOB和USPTO真实基准上分别达到0.49、0.65、0.84和0.76。
Insight: 核心创新点在于系统量化了真实训练数据对缩小合成-真实差距的决定性作用,并揭示了模型适应策略(如LoRA)的有效性高度依赖于基础模型架构,强调了针对目标任务联合选择基础模型与真实数据混合比例的必要性。
Abstract: Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS, CLEF-IP, USPTO). To identify the main source of improvement, 21 recognizers were fine-tuned on mixtures of synthetically rendered structures and labeled real depictions from patents, journal figures, and hand-drawn collections, varying the vision language model (VLM) base, the fraction of real training data, and the vision-tower adaptation strategy. Labeled real training images make the largest difference. For Qwen2.5-VL, ACS exact match rises from 0.15 with no real data to 0.37 at 9.5% and 0.46 at 50.2%; a controlled experiment across three base models reproduces the trend. A vision-tower LoRA, in contrast, does nothing for Qwen (+0.00, paired p=1.00), substantially helps InternVL3-8B (+22.8 to +34.6 pt), and modestly helps GLM-4.1V-9B (+1.0 to +9.6 pt), so its value depends on the base model. The best configuration reaches 0.96 exact match on clean renders and 0.49, 0.65, 0.84, and 0.76 on ACS, CLEF-IP, UOB, and USPTO, respectively. Gaps between base models are largest without real data (0.21), shrink to 0.06 at 70% real data, and reorder the ranking; base model and real-data mixture must therefore be selected together. Small-scale experiments on handwritten image-to-LaTeX recognition and chart-to-table conversion show that base-model rankings also vary beyond chemistry. More generally, model and adaptation choices for visual structure recognition should be evaluated on the target task.
[242] DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models cs.LG | cs.CVPDF
Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han
TL;DR: 本文提出DreOPD方法,一种用于流匹配模型的降级参考外推式策略蒸馏方法,旨在解决流匹配模型在下游任务适应中存在的任务间冲突问题。该方法通过将隐式奖励外推转化为闭式速度回归,结合策略蒸馏的稳定性与强化学习的直接优化能力,实现了更高效的任务特定优化。
Details
Motivation: 流匹配模型作为主流图像生成方法,在下游任务适应时通常依赖后训练,可能导致任务特定优化目标间的冲突;强化学习虽能直接优化任务奖励,但轨迹级优化存在高方差梯度和跨任务干扰问题,而传统策略蒸馏仍基于模仿学习,缺乏外推能力。
Result: 在单教师和多教师设置下的实验表明,DreOPD在平均性能上优于策略蒸馏和多任务强化学习基线,并在大多数指标上超越了专用教师模型。
Insight: 创新点在于将隐式奖励外推转化为闭式速度回归,实现了外推式后训练与策略蒸馏稳定性的结合;同时引入轻度降级参考以增强教师-参考对比,从而提供更清晰的外推方向,为流匹配模型的多任务优化提供了新思路。
Abstract: Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.
[243] Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation cs.LG | cs.AI | cs.CVPDF
Hossein Goli, Farzan Farnia, Amin Gohari
TL;DR: 本文提出了想象生成式AI(IGA)框架,通过引入谱熵作为多样性度量,将多样性设计纳入目标分布优化中。IGA通过熵约束投影,在保持与参考分布接近的同时,控制生成分布的谱多样性达到预设水平,从而在数据分布的熵墙之下进行多样性修复,在熵墙之上实现超越数据多样性的想象生成。
Details
Motivation: 现有生成式AI模型主要模仿数据分布,无法纠正生成器丢失的多样性,也未定义如何生成超越数据本身多样性的内容。
Result: 在合成和视觉基准测试中,IGA在熵墙之下实现了多样性修复,在熵墙之上实现了可控的谱外推。
Insight: 创新点在于将谱熵作为与参考无关的表示引导多样性度量,并提出了统一的熵约束投影框架,实现了从模仿到想象的连续调控;提出的IGA Guidance方法无需重新训练,可直接用于基于分数和扩散模型的推理时生成。
Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level. Diversity is measured by the von Neumann entropy of the generated distribution’s kernel covariance operator in a fixed representation space, providing a reference-free representation-guided measure of how broadly probability mass occupies embedding directions. The spectral entropy of the population data distribution defines an Entropy Wall. Below the wall, IGA performs diversity repair, recovering variation that a learned generator has lost while remaining within the diversity level of the data. Beyond the wall, the data distribution itself becomes infeasible, and IGA deliberately departs from it to produce distributions with greater representation-relative spectral diversity, an operational notion of imaginative generation. These regimes form a single regularization path from imitation to imagination and define an i.i.d. target distribution at each prescribed diversity level. We develop the theory of this entropy-constrained projection and show that, under a KL anchor to a pretrained generator, the optimum satisfies a self-consistent exponential-tilt relation. This characterization leads to IGA Guidance, a retraining-free inference-time method for score-based and diffusion models, including DDPM and DDIM samplers. Experiments on synthetic and vision benchmarks demonstrate diversity repair below the Entropy Wall and controlled spectral extrapolation beyond it.
cs.MM [Back]
[244] World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World cs.MM | cs.CV | cs.CY | cs.HCPDF
Adam Cole, Mick Grierson
TL;DR: 这篇论文通过艺术装置《世界模拟器》批判了AI视频模型自诩为’世界模拟器’的傲慢主张。作者将同性色情内容输入AI视频生成管道,由于模型缺乏相关训练数据,导致其产生荒诞的幻觉式输出,将亲密场景扭曲为厨房电器、怪异建筑等平庸图像。作品揭示了当前AI模型在模拟人类体验(尤其是性体验)方面的结构性盲区,并质疑了’模拟’本身的价值取向。
Details
Motivation: 针对AI视频模型被宣传为能模拟无限现实的’世界模拟器’这一现象,论文指出这些模型系统性地排除了人类具身体验(特别是性体验)的重要方面,旨在通过艺术实践揭示其宣称的普适性与实际能力之间的根本矛盾。
Result: 通过将同性色情视频输入AI视频到视频生成流程,模型因缺乏训练数据而产生了超现实的幻觉输出,将亲密行为转化为厨房用具、奇异建筑和抽象肉体的平庸场景,这从定性角度可视化了合成知识的局限性。
Insight: 创新点在于采用艺术实践的方法,通过’数据对抗性输入’(输入模型训练分布外的敏感内容)来暴露AI模型的结构性盲区;更深层的洞见是批判了以大规模过滤数据集训练的AI系统自称’世界模拟器’的认识论傲慢,并提出了对模拟范式本身的哲学性质疑,探索更具生命肯定性的感官表征可能性。
Abstract: Increasingly, AI video models are marketed as “world simulators,” suggesting their ability to model infinite realities. Despite such claims, these models systematically exclude significant aspects of embodied human experience, particularly sexuality. World Simulator is a video installation exploring the poetic friction between these universal claims and the models’ inherent blindness. To do so, the work feeds explicit gay erotica into an AI video-to-video pipeline. Lacking the training data to recognize these images, the system hallucinates surreal alternatives, transforming intimate acts into banal scenes of kitchen appliances, strange architectures, and abstract flesh. By visualizing the limits of synthetic knowledge, the work challenges the hubris of the “world simulator” label, asking how a system, trained primarily on large filtered video datasets, can claim to simulate the world while remaining structurally blind to the body. Beyond this critique, we question the value of simulation itself, asking what forms of sensual representation might offer more expansive, life-affirming possibilities.
cs.CR [Back]
[245] Adversarial Attacks on Deep OCR Systems cs.CR | cs.AI | cs.CVPDF
Wenbo Sun, Hongzong LI, Yanyun Wang, Jiahao MA, Shuxin Zhuang
TL;DR: 本文首次提出了一种针对生成式OCR视觉语言模型(如DeepSeek-OCR)的纯黑盒对抗攻击方法。该方法仅利用解码后的字符串输出,通过将攻击重构为零阶优化问题,并使用基于序列相似度的有界标量损失和随机方向有限差分方案来估计梯度,从而在图像上生成难以察觉的扰动。初步实验验证了攻击的有效性,并揭示了模型在解码时出现的严重定性故障。
Details
Motivation: 动机在于,尽管DeepSeek-OCR等先进OCR模型通过将视觉模态视为光学压缩介质提升了文档识别能力,但其增加的复杂性可能引入新的安全漏洞。本文旨在探索在纯黑盒(仅能查询解码字符串)设置下,对这些模型的对抗攻击可行性。
Result: 在Deep-OCR上的初步实验验证了仅基于字符串的攻击和评估流程的有效性,暴露了模型解码器在遭受攻击时出现的重复、截断和提示泄露等严重定性故障。实验表明,实现受控的目标重写比无目标性能退化要困难得多,因此作者避免在预注册评估完成前宣称目标攻击成功。
Insight: 创新点在于首次针对生成式OCR VLM设计了纯黑盒对抗攻击框架,将攻击问题重构为基于字符串序列相似度损失(而非传统图像分类损失)的零阶优化问题,并采用了查询成本与图像维度无关的随机方向梯度估计方法。这为评估类似复杂视觉语言模型的安全性提供了新思路和基准。
Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We recast the attack as a zeroth-order optimization problem driven by a bounded scalar loss defined directly on the string output via sequence similarity, and estimate the gradient with a random-direction finite-difference scheme whose query cost is independent of the image dimension. An Adam update with ell_infinity projection yields imperceptible perturbations for both untargeted and targeted objectives. Pilot experiments on Deep-OCR validate the string-only attack and evaluation pipeline and expose severe qualitative decoder failures, including repetition, truncation, and prompt leakage. They also show that controlled targeted rewriting remains substantially harder than untargeted degradation; we avoid claiming targeted success until the pre-registered evaluation is complete.