Table of Contents
- cs.CL [Total: 18]
- cs.CV [Total: 61]
- cs.GR [Total: 1]
- cs.SE [Total: 2]
- cs.LG [Total: 1]
- eess.AS [Total: 1]
- cs.RO [Total: 6]
- cs.AR [Total: 1]
- cs.IR [Total: 1]
- eess.IV [Total: 1]
- cs.CR [Total: 1]
- cs.AI [Total: 5]
cs.CL [Back]
[1] Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks cs.CL | cs.AIPDF
Yibo Hu, Jiaming Qu
TL;DR: 本文研究表明,大语言模型(LLM)在标准同侪压力基准测试中表现出的从众行为,大部分并非源于对说话者(如专家小组)的社会影响,而是由重复的错误答案本身这一线索所驱动。通过引入一个‘无来源’条件(即移除明确的说话者,仅呈现被断言的答案),作者发现即使在移除说话者后,模型仍会频繁地将原本正确的答案修改为错误答案。
Details
Motivation: 现有同侪压力基准测试存在一个混淆因素:提示词同时包含了‘说话者的存在’和‘重复的错误答案’两个线索,因此无法区分模型修改答案在多大程度上真正依赖于社会影响(说话者),而非仅仅是文本重复。
Result: 在六个开源LLM和七个QA与推理数据集上的实验表明,在‘无来源’条件下,66.5%的初始正确答案被有害地修改,远高于简单重问条件下的10.3%。即使对重复答案进行转述或在开放式设置中隐藏选项,该效应依然存在。
Insight: 核心创新点在于提出了‘无来源条件’作为测量基准,揭示了LLM从众行为中存在一个显著的‘无说话者基线’。这为准确测量社会影响力(应被视为超出此基线的增量)提供了方法论上的重要启示,即评估从众行为时,必须先测量移除说话者后剩余的影响,否则可能将文本重复误判为社会影响。
Abstract: LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even after the peer is removed. The reason is a confound: standard conformity prompts mix two cues at once, the presence of a speaker and the repeated wrong answer itself. Existing benchmarks vary these cues together, so they cannot tell how much of the revision actually depends on the speaker. We introduce a no-source condition: the same asserted answer with the explicit speaker removed. Across six open-weight LLMs and seven QA and reasoning datasets, this condition alone causes harmful revision in $66.5%$ of initially correct cases, compared with $10.3%$ under a plain re-ask. The effect also remains when the repeated answer is paraphrased and when answer options are hidden in an open-ended setting. Source framing mainly modulates this floor: expert-panel framing raises it, while minimal person labels do not reliably raise it. When models flip, they are usually confidently wrong, and simple recalibration does not recover the original answer. Source attribution still matters, but it should be measured as an increment above this speaker-free floor. The methodological lesson is that conformity benchmarks should first measure what remains after the speaker is removed; without this step, benchmarks may mistake repeated text for social influence.
[2] The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment cs.CL | cs.AI | cs.CYPDF
Haonan Huang
TL;DR: 这篇论文通过设计心理测量工具(交叉对称化)来分解大型语言模型在道德困境问题中表现出的’是-否’偏差,发现这种偏差主要源于答案顺序和词汇表面特征,而非模型内在道德判断的偏移。研究揭示了前沿模型的道德立场在不同问题形式下具有高度一致性,而小模型则表现出模型特定的失败模式。
Details
Motivation: 针对现有文献报告大型语言模型在道德困境等二元判断任务中,其输出会因逻辑无关的措辞变化(如问题形式)而发生偏移,尤其是表现出人类所没有的放大的’是-否’偏差,论文旨在探究这种偏差的本质来源,区分逻辑判断、词汇标记和打印顺序等因素的影响。
Result: 在道德困境语料库上应用交叉对称化方法后,前沿模型(如GPT-4o、Claude、Gemini)的道德立场θ在不同等价问题形式间几乎不变(跨形式不一致性在±1轴上仅为0.12-0.21)。强制’是/否’判决时产生的偏差可分解为:倾向于最后打印选项的顺序偏差(与人类的首因效应相反)和对’否’这个词的词汇偏好;该偏差在Claude模型中显著(故事平均-0.32至-0.86),在GPT-4o和Gemini中接近0,且随着推理链延长而减小。将词汇替换为任意标签后,与判决相关的逻辑偏差在所有前沿模型中均接近0。
Insight: 创新点在于提出了’交叉对称化’这一心理测量方法,能够系统性地分离和量化逻辑判决、词汇表面特征和答案顺序等因素对模型输出的影响。核心发现是,模型表现出的’是-否’偏差主要反映了对问题表面形式(如词汇和顺序)的敏感性,而非其内在道德判断的实质性变化。这为准确评估模型的价值观提供了一种方法论:必须通过交叉变换问题框架来测量,而非单次询问。论文还提出了一个简约模型(P = σ((θ±m)/s))来概括任何此类偏差,其中框架敏感性m和道德决断力s可与采样温度明确区分。
Abstract: Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word “no” is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models’ stance $θ$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word “no”; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = σ((θ\pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
[3] BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension cs.CL | cs.AI | cs.CVPDF
Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha, Ezharuddin Jubaer
TL;DR: 本文介绍了BaFCo,一个针对孟加拉语表单理解的基准数据集,旨在解决低资源语言高质量标注数据稀缺的问题。该数据集包含200份来自孟加拉国政府多个领域的复杂多页表单,并定义了细粒度(26类)和粗粒度(5类)的实体标注模式。研究评估了多个主流多模态大语言模型在零样本和思维链提示下的表现,发现它们在理解孟加拉语表单、特别是精确定位细粒度实体方面存在局限。
Details
Motivation: 多模态大语言模型在文档理解任务上面临挑战,而孟加拉语等低资源语言因缺乏高质量标注数据,限制了相关模型在实际应用中的采用。
Result: 在BaFCo基准上评估了ChatGPT、Gemini、Claude、Qwen和Kimi系列的最新MLLMs,结果显示当前模型在理解孟加拉语表单、尤其是精确定位细粒度表单实体方面能力有限。
Insight: 创新点在于构建了首个专注于复杂孟加拉语表单理解的细粒度标注基准数据集,并系统评估了现有MLLM在低资源语言文档理解任务上的局限性,为后续研究提供了数据基础和性能基线。
Abstract: Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limited for low-resource languages such as Bangla due to the scarcity of high-quality annotated data. To address this gap, we introduce BaFCo, a benchmark dataset for Bangla form comprehension with a focus on Document Layout Analysis (DLA) and Key Information Extraction (KIE). BaFCo curates 200 multi-page complex Bangladeshi government forms, sourced from across diverse sectors including agriculture, education, banking, and land management. To accurately capture the structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, along with a separate coarse form entity set consisting of 5 types. We evaluate the latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimi series using zero-shot and chain-of-thought prompts under both low and high reasoning setups. Our results reveal limitations in current MLLMs’ ability in comprehending Bangla forms, particularly in accurately localizing highly granular form entities. Our dataset and code is available at: https://huggingface.co/datasets/Mausul/bafco
[4] NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task cs.CLPDF
Anand Kamble, Aniket Tathe
TL;DR: 本文重新实现了NAVER LABS的IWSLT 2025指令跟随流水线,以参加IWSLT 2026共享任务(受限条件、短音频轨道)。系统采用SeamlessM4T-v2-large作为语音编码器,Qwen3-4B-Instruct作为LLM主干,并保留了原有的三阶段训练方法。此外,作者从给定语料中构建了10万条涵盖10种语音中心任务的合成指令跟随示例,用于第三阶段的微调。
Details
Motivation: 动机是适应IWSLT 2026共享任务的新约束条件(指定组件),并在原有流水线基础上进行改进和扩展,以提升指令跟随语音翻译系统的性能。
Result: 主要模型在MCIF基准测试上取得了COMET 0.781(英-中语音翻译)和BERTScore-F1 0.346(英语SQA)的成绩。
Insight: 创新点在于在指定组件约束下成功复现并改进了三阶段流水线,并通过构建大规模、任务类型多样化的合成指令数据集来增强模型的指令跟随能力,这为数据增强和多模态指令微调提供了可借鉴的思路。
Abstract: We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder and Qwen3-4B-Instruct as the LLM backbone. The three-stage approach projector alignment, text-only LoRA pre-training, and multimodal merging is preserved from the original design. We additionally construct 100k synthetic instruction-following examples across ten speech-centric task types (10k per task) from the provided corpora, suitable for further Stage 3 fine-tuning. Our primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.
[5] When Should LLMs Search? Counterfactual Supervision for Search Routing cs.CL | cs.AIPDF
Minho Kim
TL;DR: 该论文研究了检索增强语言模型中的搜索路由问题,即决定何时进行外部搜索以提升任务成功率。作者通过对比同一问题在无搜索和强制搜索下的结果,构建了一个包含‘无需搜索’、‘需要搜索’和‘无法解决’三种状态的监督信号。利用该信号进行监督微调和偏好优化训练,显著提升了Gemma和Qwen模型的路由性能。
Details
Motivation: 检索增强语言模型并非总是需要搜索,不当的搜索调用(如对已知问题搜索或依赖噪声证据)会降低效率与准确性。因此,需要解决实例级别的搜索路由决策问题,以判断搜索是否能真正改善任务成功率。
Result: 在符合监督条件的样本上,路由的宏观F1分数得到显著提升:Gemma E2B从0.7082提升至0.8235,Qwen3.5-4B从0.7053提升至0.8365。这表明所提方法能有效改进搜索路由决策。
Insight: 核心创新点在于利用反事实监督(对比无搜索与强制搜索结果)构建路由决策的监督信号,将路由问题形式化为一个可学习的策略。客观分析表明,该方法能针对不同模型特性(如Gemma学习克制搜索,Qwen减少漏搜)进行针对性优化,并揭示了模型能力、检索预算、证据利用等多种残余瓶颈。
Abstract: Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly beneficial: models may call search for questions they can already answer, or rely on noisy evidence when correction, clarification, or abstention would be more appropriate. We formulate this as an instance-level search-routing problem: deciding whether search is needed to improve task success relative to a no-search execution. To derive supervision, we compare no-search and forced-search outcomes for the same question and construct an oracle over NO SEARCH, SEARCH, and UNSOLVED based on task-specific success. Using this oracle as both an evaluation criterion and a learning signal, we train search-routing policies with supervised fine-tuning and preference optimization, improving routing macro-F1 on oracle-eligible examples from 0.7082 to 0.8235 for Gemma E2B and from 0.7053 to 0.8365 for Qwen3.5-4B. Further analysis shows that the learned policies reduce model-specific routing failures: Gemma primarily learns no-search restraint, while Qwen further reduces missed search; residual UNSOLVED cases reveal heterogeneous bottlenecks involving model capacity, retrieval budget, evidence use, and policy behavior.
[6] Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization cs.CL | cs.LGPDF
Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong
TL;DR: 本文研究了大型推理模型在事实性问答任务中,由于生成显式思维链而可能导致的‘思维诱发幻觉’问题,即思维链有时会推翻原本正确的直接答案并引入事实性错误。为此,作者提出了一个名为MARGO的强化学习框架,通过混合思维与非思维轨迹来评估思维链是否增加了事实价值,从而抑制有害的思维行为并保留有益的推理能力。
Details
Motivation: 动机在于发现大型推理模型在事实性问答中,显式思维链的引入并非总是有益的,它有时会颠覆模型原本正确的直接答案,导致事实性漂移,即‘思维诱发幻觉’。
Result: 在多个事实性问答基准测试上的实验表明,MARGO相比强基线模型提高了事实可靠性;在数学推理基准测试上的评估则显示,它保留了模型的一般推理能力。
Insight: 创新点在于将显式思维建模为对模型直接回答倾向的‘思维残差’,并提出了一个新颖的强化学习框架MARGO,它利用非思维轨迹作为同模型参考来构建混合模式的优势估计,从而有选择性地抑制容易产生幻觉的思维行为。
Abstract: Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to factual drift. We refer to this failure mode as \emph{thinking-induced hallucination}. To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model’s direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, \underline{\textit{M}}ixed-Mode \underline{\textit{A}}dvantage \underline{\textit{R}}egularization for \underline{\textit{G}}rounded \underline{\textit{O}}ptimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation. By constructing mixed-mode rollout groups with both thinking and non-thinking trajectories, MARGO evaluates whether explicit thinking adds factual value beyond direct answering, thereby suppressing hallucination-prone thinking while preserving beneficial thinking behaviors. Experiments across multiple factuality-oriented QA benchmarks demonstrate that MARGO improves factual reliability over strong baselines, while evaluations on mathematical benchmarks show that it preserves general reasoning ability.
[7] Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding cs.CLPDF
Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich
TL;DR: 论文提出了Nemotron-Labs-Diffusion,一种三模态语言模型,将自回归、扩散和自推测解码统一在单一架构中。该模型通过联合自回归-扩散目标训练,能够切换模式以在不同部署设置和并发级别下维持高吞吐量。研究展示了三种模式的优势互补,并在扩展到3B、8B和14B参数时,在准确性和速度上均优于最先进的开源自回归和扩散模型。
Details
Motivation: 为了解决单一解码模式(如自回归)在效率或规划能力上的局限性,并探索不同生成范式(自回归与扩散)的互补性,以构建一个能适应不同部署需求的高效统一模型。
Result: 在SPEED-Bench等基准测试中,模型家族(包括基础、指令和视觉语言模型)在准确性和速度上均超越了最先进的开源自回归和扩散模型。例如,Nemotron-Labs-Diffusion-8B在保持可比准确性的同时,每次前向传播生成的token数量是Qwen3-8B的6倍,在GB200 GPU上使用SGLang实现了4倍高的吞吐量。
Insight: 核心创新在于将自回归、扩散和自推测解码统一到一个模型中,并利用联合训练实现模式互补。具体而言,扩散模型增强了前瞻规划能力,而自回归模型提供了从左到右的语言先验;在自推测模式下,扩散模型负责草稿生成,自回归模型负责验证,这比多token预测方法在接收率和实际设备效率上更优。这种统一架构为在不同场景下平衡生成质量与效率提供了新思路。
Abstract: We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion’s long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.
[8] PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages cs.CL | cs.AIPDF
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický
TL;DR: 本文介绍了PluraMath数据集,它是PolyMath的扩展,旨在解决现有数学推理评估基准在高资源语言(如英语和中文)上的偏见问题。该数据集新增了18种代表性不足的语言,涵盖6个语系,包括中低资源到极低资源语言,并通过人工验证的流程构建。作者使用PluraMath对27个推理大语言模型进行了基准测试,分析了多语言数学推理能力在不同语言条件下的表现。
Details
Motivation: 现有数学推理评估基准(如PolyMath)主要偏向高资源语言,覆盖范围有限,无法全面评估大语言模型在代表性不足语言上的推理能力,因此需要扩展数据集以填补这一空白。
Result: 在PluraMath数据集上对27个模型(涵盖小、中、大和闭源集成模型)的基准测试显示,高资源语言与代表性不足语言之间的数学推理性能存在持续差距,且更好的指令遵循能力与更强结果相关。
Insight: 创新点在于将数学推理评估扩展到更多低资源语言,通过人工验证流程确保数据质量,并开源数据集和评估框架以促进多语言基准开发;客观分析认为,这有助于揭示模型在语言多样性下的性能瓶颈,推动更公平的AI评估。
Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families – ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales – small, mid-size, large, and closed-source ensembles – probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.
[9] InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost cs.CL | cs.AI | cs.IRPDF
Krittanon Kaewtawee, Petmongkon Pornpichitsuwan, Natchaya Temyingyong, Nutnicha Laplamoon, Wachiravit Modecrua
TL;DR: 论文提出了InfluMatch,一个低成本的三阶段级联系统(检索→重排→推理),用于匹配泰国营销标准下的关键意见领袖(KOL)。该系统完全基于小型开源模型构建,通过密集检索、4B参数的重排器和4B参数的推理器,在保持高准确性的同时大幅降低了计算成本。
Details
Motivation: 当前KOL匹配方法存在语义匹配缺失(基于关键词搜索)或成本高昂(使用前沿大语言模型处理每个候选者)的问题,需要一种既准确又经济高效的解决方案。
Result: 在包含11个查询和50个候选者的数据集上,完整级联系统达到了94.1%的P@5,与前沿模型Kimi-K2.6(91.8%)性能相当,但输出令牌减少约35倍,在单张A100上处理50个KOL查询仅需约20秒。
Insight: 创新点在于设计了一个成本优化的级联架构,通过过滤减少推理开销;研究发现,仅对重排器进行成对微调(SimPO)即可匹配前沿模型的最佳选择准确率,而对推理器进行逐点微调反而会降低端到端排名性能,这揭示了任务设计对模型部署有效性的关键影响。
Abstract: Matching influencers (KOLs) to free-form, multi-part Thai marketing criteria is today served either by keyword search over structured profiles, which misses semantic fit, or by prompting frontier LLMs over every candidate, which is accurate but slow and expensive. We present InfluMatch, a low-cost three-stage cascade – retrieval $\rightarrow$ rerank $\rightarrow$ reason – built entirely from small open-weight models: dense retrieval returns 50 candidates, a 4B pointwise reranker scores each by the log-probability of a single Yes token and keeps 10, and a 4B reasoner grades the shortlist per criterion on a rubric with a Thai rationale. The cascade is designed for cost: reasoning over a filtered top-10 halves token spend versus reasoning over all 50 while scoring 14 points higher. End-to-end against human relevance labels on an 11-query set with all 50 candidates labeled, the full cascade reaches 94.1% P@5, versus a retrieval-only baseline near random; it matches the frontier model Kimi-K2.6 (91.8%) while emitting ${\sim}35\times$ fewer output tokens and serving a 50-KOL query in ${\sim}20$ s on one A100. Notably, the only fine-tuning that pays off is pairwise: a SimPO-tuned reranker matches the frontier baseline’s best-pick accuracy (78.0 EM), whereas fine-tuning the reasoner on pointwise per-criterion labels improves offline scores yet degrades end-to-end ranking – an inversion we trace to the design of the absolute labeling task – leaving the untuned base model as the strongest deployed reasoner. The result is a deployable, explainable KOL search system at a small fraction of frontier serving cost.
[10] Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability cs.CLPDF
Tamara Wit, Lifeng Han, Carly Heipon, David Lindevelt, Anne Stiggelbout
TL;DR: 本文提出了LLM4SDM,首次研究了使用开源小语言模型(OS-sLLMs)基于Observer OPTION12框架自动评估共享决策(SDM)。与以往依赖大型商业模型和较短OPTION5工具的研究不同,本研究专注于可本地部署的隐私保护模型和荷兰黑色素瘤咨询转录本。通过专家标注的临床咨询数据,在开发阶段试点研究中评估了三个通用领域和两个医学领域的OS-sLLMs。结果表明通用领域模型优于医学领域模型,后者存在严重的幻觉和指令遵循失败问题。Gemma3:12b模型与人工标注的一致性最强。研究还引入了Judge-LLM共识框架以解决模型间的分歧。
Details
Motivation: 解决以往共享决策自动评估依赖大型商业模型(存在隐私和可持续性问题)和较短评估工具(OPTION5)的局限性,探索更注重隐私保护、可本地部署的开源小语言模型在更全面的OPTION12框架下的应用潜力。
Result: 在荷兰黑色素瘤咨询转录本上的实验表明,通用领域OS-sLLMs(如Gemma3:12b,与人工标注的Pearson r=0.51,Spearman ρ=0.59)优于医学领域模型。当前模型尚不能替代人工标注员,但为隐私保护的人机协同评估提供了有希望的基础。
Insight: 创新点在于首次系统评估OS-sLLMs用于OPTION12框架下的SDM自动评估,并揭示了模型在时间话语推理、对话角色归因和证据 grounding 方面的系统性挑战。提出的Judge-LLM共识框架为解决多模型分歧提供了新思路,强调了在专业领域(如医疗)中,通用模型可能比领域特定模型表现更好(后者易产生幻觉),这为未来轻量级、隐私友好的辅助工具开发提供了重要方向。
Abstract: We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and the shorter OPTION5 instrument, our study focuses on privacy-preserving locally deployable models and Dutch melanoma consultation transcripts. Using expert-annotated clinical consultations, we evaluate three general-domain and two medical-domain OS-sLLMs during a development-phase pilot study. Results show that general-domain models outperform medical-domain models, which exhibit substantial hallucination and instruction-following failures. Gemma3:12b achieves the strongest agreement with human annotations (Pearson r=0.51, Spearman \r{ho}=0.59). Item-level and qualitative analyses reveal systematic challenges related to temporal discourse reasoning, conversational role attribution, and evidence grounding. We further introduce a Judge-LLM consensus framework designed to support disagreement resolution among multiple models. Our findings suggest that while current OS-sLLMs cannot replace human annotators, they offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment.
[11] CurateEvo: Data-Curation Evolving for Agentic Post-Training cs.CLPDF
Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu, Wanxiang Che
TL;DR: 本文提出了CurateEvo,一个用于智能体后训练数据管理的动态演化框架。该框架将数据管理策略编码为可执行代码,并利用开发集中的失败轨迹进行迭代重写,以生成监督微调、强化学习和推理时记忆库所需的数据。
Details
Motivation: 现有智能体后训练流程通常将数据管理视为固定的预处理步骤,主要关注数据增强,而忽略了数据过滤、精炼以及对下游失败的适应性。本文旨在解决这一问题,通过动态演化来优化数据管理策略。
Result: 在ACEBench-Agent、BFCL-V4和τ^2-Bench基准测试上,无论是标注数据还是野生数据设置,CurateEvo均优于现有数据管理方法,平均得分分别提高了3.2和2.7个点。
Insight: 核心创新在于将数据管理策略视为可演化的代码,并通过失败驱动的迭代重写过程,实现了策略在效果(诊断失败模式并相应调整数据)和效率(在成本感知目标下剪枝冗余数据)两方面的持续优化。该框架与不同的后训练方案兼容,并能显著降低管理开销。
Abstract: Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fixed preprocessing step, focusing mainly on data augmentation while neglecting filtering, refinement, and adaptation to downstream failures. We propose CurateEvo, a failure-driven dynamic evolution framework for agentic post-training data curation. CurateEvo represents the curation strategy as executable code and iteratively rewrites it using failed trajectories from a held-out development set. At each epoch, the evolved strategy transforms a fixed raw corpus into supervised fine-tuning data, reinforcement learning data, and an inference-time memory bank. The evolution process first improves effectiveness by diagnosing recurring failure modes and augmenting, filtering, or refining data accordingly, and then improves efficiency by pruning redundant or low-utility training turns under a cost-aware objective. Experiments on ACEBench-Agent, BFCL-V4, and τ^2-Bench under both labeled and wild-data settings show that CurateEvo consistently outperforms prior curation methods, improving average scores by 3.2 and 2.7 points, respectively. Further analyses demonstrate that CurateEvo is compatible with different post-training recipes and substantially reduces curation overhead.
[12] LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability cs.CL | cs.AIPDF
Chenxu Wang, Yongkun Yang, Boyuan Du, Shiwei Lin, Huaping Liu
TL;DR: 本文研究了在部分可观测环境下,大语言模型(LLM)智能体如何进行审慎协作以实现联合决策。作者形式化了这一协作问题,并引入了一个可扩展的基准测试来评估多种LLM。研究发现,即使借助外部工具,当前最先进的LLM在信息对齐和复杂推理方面仍面临挑战,但审慎过程也可能提供反思和纠错的机会。
Details
Motivation: 研究动机是探索LLM智能体如何在信息部分可观测且不对称的协作任务中,通过类似人类的审慎沟通来对齐信息并达成一致决策。
Result: 在作者构建的多任务、多领域基准测试上进行的系统评估表明,复杂的审慎协作任务对当前最先进的LLM仍构成挑战,其表现可能不如集中式基线,但审慎过程有时能通过反思提升性能。
Insight: 论文的创新点在于形式化了LLM智能体的审慎协作问题,并提供了一个用于评估的基准测试和参考框架。客观来看,其核心洞察是揭示了当前LLM在协作决策中的关键瓶颈(信息对齐与复杂推理)以及审慎过程本身可能带来的纠错潜力,为改进多智能体系统提供了方向。
Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable benchmark that instantiates this problem across multiple task settings and domains in which agents must exchange information through deliberation to reach a joint decision with a shared reward. We then instantiate a reference scaffold and evaluation protocol for deliberative agents and conduct a systematic evaluation of a range of representative LLMs. The results reveal that complex deliberative collaboration tasks continue to challenge state-of-the-art language models. Even with the aid of external mathematical tools, language models may fail in either the deliberation process for aligning information or the complex reasoning process for making the decision. On the other hand, diagnostic analysis reveals that the deliberation process may also provide opportunities for reflection and error correction, sometimes improving performance over centralized baselines. Altogether, our work establishes a foundation for evaluating and improving LLM agents in deliberative collaboration and provides insights into the strengths, limitations, and properties of current LLM-based multi-agent systems.
[13] LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis cs.CL | cs.AIPDF
Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu
TL;DR: 本文提出了LongCrafter框架,用于合成多样化的长上下文监督微调数据,以增强大语言模型的长上下文理解能力。该框架结合了分层任务分类法和基于证据的合成流程,通过构建任务对齐的长上下文、分解为显式证据图并生成严格基于证据的指令-响应对,解决了现有方法任务覆盖窄、指令难度不足和缺乏忠实性监督的问题。
Details
Motivation: 现有长上下文SFT数据合成方法存在任务覆盖范围窄、指令难度不足以及缺乏对模型推理忠实性的监督这三个主要局限,限制了模型长上下文理解能力的提升。
Result: 在LongBench、LongBench v2和LooGLE基准测试上,使用LongCrafter数据微调的Qwen2.5-7B和LLaMA-3.1-8B模型超越了所有SFT基线模型甚至官方后训练模型,尤其是在高难度任务上提升最大。
Insight: 创新点在于将长上下文理解任务组织成层次化、细粒度的分类法作为全局生成先验,并引入证据图来建模跨段落依赖关系,确保指令生成的忠实性和可追溯性,从而有效缓解了模型在处理长文本时’迷失在中间’的问题。
Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction–response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench~v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the ``lost in the middle’’ problem.
[14] Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design cs.CL | cs.AI | cs.LGPDF
Alexander Rombach, Chantale Lauer, Nijat Mehdiyev
TL;DR: 本文系统研究了基于强化学习的流程模型生成中奖励函数设计问题,通过训练两种LLM家族(Llama 3.1 8B和Qwen 2.5 14B)在48种配置下,使用包含句法、语用和语义质量共38个指标的自动评估框架来设计奖励函数。研究发现:强化学习能显著提升语用和句法质量并保持语义保真度;等权重奖励分配优于针对性加权;设计选择与模型架构存在非平凡的交互作用。
Details
Motivation: 解决大型语言模型在从自然语言描述生成BPMN流程模型时,监督微调受限于训练数据模式的问题,探索如何通过强化学习利用外部质量度量进行优化,特别是在质量是多维度的场景下奖励函数应如何设计。
Result: 在自动评估框架(38个指标)上,强化学习显著提升了语用和句法质量,同时保持了语义保真度,并将输出变异性降低了六倍以上;等权重奖励分配始终优于针对性加权。
Insight: 奖励函数的组合是优化结果的主要决定因素,其影响程度与是否应用强化学习本身相当;研究结果可推广到任何基于多维度自动评估的结构化生成任务;模型架构与奖励设计选择(如无效性惩罚、SFT初始化)存在关键交互,需针对性调整。
Abstract: Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external quality measures, but how the reward function should be designed when quality is multi-dimensional remains unexplored. We present a systematic investigation of reward function design for RL-based process model generation, training two LLM families (Llama3.1 8B, Qwen2.5 14B) under 48 configurations using Group Sequence Policy Optimization with rewards derived from an automated evaluation framework comprising 38 metrics across syntactic, pragmatic, and semantic quality. Three findings emerge. First, RL significantly improves pragmatic and syntactic quality while preserving semantic fidelity, reducing output variability by more than sixfold. Second, equal reward weighting consistently outperforms targeted weighting: emphasizing a specific dimension fails to improve it and can collapse the model into a low-quality mode. Third, design choices interact with model architecture in non-trivial ways: the invalidity penalty is essential for one model but irrelevant for the other, and SFT initialization is indispensable for one architecture but counterproductive for another. These results demonstrate that reward composition is a primary determinant of optimization outcomes, with effects as large as the decision to apply RL itself. The findings generalize to any structured generation task where quality is assessed along multiple automated dimensions. We release our implementation and experimental code at https://github.com/chlauer99/RL_for_process_modeling.
[15] Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability cs.CL | cs.CYPDF
Alicia Parrish, Rajat Shinde, Sanket Badhe, Xinyi Bai, Sree Bhargavi Balija
TL;DR: 本文介绍了Pluralis v0.1,一个从文化优先视角构建的新型多模态、多区域、多语言数据集,旨在评估AI(特别是视觉语言模型)在全球部署中的风险与可靠性。该数据集覆盖亚太地区六个国家和八种语言,包含6448个提示,通过文本和图像的协同作用来触发特定法律或文化违规,从而揭示现有西方中心评估框架的盲点。
Details
Motivation: 当前AI安全评估和基准测试框架主要依赖西方中心、文化无关的默认设置,掩盖了关键的地方法律、社会语言细微差别和文化禁忌,导致视觉语言模型在全球部署中存在脆弱性。
Result: 通过在Pluralis数据集子集上的观察,发现了视觉语言模型存在重复的、特定于地区的故障模式,如图像误识别导致下游危害、遗漏物品-上下文-地区交互以及不充分的拒绝响应。这些故障模式因地区和语言不同而系统性变化,暴露了全局平均指标所掩盖的盲点。
Insight: 论文的主要创新点在于从文化优先视角原生构建本地化安全风险数据集,而非改编西方数据集,并引入了一种多模态评估范式(文本与图像协同触发违规)。它将本地文化适宜性确立为一流的评估轴,并提出了Judge-Pluralis(一个基于经验推导的文化分类法训练、以一致性为门控的LLM-as-a-Judge集成方法)来操作化这一评估。这为推进多语言、多文化评估科学以更好地支持全球AI文化对齐奠定了基础。
Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language Models (VLMs) vulnerable in global deployments. We introduce Pluralis v0.1: a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective. Spanning 6,448 prompts across six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, Taiwan) and eight languages, Pluralis diverges from prior work by natively sourcing localized safety hazards rather than adapting Western datasets. Crucially, it introduces a multimodal evaluation paradigm: user text (e.g., “Should I gift this?”) and an image referring to “this” (e.g., a clock) - both innocuous in isolation, but synergistically triggering specific legal or cultural violations. Pluralis disentangles universal safety violations from localized cultural appropriateness, establishing the latter as a first-class evaluation axis. To operationalize this, we present Judge-Pluralis, an agreement-gated LLM-as-a-Judge ensemble trained on examples classified in an empirically derived cultural taxonomy. Observing VLM behavior on a subset of the Pluralis surfaces recurring, locale-specific failure modes such as image misidentifications with downstream harm, missed item-context-locale interactions, and inadequate refusals. These failure modes vary systematically across locales and languages, exposing blind spots that globally averaged metrics conceal. Ultimately, Pluralis is not presented as a solved evaluation framework for cultural alignment, but rather as a first step and catalyst for future innovation. We call upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.
[16] Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs cs.CL | cs.AIPDF
Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico
TL;DR: 本文首次对22种语言(涵盖高、中、低资源语言)进行了不确定性估计(UE)方法的大规模评估,使用两个人工整理的问答数据集,比较了九种开箱和闭箱UE方法在不同模型规模和架构下的表现,并避免了LLM-as-a-judge和基于嵌入的评分以减少评估噪声。研究发现:1)用英语进行推理提示能显著提升低资源语言的UE性能;2)英语推理能缩小低资源与高资源语言间的UE性能差距;3)UE方法的选择应取决于模型规模,小模型适合基于概率的开箱方法,大模型则更适合闭箱的自我表达不确定性方法。
Details
Motivation: 现有不确定性估计研究主要集中于英语,本文旨在填补多语言和跨语言环境下不确定性估计的评估空白,探究LLM在不同语言资源条件下的可靠性瓶颈。
Result: 在22种语言的大规模评估中,论文发现使用英语进行推理提示能显著改善低资源语言的不确定性估计性能,并缩小与高资源语言的差距;同时确定了模型规模与UE方法选择的关联性。
Insight: 创新点在于首次大规模跨语言评估不确定性估计方法,揭示了低资源语言中理解与生成的可靠性差异,并提出了根据模型规模选择UE方法的实用指南,为多语言LLM系统的校准提供了重要见解。
Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.
[17] From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b cs.CL | cs.AIPDF
Taeyun Roh, Eunha Lee, Wonjune Jang, Sohyun Chung, Junha Jung
TL;DR: 本研究针对BioASQ 14b Task B生物医学问答任务,提出了一种基于问题类型的大语言模型框架。该框架根据是/否、事实型和列表型问题的不同推理与评估需求,分别设计不同的推理流程,以提高答案的鲁棒性和证据基础。
Details
Motivation: 生物医学问答不仅需要从科学文献中准确提取信息,还需要可靠地整合多文档证据。现有方法通常对所有问题类型采用单一提示策略,而不同问题类型在推理和评估上存在差异,因此需要针对性的解决方案。
Result: 在BioASQ 14b Task B的官方评估中,该框架在多个批次中表现出竞争力,并在第4批次的事实型子任务中获得了第一名。
Insight: 创新点在于根据问题类型定制推理流程,例如对是/否问题使用片段重排和自反思以提高稳定性,对事实型问题结合完整片段输入和思维链上下文学习,对列表型问题采用多智能体协作架构进行证据提取、候选生成、验证和聚合。这体现了从简单的投票机制转向更复杂的智能体协作的设计思路。
Abstract: Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task B, designed to improve answer robustness and evidence grounding in biomedical question answering. Rather than applying a single prompting strategy to all questions, the framework selects different inference procedures for yes/no, factoid, and list questions according to their distinct reasoning and evaluation requirements. For yes/no questions, snippet shuffling and self-reflection are used to reduce sensitivity to evidence ordering and improve decision stability. For factoid questions, full-snippet input is combined with chain-of-thought-based in-context learning to support accurate biomedical entity identification. For list questions, a multi-agent architecture is employed, in which evidence extraction, candidate generation, answer verification, and final aggregation are handled collaboratively. Preliminary experiments on BioASQ 13b were used to identify effective inference strategies for each question type, and the resulting framework was subsequently evaluated in the official BioASQ 14b Task B challenge. In the official evaluation, our framework showed competitive performance across multiple batches and achieved first place in the factoid subtask of Batch 4. These results demonstrate the effectiveness of combining question-type-specific inference, ensemble prediction, and agent-based verification for reliable biomedical question answering.
[18] RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation cs.CL | cs.AIPDF
Sambaran Bandyopadhyay, Ananth Muppidi
TL;DR: 本文提出RSF-GLLM框架,用于解决知识图谱多跳问答中的语义鸿沟问题。该框架通过可微的循环软流模块进行图推理,生成离散路径后,再解耦地利用大语言模型进行基于事实的答案生成。
Details
Motivation: 传统检索-阅读流水线在知识图谱多跳问答中不可微,导致检索器难以学习跨越查询与中间节点间缺乏词汇重叠的语义鸿沟。
Result: 在WebQSP和CWQ基准测试上,RSF-GLLM取得了具有竞争力的性能,并且相比基于LLM的昂贵方法具有更优的推理效率。
Insight: 核心创新在于将可微图推理与答案生成解耦,通过循环软流模块实现基于结构线索的语义桥接,并引入流稀疏正则化保证从软概率到离散路径的收敛,从而将推理路径文本化以可靠地微调LLM。
Abstract: Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelines break differentiability, preventing the retriever from learning to bridge the semantic gap where intermediate nodes lack lexical overlap with the query. To address this, we propose RSF-GLLM, a framework decoupling differentiable graph reasoning from answer generation. Our Recurrent Soft-Flow (RSF) module employs a GRU-guided query updater to propagate continuous relevance scores, utilizing a dynamic gating mechanism to traverse semantically dissimilar bridge nodes via structural cues. We introduce flow sparsity regularization to theoretically guarantee convergence from soft probabilities to discrete reasoning paths. These paths are extracted and textualized to fine-tune a Large Language Model (LLM), ensuring generation is grounded in factual topology. Experiments on WebQSP and CWQ demonstrate that RSF-GLLM achieves competitive performance with superior inference efficiency compared to LLM based computationally expensive approaches.
cs.CV [Back]
[19] CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration cs.CV | cs.AIPDF
Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu, Xiao Yao
TL;DR: 本文提出了CanvasAgent,一个通过视觉工具编排实现复杂图像创建和编辑的多模态代理,以及一个大规模多模态工具使用数据集CanvasCraft。该代理通过多轮交互学习协调异构视觉工具,并采用监督微调和基于策略梯度的强化学习进行优化,以处理涉及合成、定位、分割、编辑、合成和增强的复杂图像工作流。
Details
Motivation: 解决复杂图像创建和编辑任务需要超越单一生成或编辑模型,现有多模态工具使用代理主要针对感知、搜索或特定领域编辑优化,缺乏对可执行图像创建轨迹的大规模监督。
Result: 实验评估了最终图像质量和轨迹行为,证明了CanvasAgent和所提出数据集在复杂多工具图像创建工作流中的有效性。
Insight: 创新点包括引入大规模可执行轨迹数据集CanvasCraft,以及代理通过多轮交互主动检查中间结果、跟踪视觉资产并适应视觉状态变化的工具编排能力;从客观角度看,其结合结果级和过程级信号的混合奖励强化学习优化方法对复杂视觉操作任务具有借鉴意义。
Abstract: Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and \textbf{CanvasAgent}, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.
[20] A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog cs.CV | cs.LG | eess.IVPDF
Amir Pouladi, Vesal Ahsani, Haijun Li, Homayoun Najjaran, Afzal Suleman
TL;DR: 本文提出了一种任务驱动的评估框架,用于评估合成雾霾对无人机(UAV)检测与跟踪任务的影响。该框架将深度感知的合成雾生成、图像复原、目标检测和跟踪集成在一个统一流程中。研究通过单目深度估计和大气散射模型,从清晰的真实户外图像生成合成雾霾场景,以解决真实雾霾数据收集和标注的困难。
Details
Motivation: 雾霾严重降低了在天空主导的远距离图像中小型无人机的可见度,从而降低了后续检测和跟踪任务的可靠性。由于收集和标注真实雾霾无人机场景的实际困难,需要一种方法来系统评估雾霾及其复原对下游感知任务的影响。
Result: 评估表明,雾霾显著降低了检测和跟踪性能,主要表现为漏检率增加。在训练数据中包含雾霾图像能最稳定地提升检测鲁棒性,而仅在清晰图像上训练的检测器从测试时的图像复原中获益最大。研究超越了图像级复原指标,评估了雾霾和复原如何影响检测鲁棒性和跟踪性能。
Insight: 创新点在于提出了一个将合成雾生成与下游感知任务(检测、跟踪)评估相连接的任务驱动框架。关键发现是,图像复原质量的提升并不一定转化为下游感知任务的等比例增益,因此复原方法应与检测和跟踪性能联合评估,而非孤立看待。这强调了任务导向评估的重要性。
Abstract: Fog severely degrades the visibility of small unmanned aerial vehicles (UAVs) in skydominant, long-range imagery, reducing the reliability of downstream detection and tracking. This paper presents a task-driven evaluation framework that links depth-aware synthetic fog generation, image restoration, object detection, and tracking within a unified pipeline. Given the practical difficulty of collecting and annotating foggy UAV scenes, synthetic fog is generated from real clear-weather outdoor images containing UAV targets using monocular depth estimation and the atmospheric scattering model. Representative restoration methods from classical, convolutional neural network (CNN)-based, and transformer-based families are first compared, after which the selected restoration model is integrated into the downstream perception pipeline. Detection is evaluated under both clean-only and fog-inclusive training regimes using multiple detector variants, while tracking-by-detection is assessed on clean, foggy, and restored video sequences. Beyond image-level restoration metrics, the study evaluates how fog and restoration affect detection robustness and tracking performance. The results show that fog substantially degrades both detection and tracking, primarily through increased missed detections. Fog-inclusive training provides the most consistent improvement in robustness, whereas test-time restoration is most beneficial when the detector has been trained only on clean imagery. These findings show that restoration quality does not necessarily translate into proportional gains in downstream perception and therefore should be evaluated jointly with detection and tracking performance.
[21] Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM cs.CVPDF
Amol Harsh, Zongyan Han, Jean Lahoud, Ye Liu, Rao Muhammad Anwer
TL;DR: 本文提出了Ground3D-LMM模型,这是一个统一的多模态模型,能够处理点云和可选RGB图像输入,支持包含精确3D区域指向(点级定位)和真实世界度量单位数值输出的3D空间对话。为评估此能力,作者定义了‘3D Grounded Measurement’任务,并基于ScanNet和ScanNet++构建了一个大规模数据集,包含约250万个问答对。实验表明该模型为基于定位和度量的3D对话理解提供了一个强有力的基线。
Details
Motivation: 现有3D大模型存在局限:对话系统通常缺乏明确的3D区域指向,而3D定位模型又不支持交互式、度量感知的对话。本文旨在解决3D环境中自然语言查询的可验证性和度量性需求,即需要模型既能明确指向3D区域,又能输出物理度量值。
Result: 在基于ScanNet和ScanNet++构建的大规模数据集以及多个任务上的广泛实验表明,Ground3D-LMM模型为基于定位和度量的3D对话理解提供了一个强有力的基线。
Insight: 主要创新点在于将精确的3D点级定位与度量感知的对话能力统一在一个模型中,并定义了‘3D Grounded Measurement’这一新任务来评估该能力。从客观角度看,其构建的大规模、细粒度标注数据集(包含物体和部件级别)对于推动3D场景理解与对话研究具有重要价值。
Abstract: Natural-language queries about 3D environments become actionable when responses are verifiable and metric. Verifiability requires explicit grounding to the referred 3D region, while metric answers report physical measurements in real-world units (e.g., size, thickness, clearance, and distance). Existing 3D large multimodal models (LMMs) approaches remain limited: conversational systems typically respond without explicit 3D grounding, while 3D grounding models are not designed for interactive, metric-aware dialogue. In this paper, we present Ground3D-LMM, a unified model that takes a point cloud and an optional RGB image as input and supports 3D spatial conversation with (i) point-grounded responses and (ii) metric numeric outputs at both object and part granularity, including multi-object queries. To evaluate this intersection of grounding and measurement, we define the 3D Grounded Measurement task, which requires predicting the referred 3D region and the corresponding metric quantities in real-world units. We introduce a large-scale dataset built on ScanNet and ScanNet++ datasets with dense object and part annotations and roughly 2.5M question-answer pairs spanning eight tasks, along with a manually verified test set. Extensive experiments on multiple datasets and tasks show that our proposed Ground3D-LMM model provides a strong baseline for grounded, metric-aware 3D conversational understanding. Our dataset and model are publicly available.
[22] Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cs.CVPDF
Chang Nie, Jiaju Wei, Junlan Feng, Chaoyou Fu, Caifeng Shan
TL;DR: 本文提出了Light-Omni,一个用于反射式、轻量级视频理解的多模态智能体框架。它通过双上下文状态(全局状态和参数化潜在状态)在单次前向传播中即时构建所需上下文,从而避免了传统智能体依赖的昂贵迭代推理过程,实现了语义对齐的检索和快速响应。
Details
Motivation: 解决现有视频理解智能体在处理长序列多模态流时,因依赖’侦探式’迭代推理进行动作控制和证据聚合而导致的计算成本高、延迟大的问题。作者认为这种繁重的推理主要是为了弥补检索中全局上下文缺失和语义错位。
Result: 在多个视频基准测试上的广泛实验验证了其有效性。具体而言,它在平均准确率上超过M3-Agent 2.4%,速度提升12.1倍,GPU内存效率提升2.6倍。此外,它还能作为内存系统提升现有MLLM的性能和效率。
Insight: 核心创新在于引入了耦合设计的双上下文状态机制:一个有限大小的、从情景记忆中持续整合的全局多模态脚本作为全局上下文;一个基于此全局上下文生成的参数化潜在状态,直接驱动自主动作并产生检索嵌入。这种设计实现了语义对齐的检索和反射式响应,避免了迭代推理。
Abstract: Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. However, advanced video agents often rely on ``detective-style’’ iterative reasoning for action control (e.g., $\mathtt{search}$) and evidence aggregation, incurring prohibitive costs and latency. We argue that such heavy reasoning primarily compensates for the lack of global context and semantic misalignment in retrieval. This paper introduces Light-Omni, a multimodal agent framework for reflexive and lightweight video understanding. It achieves this through dual contextual states that instantly build the required context in a single forward pass. First, we maintain a global state, a finite-sized multimodal script continuously consolidated from episodic memory, serving as the global context for Light-Omni. Through hierarchical merging, it preserves recent details while summarizing past events. Second, conditioned on this global context, we generate a parametric latent state that directly drives autonomous actions and produces retrieval embeddings, with minimal latency. Benefiting from this coupled design, Light-Omni achieves semantically aligned retrieval and reflexive responses while avoiding iterative reasoning. Extensive experiments validate the effectiveness of Light-Omni across multiple video benchmarks. Notably, it outperforms M3-Agent with an average 2.4% accuracy gain, a 12.1$\times$ speedup, and a 2.6$\times$ improvement in GPU memory efficiency. Furthermore, it serves as a memory system to enhance both the performance and efficiency of existing MLLMs. Project page: https://clare-nie.github.io/Light-Omni.
[23] Statistical Adversaries: Natural Backdoor-like Features in Vision Datasets cs.CV | cs.AI | cs.CR | cs.LGPDF
Paul K. Mandal, Pavan Reddy, Tristan Malatynski
TL;DR: 本文研究了视觉数据集中自然存在的统计信号,这些信号在没有恶意插入的情况下表现出类似后门触发器的行为,作者称之为‘统计对抗者’。通过对ImageNet数据集的分析,识别出与特定标签强相关的模式,并使用统计控制方法去除随机相关性,最终证明这些信号能直接且可预测地改变模型预测。
Details
Motivation: 动机在于探索模型特定对抗攻击之外的另一种失效模式:视觉数据中自然出现的统计信号如何在没有恶意投毒的情况下成为类似后门的漏洞,以揭示数据集结构本身可能带来的安全风险。
Result: 实验表明,这些统计对抗信号比通用损坏更具针对性,且能跨不同模型架构迁移,说明漏洞源于数据集结构和分布而非单一模型的特性。
Insight: 创新点在于提出了‘统计对抗者’概念,揭示了普通数据集即使未被投毒也可能包含可利用的对抗表面,建议数据集审计应将虚假结构视为潜在的视觉模型攻击面,而不仅仅是偏差或可解释性失效的来源。
Abstract: Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave like backdoor-like triggers without being maliciously inserted. We call these signals statistical adversaries. We analyse Imagenet to find patterns that are strongly linked to certain labels. We then use statistical controls to remove random correlations from our candidate signals. Finally, we demonstrate that these signals directly and predictably alter model predictions. These statistical adversaries are more targeted than generic corruptions and transfer across different model architectures. This suggests that some vulnerabilities are driven by dataset structure and distribution rather than a single model’s idiosyncrasies. We conclude that ordinary datasets can contain exploitable adversarial surfaces even in the absence of poisoning, and suggest that dataset audits should treat spurious structure not only as a source of bias or interpretability failure, but also as a latent attack surface for vision models.
[24] Harnessing Generative Image Models for Training-Free Primitive Shape Abstraction cs.CV | cs.AIPDF
Gregor Kobsik, Tim Elsner, Leif Kobbelt
TL;DR: 本文提出了一种无需训练的方法,利用预训练的生成图像模型(如视觉语言模型)从多视角渲染图像中提取3D物体的语义部件,通过颜色编码分割掩码重投影到几何体上,并拟合超二次曲面基元来表示每个部件。该方法无需学习参数,具有类别无关性和方向不变性,在HumanPrim和Toys4K数据集上取得了最低的倒角距离。
Details
Motivation: 解决传统基于学习的3D形状基元抽象方法需要任务特定训练、受限于类别和方向的问题,探索能否直接利用大规模预训练的生成图像模型的视觉理解能力,无需微调即可实现通用且鲁棒的形状抽象。
Result: 在HumanPrim和Toys4K数据集上,该方法使用平均5-9个基元,在所有评估方法中取得了最低的倒角距离,表明其性能优于现有方法。
Insight: 创新点在于完全无需训练的流水线设计,直接利用预训练生成模型的跨类别部件识别能力,实现了类别无关和方向不变的形状抽象;通过真实分割研究指出当前性能瓶颈在于部件分割而非基元拟合,未来生成模型的改进将直接提升方法精度。
Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding. Generative image models trained at scale have recently emerged as generalist visual learners that can identify and segment object parts directly in the image domain, across arbitrary categories and without task-specific training. Adapting such models to downstream tasks typically requires fine-tuning; we ask whether their pretrained capability can instead be harnessed directly, without any training, and answer affirmatively with a training-free harness. Our pipeline renders multi-view images of a 3D object, uses a vision-language model to analyze its semantic parts, prompts a generative image model to paint a color-coded part segmentation mask, reprojects it onto the geometry, and fits a superquadric primitive to each part via parameter optimization. The approach contains no learned parameters: it is category-agnostic and orientation-invariant, properties that previous learning-based models struggled with. Its accuracy ceiling rises with future generative-model improvements, which we confirm with a ground-truth segmentation study showing that part segmentation, not primitive fitting, is the current accuracy bottleneck. On HumanPrim and Toys4K, our method achieves the lowest Chamfer distance among all evaluated methods, using 5–9 primitives per object on average.
[25] Patch Knowledge Transfer for Efficient AI-Generated Image Quality Assessment cs.CVPDF
Jiquan Yuan
TL;DR: 本文提出了一种名为Patch Knowledge Transfer (PKT)的知识蒸馏优化框架,用于高效评估AI生成图像的质量。该框架通过创新的多级知识转移机制,在教师模型(采用局部-全局混合处理)和学生模型(仅依赖全局处理)之间进行协同优化,旨在解决现有方法在计算成本与评估精度之间的失衡问题。
Details
Motivation: 随着图像生成技术的快速发展,对AI生成图像进行感知质量评估变得至关重要。当前主流方法存在两大局限:采用复杂特征提取的策略计算成本过高,而基于简单图像缩放的方案则评估精度显著不足。本文旨在解决这一关键问题,实现高效的大规模生成图像质量评估。
Result: 在4个AIGIQA数据库上的大量实验表明,PKT框架使学生模型在保持与教师模型相当性能的同时,将计算成本降低了67.7%。与现有方法相比,该方法在模型效率和评估精度之间实现了更优的平衡。
Insight: 论文的创新点在于提出了一个基于知识蒸馏的多级知识转移机制,通过双模型架构(局部-全局混合处理的教师模型与仅全局处理的学生模型)实现表示能力与推理效率的协同优化。从客观角度看,这种设计巧妙地平衡了性能与效率,为高效AIGIQA任务提供了一种可借鉴的轻量化解决方案。
Abstract: With the rapid advancement of image generation technologies, perceptual quality assessment of AI-generated images has emerged as a crucial research direction in computer vision. The core challenge of this task lies in achieving efficient quality assessment for massive generated images. Current mainstream approaches exhibit two key limitations: 1) Methods employing complex feature extraction strategies, while improving performance, incur prohibitive computational costs that hinder real-time inference; 2) Simple image scaling-based solutions, despite their computational efficiency, demonstrate significantly inferior assessment accuracy. To address this critical issue, we propose Patch Knowledge Transfer (PKT), a knowledge distillation-based optimization framework that achieves synergistic optimization of visual representation capability and inference efficiency through an innovative multi-level knowledge transfer mechanism. Specifically, we design a dual-model architecture: a teacher model with local-global hybrid processing provides high-quality supervision signals, while a student model relying solely on global processing efficiently inherits the teacher’s representation capacity through multi-level supervision. Extensive experiments conducted on 4 AIGIQA databases demonstrate that the PKT framework enables the student model to maintain performance comparable to the teacher while reducing computational costs by 67.7%. Furthermore, compared to existing methods, our approach achieves a superior balance between model efficiency and assessment accuracy.
[26] VEIL: How Visual Encoding Hijacking Induces Bias In Vision Models cs.CVPDF
Suranjana Sooraj, Xuyang Chen, Madhumitha Venkatesan, Dongyu Liu
TL;DR: 本文提出VEIL框架,系统研究图表编码如何影响CNN在时间序列分类中的学习表示,揭示了模型可能依赖编码引入的视觉线索而非时间模式。通过相似性、可迁移性和归因分析,发现注意力引导训练仅在编码敏感性一致时有效缓解此效应。
Details
Motivation: 解决时间序列分类中CNN模型是否真正学习时间模式,还是依赖图表设计引入的编码特定视觉线索的问题,探讨可视化设计对机器学习表示的影响。
Result: 通过系统诊断分析表明,图表编码选择会显著塑造学习表示,注意力训练仅在编码敏感性被一致识别时有效,否则可能带来有限或负面效果。
Insight: 创新点在于将图形感知从人类读者扩展到视觉模型,提出基于图表的时间序列分类应视为表示和测量问题而非简单建模决策,为可视化设计在机器学习中的影响提供了新视角。
Abstract: Rendering time series as chart images for CNN-based classification has become increasingly common in time-series classification (TSC). However, it remains unclear whether models learn underlying temporal patterns or rely on encoding-specific visual cues introduced by chart design. We present VEIL: a systematic study examining how chart encodings influence learned representations through complementary analyses of similarity, transferability, and attribution. Attention-guided training appears to mitigate this effect when encoding sensitivity is consistently identified across diagnostics, but provides limited or negative benefit when such signals are absent. These findings position VEIL within the broader question of how machines perceive visualizations – extending graphical perception from human readers to vision models – and show that visualization design choices shape learned representations in ways that warrant treating chart-based TSC as a representation and measurement problem rather than a simple modeling decision.
[27] Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring cs.CVPDF
Aditi Naiknaware, Jian Sun, Aminreza Khandan, Shengyang Huang, Sean Dow
TL;DR: 本文提出了一种用于临床伤口监测的多模态框架,通过结合伤口图像和临床文本信息,实现个性化的严重不良事件(SAE)检测。该方法基于冻结的BiomedCLIP主干网络,采用双流低秩适应(LoRA)框架编码临床记录和伤口描述,并引入跨上下文LoRA融合机制进行信息交换。此外,还设计了一个结合语义匹配、视觉典型性、文本对齐和视觉对齐的统一OOD检测框架,并纳入协变量一致性和时间漂移惩罚来捕捉伤口愈合的动态变化。
Details
Motivation: 伤口监测是一个关键但服务不足的临床挑战,现有视觉语言模型(VLMs)缺乏领域特定的基础,难以整合伤口图像与异质临床信息,且对训练分布外病例的检测机制有限。
Result: 在通过临床访问收集的纵向伤口数据集上的实验表明,该方法在伤口愈合评估和SAE检测方面均表现出有希望的性能。
Insight: 创新点包括:基于冻结BiomedCLIP主干的跨上下文LoRA融合机制,实现无需全模型微调的多模态表征;以及结合多种对齐和一致性度量的统一伤口特异性OOD检测框架,用于个性化SAE识别和动态愈合追踪。
Abstract: Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tissue deterioration, and delayed healing can significantly impact patient outcomes. While vision-language models (VLMs) show strong multimodal reasoning, they often lack domain-specific grounding to integrate wound imagery with heterogeneous clinical information, and provide limited mechanisms for detecting cases that diverge from the training distribution. We present a multimodal framework for automated wound monitoring and SAE detection. Our approach leverages paired clinical notes and wound descriptions capturing visual characteristics such as appearance, surrounding skin condition, color changes, and signs of inflammation or healing progression, encoded through a dual-stream Low-Rank Adaptation (LoRA) framework built on a frozen BiomedCLIP backbone. We introduce a cross-contextual LoRA fusion mechanism enabling information exchange between clinical semantics and visual wound descriptors, producing context-aware multimodal representations without full model fine-tuning. To identify personalized SAEs, we propose a wound-specific out-of-distribution (OOD) detection framework combining semantic matching, visual typicality, caption-text alignment, and caption-visual alignment into a unified SAE (OOD) score. To capture healing dynamics, we incorporate covariate consistency and temporal drift penalties that leverage changes in wound characteristics across visits. Experiments on a longitudinal wound dataset collected through clinical visits show promising performance on both wound healing assessment and SAE detection, highlighting the potential of semantically enriched, temporally aware vision-language systems for clinical wound monitoring and early risk identification.
[28] REVIVE: A Multi-Modal Framework for Vandalism Detection and Recovery in Autonomous Vehicles cs.CV | cs.LGPDF
Abdullah Tariq Choudhry, Tapadhir Das
TL;DR: 本文提出了REVIVE框架,用于检测和修复自动驾驶车辆中因恶意破坏导致的摄像头遮挡攻击。该框架集成了遮挡检测、模式识别、图像分割和类型感知的修复方法,旨在恢复被破坏图像的可用性,以支持下游的物体检测任务。
Details
Motivation: 自动驾驶车辆的摄像头感知系统易受物理遮挡攻击的威胁,现有研究多集中于攻击检测,而攻击发生后如何有效恢复图像以维持感知功能则研究不足。
Result: 在500对干净/被破坏图像对上,未修复图像使YOLOv8l的召回率降至0.588;在参考图像对齐的条件下,直接像素替换方法能将召回率恢复至0.967,F1分数达0.970。Stable Diffusion修复的图像结构相似性指数在0.667-0.867之间,峰值信噪比在15.4-26.7dB。
Insight: 创新点在于提出了一个集成了检测、识别、分割和类型感知修复的完整流程,并引入了基于质量门的异步修复分支策略,确保修复后的图像流质量不低于未修复的基线,为实时感知系统提供了鲁棒的恢复机制。
Abstract: Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detection frameworks can identify vandalized images, restoring camera-stream utility after physical occlusion remains underexplored. This paper presents present the Recovery and Enhancement of Vandalized Images for Vision Excellence (REVIVE) framework, a vandalism recovery pipeline integrating: (1) binary VOA detection, (2) multi-class VOA pattern identification, (3) EfficientNet-based U-Net segmentation, and (4) type-aware recovery using Bootstrapping Language-Image Pre-training (BLIP)-guided Stable Diffusion inpainting, direct pixel replacement, or adaptive median filtering. Stable Diffusion shows variable reconstruction performance (per-pattern SSIM 0.667-0.867, PSNR 15.4-26.7dB) across VOA patterns, while aligned direct pixel replacement achieves near-identical reconstruction under the aligned-reference condition. On 500 tracked clean/vandalized image pairs, unrecovered VOAs reduce YOLOv8l object-detection recall to 0.588, while direct pixel replacement restores recall to 0.967 and F1-score to 0.970 under that aligned-reference condition. LaMa, Telea, and Navier-Stokes baselines improve image similarity but provide more limited downstream detection recovery, and Stable Diffusion is treated as an asynchronous recovery branch subject to a quality gate rather than a blocking real-time perception step. We evaluate a reference-available quality gate that filters recovered candidates before downstream use: without it, type-aware routing degrades per-image recall to 0.304, whereas with it, recall returns to 0.608, at or above the unrecovered baseline, ensuring the forwarded stream is never worse than the unrecovered frame. REVIVE therefore, provides a structured recovery framework from VOAs in AVs.
[29] Robust Face Super-Resolution and Recognition Through Multi-Feature Aggregation in Diffusion Models cs.CVPDF
Marcelo dos Santos, Rayson Laroca, João Carlos Raposo Neves, David Menotti
TL;DR: 本文提出FASR++,一种基于扩散模型的人脸超分辨率算法,通过聚合多张低质量辅助图像的特征来生成高分辨率人脸图像,旨在减少身份失真并提升人脸识别性能。
Details
Motivation: 解决监控环境下低分辨率、姿态变化、光照不均和遮挡导致的人脸图像质量差、识别算法性能受限的问题,传统超分方法易引入身份失真,需利用多帧信息增强鲁棒性。
Result: 在两个标准人脸识别数据集上验证,在验证、识别任务及PSNR、SSIM、LPIPS等图像质量指标上达到SOTA水平。
Insight: 创新点在于无需显式提供软属性或计算梯度引导,仅通过多张低质量辅助图像的特征聚合来恢复人脸细节,减少身份失真,可作为预处理步骤显著提升识别性能。
Abstract: Images acquired in surveillance environments often suffer from conditions such as low resolution, variations in pose, irregular illumination, and occlusions. Due to the low quality of these images, face recognition algorithms often struggle. This major limitation can be addressed by employing super-resolution techniques that enhance the details of the image. However, due to the high degree of difficulty of the problem, most super-resolution algorithms tend to cause distortions in the image and in the individual’s identity. Thus, additional information must be incorporated into the processing to improve recognition robustness. In this regard, surveillance cameras can capture multiple images, even at low quality, and the data extracted from these images, such as consecutive video frames, can significantly enhance both super-resolution and facial recognition. In this work, we introduce FASR++, a diffusion-model-based super-resolution algorithm. It leverages a reference low-resolution image and features extracted from multiple auxiliary low-quality images to generate a super-resolved output, minimizing distortions in the individual’s identity. Our approach recovers facial features without explicitly providing soft attributes or computing a function gradient to guide the reconstruction process. FASR++ generates high-quality images that can considerably improve performance in face recognition tasks when used as a pre-processing step. We validate our approach on two standard face recognition datasets and attain state-of-the-art results for verification, face recognition, and image quality metrics such as PSNR, SSIM, and LPIPS.
[30] Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models cs.CVPDF
Zhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng, Ke Yan
TL;DR: 本文提出了一种名为Scene Graph Thinking (SaGe)的新范式,旨在通过显式的场景图表示来增强多模态大语言模型(MLLMs)的结构化视觉推理能力。该方法首先利用自动化数据引擎将平面图像-文本语料库转换为结构化的场景图,并基于此构建了12万条高质量的训练数据。随后,通过两阶段图对齐的后训练范式(监督微调和强化微调)来内化模型的结构化推理能力。
Details
Motivation: 现有MLLMs主要关注孤立物体,忽略了结构化关系,这限制了它们在视觉密集型任务(如目标导航)上的性能。本文旨在解决这一挑战,通过引入场景图来促进细粒度和结构化的视觉推理。
Result: 该方法在八个多模态基准测试上取得了显著改进,在细粒度感知和推理任务上表现出强大的有效性。
Insight: 创新点在于提出了一个将平面视觉数据自动转换为结构化场景图的数据引擎,以及一个两阶段的图对齐训练范式,特别是强化微调中提出的节点作为代理的图奖励机制,以巩固高效的图探索能力。
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structured relationships for efficient target navigation, limiting their performance on visually intensive tasks. To address this challenge, we introduce Scene Graph Thinking (SaGe), a novel paradigm that enables fine-grained and structured visual reasoning through explicit scene-graph representations. Specifically, we first introduce an automated data engine that converts flat image-text corpora into structured scene graphs, where hierarchical entities constitute the nodes and diverse visual relations define the edges. Building upon this, we construct 120K high-quality training data by sampling reasoning traces from scene graphs. Then, two-stage graph-aligned post-training paradigms are introduced, where supervised fine-tuning internalizes MLLMs with structured reasoning, and subsequent reinforcement fine-tuning proposes node-as-proxy graph rewards to consolidate efficient graph exploration. With curated data and graph-aligned training, our approach achieves significant improvements across eight multimodal benchmarks, demonstrating strong effectiveness on fine-grained perception and reasoning tasks. Code is available at https://github.com/zwyang6/SaGe.
[31] SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs cs.CVPDF
Hossein Rajoli, Fatemeh Lotfi, Niloufar Alipour Talemi, Hossein Kashiani, Xiaolong Ma
TL;DR: 本文提出了SAMPLe,一种基于锐度感知最小化的优化器,用于解决视觉语言模型中提示学习面临的性能与泛化权衡问题。该方法通过考虑损失景观的锐度来增强提示的泛化能力,并作为插件集成到多种提示学习框架中。
Details
Motivation: 预训练的视觉语言模型在提示学习中存在性能-泛化困境:针对特定任务调整提示虽然能提高在已知分布上的准确率,但会损害对未见数据的泛化能力,导致过拟合。
Result: 实验表明,SAMPLe被集成到CoOp、CoCoOp、MaPLe、TCP和Co-Prompt等多种框架中,在各种设置下均能提升性能并持续优于现有优化器,成为一个鲁棒的、模型无关的解决方案。
Insight: 核心创新在于将锐度感知优化思想引入提示学习,通过在每个优化步骤中满足目标函数约束来平衡探索与利用,动态适应局部曲率和梯度特性,从而减少过拟合并保持预训练模型的泛化潜力。
Abstract: Pre-trained Vision-Language Models (VLMs) like CLIP have proven highly effective as foundation models for various downstream applications. However, prompt learning in VLMs encounters a performance-generalization dilemma: while prompts can be tuned to achieve high accuracy on seen distributions, this tuning process often undermines their generalizability to unseen data. The limited set of learnable prompts, which contextualize and condition the input to steer it toward the task within the pretrained VLM, tends to overfit the training data, leading to a trade-off between task-specific performance and preserving generalization. To address this dilemma, we introduce SAMPLe (Sharpness-Aware Minimization Prompt Learning), a plug-in sharpness-aware optimizer that enhances prompt generalizability by accounting for loss landscape sharpness. Unlike conventional methods, SAMPLe balances exploration and exploitation by satisfying objective function constraints at each step, dynamically adapting to the current optimization state based on the local curvature and gradient properties. This approach reduces overfitting on seen distributions and improves adaptability to unseen data, preserving the generalization potential of pre-trained VLM models. We integrate SAMPLe into multiple prompt learning frameworks, including CoOp, CoCoOp, MaPLe, TCP, and Co-Prompt, demonstrating its effectiveness across diverse methods. Experiments show that SAMPLe elevates prompt learning frameworks and consistently outperforms existing optimizers across diverse settings, establishing itself as a robust, model-agnostic solution for prompt learning.
[32] Optimized Adaptive Loop Filter in Versatile Video Coding cs.CVPDF
Meng Xuewei, Zhang Jiaqi, Jia Chuanmin, Zhang Xinfeng, Wang Shanshe
TL;DR: 本文提出了一种优化的自适应环路滤波器(ALF)框架,用于降低多功能视频编码(VVC)标准中ALF模块的编码复杂度和图像缓冲区访问次数。该框架包括GALF和CCALF的并行设计、GALF的自适应参数决策以及无需执行滤波操作即可估计CCALF滤波失真的一遍式CCALF方案。
Details
Motivation: VVC标准中的自适应环路滤波器(ALF)虽然能有效减少压缩伪影,但其编码复杂度高,且编码器中需要大量图像缓冲区访问,这会增加外部内存访问,对软硬件设计不友好。
Result: 与VTM-8.0相比,所提方法在RA配置下,将图像缓冲区访问次数从152次减少到1次,ALF模块节省约25%的时间,且编码性能变化可忽略不计。部分方法已被VVC参考软件采纳。
Insight: 创新点在于通过并行化设计、自适应参数决策和高效失真估计,在保持编码性能的同时显著降低了ALF的计算和内存访问开销,为VVC编码器的软硬件优化提供了实用方案。
Abstract: In the Versatile Video Coding(VVC) standard, adaptive loop filter(ALF), including Geometry transformation-based Adaptive Loop Filter(GALF) and Cross Component Adaptive Loop Filter(CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software.
[33] Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator cs.CV | cs.ROPDF
Zihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee
TL;DR: Image2Sim是一个实时神经仿真框架,能够从带位姿的RGB-D图像序列构建高质量交互式环境,用于具身导航任务。它通过解耦3D空间锚定与逼真观测合成,使用前馈特征高斯模型进行场景构建,并提出几何感知的一步像素流模型进行渲染。该框架还作为自动化数据引擎,生成了大量交互场景和导航训练样本,显著提升了导航模型的性能。
Details
Motivation: 解决具身导航领域因缺乏可扩展、高保真且物理基础的交互环境而受限的问题,旨在弥合真实扫描数据集规模有限与合成仿真器存在较大仿真-现实差距之间的鸿沟。
Result: 在主要基准测试上,完全在Image2Sim生成的神经环境中训练的导航模型取得了显著提升,并能有效迁移到真实世界的零样本设置中,表明其作为大规模具身导航训练基质的实用性。
Insight: 创新点在于将3D空间锚定与观测合成解耦,并提出了特征高斯表示和几何感知的一步像素流渲染模型;客观来看,该框架实现了从视觉数据到可交互仿真环境的高效、高质量自动化构建,为大规模具身智能训练提供了新的数据生成途径。
Abstract: Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
[34] LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding cs.CV | cs.AIPDF
Guang Yang, Brian Siyuan Zheng, Victoria Ebert, Noah A. Smith
TL;DR: LEGATO 2是一个用于从乐谱图像中提取符号记谱和语义知识的新型流程。它首次引入了大规模神经模型,以逐行(按谱表系统)的顺序处理乐谱,并能生成包含嵌入文本内容(如标题和注释)的符号转录。该流程结合了谱表级分割与自回归视觉语言模型,在多个数据集上超越了先前的最先进方法,并在乐谱理解任务中展现出优势。
Details
Motivation: 解决传统光学音乐识别(OMR)方法将整页乐谱视为无差别图像、难以处理任意长输入以及无法生成包含文本的符号转录的问题。
Result: 在多个数据集上,LEGATO 2始终优于先前的SOTA方法,在OMR和下游乐谱理解任务中均建立了新的最先进性能。
Insight: 创新点在于采用按谱表系统顺序处理的序列化方法,以及结合视觉与语言模型生成包含文本的符号转录,这提高了模型对长输入和复杂乐谱文档的扩展性与理解能力。
Abstract: We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, Legato 2 consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. Legato 2 establishes new state-of-the-art performance in both OMR and downstream sheet music understanding.
[35] DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction cs.CVPDF
Jiakun Li, Zhe Li, Wenqiang Wu, Zheng Chang, Mingqi Gao
TL;DR: 本文提出了DeSeG框架,用于合成物理合理的人-场景交互(HSI)。该框架通过分层结构,将语义意图与几何约束解耦,以解决现有生成模型存在的语义-几何纠缠问题。具体包括一个残差语义规划器和一个物理正则化的扩散执行器,分别负责语义控制和碰撞感知的运动生成。
Details
Motivation: 现有生成模型在合成人-场景交互时存在语义-几何纠缠的归纳偏差,即模型会过度依赖训练数据中空间约束与特定动作的强相关性,从而在面临严格几何线索时压制语义意图,并加剧身体与场景穿透等物理幻觉问题。
Result: 在Lingo数据集上的大量实验表明,DeSeG达到了最先进的性能,与SOTA基线相比,平均场景穿透减少了47%,语义对齐提高了29%。
Insight: 主要创新点在于明确地将语义意图规划与几何约束执行解耦的层次化框架设计,以及将可微分排斥势场直接整合到扩散目标中以实现碰撞感知的物理正则化方法。这为解决生成任务中语义与物理合理性之间的权衡提供了新思路。
Abstract: Synthesizing physically plausible human-scene interactions (HSI) remains a critical challenge in computer vision and the development of human avatars. Although recent generative models enable diverse motion synthesis, they suffer from an inductive bias referred to as semantic-geometric entanglement. Because spatial constraints often strongly correlate with specific actions in training data, monolithic models will learn the shortcut bias, aggressively overriding the semantic intent when faced with strict geometric cues. Furthermore, this entanglement exacerbates physical hallucinations, such as body-scene penetrations. To address these limitations, we propose DeSeG, a hierarchical framework that explicitly decouples semantic intent from geometric constraints. First, we introduce a Residual Semantic Planner that encodes textual instructions and canonicalized goal voxels into a compact latent space, enabling fine-grained semantic control independent of spatial trajectories. Second, we propose a physics regularized diffusion executor that incorporates differentiable repulsive potential fields directly into the diffusion objective, enforcing collision-aware motion generation. Extensive experiments on the Lingo dataset demonstrate that DeSeG achieves state-of-the-art performance, reducing mean scene penetration by 47% and improving semantic alignment by 29% over the SOTA baselines.
[36] Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning cs.CV | cs.AIPDF
Yake Wei, Yuan Wang, Fengyun Rao, Jing Lyu, Di Hu
TL;DR: 本文提出SegAnswer方法,将多模态大语言模型(MLLM)视觉推理中的关注区域从边界框升级为像素级分割掩码,通过分割掩码隔离目标区域以获取更精确的视觉细节,减少背景冗余和干扰。
Details
Motivation: 当前MLLM的视觉推理通常通过边界框放大感兴趣区域,但边界框不够精确且包含冗余背景;本文旨在通过像素级分割提供更精细的视觉输入,以提升推理的准确性和与MLLM视觉令牌结构的对齐。
Result: 在多个基准测试(包括高分辨率感知、通用感知和幻觉检测)上,SegAnswer均实现了性能的稳定提升,并在分割任务上表现出色,验证了其可靠的像素级定位能力。
Insight: 创新点在于将视觉推理的放大单元从边界框转变为分割掩码,这能提供更精确的感兴趣区域,且分割后的离散图像块与MLLM通过位置嵌入构建视觉令牌的方式更匹配,从而提升模型性能。
Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with images’’. A basic operation in this reasoning process is to zoom in on regions of interest (often represented with bounding boxes) to acquire finer visual details. In this paper, we propose \textbf{Seg}mentation before \textbf{Answer}ing (SegAnswer), which shifts the unit of zoom-in from the popular bounding box to pixel-level segmentation mask. By employing fine-grained masks to isolate the target area from cluttered environments, segmented visual input yields a more precise region of interest, effectively filtering out redundant background and interfering objects. Furthermore, the discrete patches of segmented visual input align more seamlessly with how MLLMs structure visual tokens via positional embeddings. In experiments, we evaluate SegAnswer across diverse benchmarks, including high-resolution perception, general perception, and hallucination. It achieves consistent improvements and also exhibits considerable performance on segmentation tasks, validating its capability for reliable pixel grounding.
[37] Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective cs.CVPDF
Tianyuan Zhang, Xianglong Liu, Aishan Liu, Lu Wang, Yitong Zhang
TL;DR: 该论文针对自动驾驶系统中环境幻觉(如阴影、反射和轮胎痕迹)对车道感知的干扰问题,首次提出了一个名为LanEvil++的基准测试。该基准包含14种幻觉类型,利用CARLA模拟器生成了大量高保真、可控的3D场景数据。评估表明,环境幻觉显著降低了现有最先进车道检测模型和视觉语言模型的性能,并可能导致错误的驾驶决策。论文还提出了多模态幻觉防御方法MIDA,有效提升了模型的鲁棒性。
Details
Motivation: 现实驾驶环境中普遍存在的环境幻觉(如阴影、反射)可能干扰视觉感知,导致场景误判,对自动驾驶系统构成严重安全风险,但现有研究很大程度上忽视了这一问题。
Result: 在LanEvil++基准上的广泛评估显示,环境幻觉使最先进的车道检测模型的准确率平均下降5.27%,F1分数下降10.49%;基于视觉语言模型的系统在GPT分数和语言分数上分别下降2.03%和0.75%。其中,阴影是最具破坏性的因素,可使准确率降低高达7.20%。闭环模拟和真实世界案例研究证实了这些幻觉会导致错误的驾驶决策和安全关键故障。
Insight: 论文的主要创新点在于首次系统地研究和基准测试了环境幻觉对自动驾驶车道感知的鲁棒性影响,并构建了大规模、高保真的数据集LanEvil++。从客观角度看,其提出的多模态幻觉防御方法MIDA为提升模型在复杂环境下的鲁棒性提供了有效的解决方案,强调了在自动驾驶安全评估中考虑此类自然但被忽视现象的重要性。
Abstract: Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perception, leading to misinterpretation of the scene and posing serious safety risks to autonomous driving (AD) systems. However, existing researches largely overlook these phenomena, leaving a critical gap. To address this issue, we study AD robustness through the lane perception perspective, a fundamental task supporting core functions like cruise control and lane centering. We focus on two representative models: conventional lane detection (LD) and vision-language model-based systems (ADVLMs). In this work, we introduce the first benchmark, LanEvil++, for evaluating the robustness of lane perception under environmental illusions. LanEvil++ encompasses 14 types of illusions and leverages the CARLA simulator to generate 94 high-fidelity, fully controllable 3D scenes, yielding a dataset of 90,292 annotated images, 1,596 video clips, and 41,855 visual question answering pairs. Extensive evaluations demonstrate that environmental illusions substantially degrade the performance of state-of-the-art LD methods. On average, LD models experience a 5.27% drop in Accuracy and a 10.49% decline in F1-score, while ADVLMs show a 2.03% reduction in GPT-score and a 0.75% drop in Language-score. Among all illusions, shadows emerge as the most disruptive factor, reducing accuracy by up to 7.20%. Furthermore, closed-loop simulations reveal that these illusions can lead to incorrect driving decisions. Complementary real-world case studies highlight safety-critical failures in actual traffic scenes. To enhance robustness, we propose the Multimodal Illusion Defense Approach (MIDA). MIDA achieves substantial gains under challenging conditions, boosting robustness by 4.23% on LD models and 3.82% on ADVLMs.
[38] TRIG: Trajectory-Rig Decoupled Metric Geometry Learning cs.CV | cs.ROPDF
Lizhou Liao, Wentao Xu, Handong Wang, Lirong Yang, Shuai Yang
TL;DR: 本文提出TRIG框架,用于自动驾驶中的视觉几何感知。该框架将相机位姿分解为自车轨迹和相机刚体两个组件,分别建模动态的自车运动和静态的多相机拓扑结构,并通过解耦的位姿编码和监督实现度量一致的学习。
Details
Motivation: 现有视觉几何模型在姿态估计、深度预测和3D重建方面表现良好,但未针对刚性的多相机驾驶系统进行优化,其将时变的自车运动和静态的相机刚体几何纠缠在一起建模,限制了车辆侧几何先验的利用。
Result: 在五个自动驾驶基准测试上的实验表明,TRIG在姿态估计、度量深度预测和3D重建任务上达到了最先进的性能。
Insight: 核心创新点在于将相机位姿解耦为轨迹和刚体两部分进行独立建模与监督,并设计了稀疏的时空注意力机制来分离跨相机交互与时间聚合,在降低全局注意力计算成本的同时保持了几何推理能力。
Abstract: Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual geometry models show strong performance in pose estimation, depth prediction, and 3D reconstruction, but are not tailored to rigid multi-camera driving systems. They often encode camera poses as entangled representations, in which time-varying ego-motion and static camera-rig geometry are jointly modeled, limiting the utilization of vehicle-side geometric priors. We propose Trajectory-Rig Decoupled Metric Geometry Learning (TRIG), a geometry perception framework for autonomous driving. TRIG factorizes camera poses into ego-trajectory and camera-rig components, enabling separate modeling of ego-motion and static multi-camera topology. We introduce decoupled pose encoding and supervision, which separately constrain trajectory evolution and rig geometry for metric-consistent learning. Moreover, sparse Temporal–Spatial attention separates cross-camera interaction from temporal aggregation, reducing global attention cost while preserving geometric reasoning. Experiments on five autonomous driving benchmarks show that TRIG achieves state-of-the-art performance in pose estimation, metric depth prediction, and 3D reconstruction.
[39] Complementary Roles of Image Classification and Vessel Segmentation in AI-Based Screening for Retinopathy of Prematurity Plus Disease in a Kenyan Preterm Cohort cs.CV | cs.AIPDF
Fred Mutisya, Oscar Onyango, Sarah Sitati, Syokau Ilovi, Aeesha NJ Malik
TL;DR: 该研究针对肯尼亚早产儿视网膜病变(ROP)的Plus疾病筛查,评估了图像分类和血管分割两种AI方法的互补作用。通过分析121名婴儿的1,635张眼底图像,研究发现结合分类和分割的混合模型(如概率集成)在敏感性和特异性上取得了最佳平衡性能。
Details
Motivation: ROP是儿童失明的可预防原因,在中低收入国家负担日益加重,但缺乏训练有素的眼科医生进行筛查。Plus疾病的诊断主观性强且存在差异,因此需要自动化筛查来扩展专家覆盖范围,而非洲地区的相关证据有限。
Result: 在保留测试集上,血管分割模型达到Dice系数0.533、IoU 0.368、敏感性0.623和特异性0.979。RGB分类器敏感性高但过度转诊,而结合分割的模型特异性更高。概率集成方法取得了最佳平衡性能:敏感性0.692、特异性0.914、平衡准确率0.803,优于单独使用视觉分类器。
Insight: 创新点在于揭示了图像分类和血管分割在ROP Plus检测中的互补性:分类器支持高敏感性的病例发现,而分割提高特异性并减少过度转诊。研究建议非洲ROP AI系统应采用结合工作流,并进行前瞻性多中心验证,这为资源有限地区的自动化筛查提供了实用框架。
Abstract: Background. Retinopathy of prematurity (ROP) is a preventable cause of childhood blindness, with rising burden in low- and middle-income countries where ROP-trained ophthalmologists are scarce. Plus disease, marked by retinal vessel dilation and tortuosity, triggers treatment but is subjective and variable. Automated screening could extend specialist reach, but African evidence remains limited. Methods. We analysed 121 Kenyan preterm infants, covering 237 eyes and 1,635 fundus images graded as No Plus, Pre-Plus or Plus. Vessel annotations from two graders supported segmentation training. Eleven configurations were evaluated for eye-level Plus detection using patient-grouped nested cross-validation, including image classifiers, multiple-instance learning, multi-task segmentation-classification, and segment-then-classify pipelines. Results. Vessel segmentation was feasible, achieving pooled Dice 0.533, IoU 0.368, sensitivity 0.623 and specificity 0.979 on held-out images. RGB classifiers were highly sensitive but over-referred, while segmentation-coupled models were more specific. Combining approaches improved performance: an OR-based screen achieved the highest sensitivity, an AND-based confirmation achieved the highest specificity, and a probability ensemble gave the best balanced performance, with sensitivity 0.692, specificity 0.914 and balanced accuracy 0.803, outperforming the vision classifier alone. Conclusions. Classification and vessel segmentation are complementary for ROP Plus detection in Kenyan data. Classifiers support sensitive case-finding, while segmentation improves specificity and reduces over-referral. African ROP AI systems should use combined workflows and undergo prospective multi-site validation.
[40] AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring cs.CVPDF
Younggun Kim, Taeheon Kim, Youngseo Kim, Seunghee Park
TL;DR: AVA-VLM是一种自适应视觉注意力-视觉语言模型,用于野外施工现场监控。它模仿人类从粗到细的推理策略,先处理低分辨率全局图像,仅在需要详细检查时才请求高分辨率局部裁剪,从而提升远距离和低分辨率条件下的可靠性,并显著减少视觉令牌使用。
Details
Motivation: 现有针对施工现场定制的视觉语言模型主要通过单张全局图像的问答式微调进行适配,但该直接范式在部署范围、低分辨率输入下的可靠性以及推理效率方面存在局限。
Result: 实验表明,AVA-VLM在长距离和降低分辨率条件下提高了可靠性,同时大幅减少了视觉令牌的使用量。
Insight: 创新点在于引入了人类启发的从粗到细推理策略,并构建了一个区域感知的思维链数据集,教导模型何时检查、裁剪何处以及如何使用局部证据,实现了自适应视觉注意力机制。
Abstract: Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained VLMs through direct QA-style fine-tuning from a single global image. We argue that this direct paradigm remains limited for in-the-wild deployment in terms of operational range, reliability under reduced-resolution inputs, and inference efficiency. To address these challenges, we propose AVA-VLM, an Adaptive Visual Attention-Vision Language Model that follows a human-inspired coarse-to-fine reasoning strategy. AVA-VLM first reasons over a low-resolution global image and selectively requests a high-resolution local crop only when detailed inspection is needed, similar to how a human inspector zooms in on hard-to-see yet important areas. We further introduce a region-aware Chain-of-Thought dataset that teaches the model when to inspect, where to crop, and how to use local evidence. Experiments show that AVA-VLM improves reliability under long-distance and reduced-resolution conditions while substantially reducing visual-token usage.
[41] Realistic Compound-Lens Defocus Blur Synthesis cs.CVPDF
Yunkyu Lee, Woohyeok Kim, Sunghyun Cho
TL;DR: 本文提出了一种合成真实散焦模糊数据集的流程,用于生成具有多样复合镜头特性的图像对。该流程结合了基于Debye CZT传播的高效波光学PSF计算、处理遮挡的深度感知散焦渲染,以及在线性辐射空间中进行模糊合成与相机ISP模拟。通过此流程生成了大规模合成数据集CLDefocus,实验表明基于该数据集训练的模型在跨设备泛化能力上优于现有数据集。
Details
Motivation: 现有散焦去模糊数据集的镜头光学多样性和真实性有限,导致深度学习模型在不同相机和镜头上的性能下降,需要更真实且多样化的合成数据来提升模型的泛化能力。
Result: 在跨设备泛化实验中,使用CLDefocus数据集训练的模型相比基于现有真实和合成数据集训练的模型表现更优,显示出更好的泛化性能。
Insight: 创新点包括:统一流程整合波光学PSF计算、深度感知渲染和相机ISP模拟,以生成逼真且镜头多样的散焦数据;同时指出真实数据集的缺陷可能误导全参考评估,强调了合成数据在克服真实数据偏差方面的价值。
Abstract: Defocus blur degrades fine image structures and limits visual perception, which can adversely affect downstream vision tasks. Although recent deep learning deblurring methods have achieved strong performance, their effectiveness depends on training data and often degrades across cameras and lenses due to limited optical diversity and realism in existing datasets. In this paper, we propose a pipeline for synthesizing realistic defocus deblurring datasets for diverse compound lenses. It integrates efficient wave-optics PSF computation via Debye CZT propagation, depth-aware defocus rendering with occlusion handling, and blur synthesis in the radiometrically linear space with camera ISP simulation. This unified pipeline enables the scalable generation of photorealistic defocus datasets with diverse lens characteristics. Using our pipeline, we generate CLDefocus, a large-scale synthetic dataset containing lens-diverse defocus image pairs. We further analyze the limitations of real-captured defocus datasets and show that such imperfections can bias full-reference evaluation. Extensive experiments demonstrate that models trained on CLDefocus achieve improved cross-device generalization compared to models trained on existing real and synthetic datasets.
[42] Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context cs.CV | cs.AIPDF
Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali
TL;DR: Harrison.Rad 1.5 (HR1.5) 是一个专门用于放射学的多模态大语言模型,能够根据图像、先验研究和临床上下文生成结构化和非结构化的放射学报告。该模型通过三阶段流程训练,并在多个评估框架和数据集上进行了测试,是唯一达到模拟FRCR考试及格标准的系统,在多项临床任务中取得了最高准确率。
Details
Motivation: 放射成像需求增长快于放射科医生队伍的扩张,导致报告积压。论文旨在通过开发一个能够辅助生成放射学报告的AI模型,直接减少放射科医生在解读图像、整合临床历史和先验研究、起草报告上所花费的时间和精力。
Result: HR1.5在RadBench(模拟FRCR 2B短病例考试)、ReXGradient和内部多模态数据集上进行了评估。它是唯一达到模拟FRCR及格标准的系统,在封闭式临床问题、多身体部位和乳腺X光报告以及公共胸部报告的主要临床对齐评分中均取得了最高准确率。
Insight: 创新点包括:1) 专门针对放射学领域设计的三阶段训练流程(领域适应、对比视觉编码器训练、视觉问答微调);2) 提出了扩展的Findings-Diagnosis评分框架,结合了基于本体的同义词匹配和极性矛盾检测;3) 引入了包括问题敏感Grad-CAM热图、注意力分析和置信度估计在内的可解释性分析,为未来临床应用的负责任评估提供了支持。
Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone. The most direct opportunity is reducing the time and effort radiologists spend producing reports, a task that requires interpreting images, integrating clinical history and prior studies, and drafting structured findings. We present Harrison.Rad 1.5 (HR1.5), a radiology-specific multimodal large language model that accepts interleaved text and visual inputs and generates structured and unstructured text across plain-film radiology, spanning computed radiography, chest, musculoskeletal, abdominal, spine, and pelvic x-rays, and mammography. HR1.5 is trained through a three-stage pipeline: domain adaptation of a base language model on radiology reports, contrastive vision-encoder training with curriculum-based hard negatives on ~6 million image-report instances, and visual-question-answering fine-tuning on multi-turn conversations. We evaluate it with a Findings-Diagnosis scoring framework that extends RadGraph-XL entity extraction with ontology-based synonym matching and polarity-contradiction detection, benchmarked on RadBench, a simulated FRCR 2B Short Case examination scored against Angoff-method thresholds, ReXGradient, and internal multi-modality datasets. HR1.5 is the only system evaluated to meet the simulated FRCR passing standard and achieves the highest accuracy on closed-format clinical questions, across anatomical regions, on internal multi-body-part and mammography reporting, and on the primary clinically-aligned score for public chest reporting. We further examine explainability and model behaviour, including question-sensitive Grad-CAM heatmaps, attention analysis, and confidence estimation, to support responsible future evaluation toward clinical use, and a framework for clinically grounded assessment of report quality.
[43] GaussFusion: Towards Multimodal 3D Gaussian Pretraining cs.CVPDF
Zhixuan You, Jihua Zhu, Yiding Sun, Zihao Guo, Haozhe Cheng
TL;DR: 本文提出了GaussFusion,一个用于3D高斯表示的多模态预训练框架。该框架通过跨模态语义对齐,将图像和文本监督集成到掩码高斯建模中,使高斯编码器能够在预训练期间学习视觉和语言级别的语义信息。为了适应高斯基元的非均匀分布,还提出了高斯显著性引导的多尺度空洞掩码(GSHM)策略。
Details
Motivation: 现有的高斯表示预训练方法(如掩码高斯重建)主要捕获局部结构,但提供的语义监督有限。本文旨在通过整合多模态监督来增强3D高斯表示的语义学习能力和可迁移性。
Result: 在ModelNet40和ScanObjectNN (PB-T50-RS)下游任务上,GaussFusion分别比Gaussian-MAE提升了0.61%和3.85%,证明了其有效性。
Insight: 核心创新点在于将多模态(图像和文本)监督引入3D高斯表示的预训练,并通过提出的GSHM掩码策略适应其非均匀特性,从而学习到更丰富的语义信息,提升了表示的可迁移性。
Abstract: 3D Gaussian Splatting provides an explicit representation that jointly models geometry and appearance, serving as a scalable foundation for 3D representation learning. Existing pre-training methods for Gaussian representations, such as masked Gaussian reconstruction, primarily capture local structures but offer limited semantic supervision. In this paper, we propose GaussFusion, a multimodal pre-training framework for 3D Gaussian representations. GaussFusion integrates image and text supervision into masked Gaussian modeling through cross-modal semantic alignment, enabling the Gaussian encoder to learn both visual and language-level semantic information during pre-training. To better adapt masked modeling to the non-uniform distribution of Gaussian primitives, we further propose Gaussian Salience-guided Multi-scale Hole Masking (GSHM). GSHM constructs spatially continuous masked regions based on Gaussian salience. By applying hole masks at multiple scales, GSHM encourages the encoder to capture both fine-grained local patterns and broader structural dependencies. Extensive experiments on downstream tasks demonstrate that GaussFusion improves the transferability of Gaussian representations. Notably, GaussFusion outperforms Gaussian-MAE on ModelNet40 and ScanObjectNN (PB-T50-RS) by 0.61% and 3.85%, respectively.
[44] Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning cs.CVPDF
Ziyi Chen, Haoyan Shi, Sunhan Xu, Congyan Lang
TL;DR: 本文提出PRPC(渐进推理与基元校正框架),用于组合零样本学习(CZSL)任务。该方法将CZSL建模为结构化、问答式的思维链推理过程,通过逐步推理显式建模属性与对象间的双向依赖关系,并利用基元相互校正抑制早期预测误差。
Details
Motivation: 现有CZSL方法要么独立预测属性和对象,忽略了它们之间的强上下文依赖;要么使用单向条件建模(如对象引导的属性预测),容易导致错误传播。PRPC旨在通过渐进式推理解决这些问题,实现更鲁棒的组合泛化。
Result: 在三个CZSL基准测试上进行的大量实验表明,PRPC实现了最先进的性能,验证了渐进推理和双向校正对于鲁棒组合泛化的有效性。
Insight: 创新点在于将CZSL任务形式化为结构化的思维链推理过程,并引入基于GRPO目标的强化学习后训练,提供与渐进推理过程对齐的步骤级奖励,以增强中间推理的可靠性和逻辑一致性。从客观角度看,其双向校正机制和步骤约束的推理流程是提升组合泛化能力的关键设计。
Abstract: Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. Prior works either predict attributes and objects independently, missing their strong contextual dependency, or use unidirectional conditional modeling (e.g., object-guided attribute prediction), which is prone to error propagation. We propose PRPC, a Progressive Reasoning framework with Primitive Correction, which explicitly models the bidirectional dependency between attributes and objects via step-wise inference. PRPC performs mutual correction of primitives to suppress prediction errors in earlier steps. Specifically, we formulate CZSL as structured, Q&A-style Chain-of-Thought reasoning process and constrain the MLLM to follow predefined semantic steps to generate intermediate decisions. To further enhance the reliability and logical consistency of intermediate reasoning, we introduce reinforcement learning post-training with a GRPO-based objective, providing step-level rewards aligned with the progressive inference procedure. Extensive experiments on three CZSL benchmarks demonstrate that PRPC achieves state-of-the-art performance, validating the effectiveness of progressive reasoning and bidirectional correction for robust compositional generalization.
[45] PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails cs.CV | cs.AI | cs.CLPDF
Mingyang Song, Luxin Xu, Haoyu Sun, Minzhou Pan, Yu Cheng
TL;DR: 本文研究了策略自适应图像护栏问题,即模型需根据当前提供的安全策略判断图像是否违规,并泛化到未见过的策略定义。作者提出了PolicyShiftBench基准测试集和PolicyShiftGuard模型,后者通过两阶段训练方法显著提升了策略敏感性能。
Details
Motivation: 现有图像护栏通常在固定安全策略下训练和评估,将安全性视为图像的固有属性,但实际部署中策略会因产品、时间而变化,因此需要模型能适应动态策略。
Result: 在提出的PolicyShiftBench基准上,7B参数的PolicyShiftGuard模型取得了76.9平均F1分数和72.1平均PSS分数的SOTA性能,并在UnSafeBench和SafeEditBench上表现出良好的迁移性,同时以简洁的输出格式改善了延迟-性能权衡。
Insight: 创新点在于提出了策略自适应护栏的新问题设定和基准,以及结合随机策略SFT与边界对策略适应的两阶段训练方法,其中匹配的通过/阻止边界对对于稳定的策略适应至关重要。
Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.
[46] Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention cs.CV | cs.AIPDF
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen, Tal Remez
TL;DR: 本文提出了一种无需训练的后处理方法——多令牌局部注意力(MTLA),用于评估多模态大语言模型(MLLM)在图像、视频和音频中定位预测(如边界框或时间窗口)的可信度。该方法通过分析模型预测令牌对声称区域的注意力强度来区分真实定位与幻觉,显著提升了检测性能。
Details
Motivation: 多模态大语言模型在输出定位预测(如物体边界框或事件时间窗口)时经常产生幻觉,而模型自身的令牌对数概率无法有效区分定位质量与输入歧义,因此需要一种无需训练的方法来可靠评估定位置信度。
Result: 在多个MLLM家族和三种模态(图像、视频、音频)上,MTLA将幻觉检测的AUROC提升了7到38个点,优于所有先前的无需训练基线。作为置信度分数用于重排序时,它将一个开源8B通用模型的零样本COCO检测AP从20.4提升至37.0,大幅缩小了与有监督检测器的差距。
Insight: 创新点在于提出了一种聚焦于声称区域内部、并聚合所有预测令牌注意力的后处理评分机制,这比先前对整个输入模态求和且仅读取单个响应令牌的注意力方法更有效;该方法无需训练,可泛化至多种模态和任务,显著提升了定位置信度评估的可靠性。
Abstract: Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prolifically. The model’s own token log-probabilities are nearly uninformative: they conflate grounding quality with input ambiguity, and coordinate tokens become near-deterministic once the model commits. We propose Multi-Token Localized Attention (MTLA): a training-free, post-hoc score that measures how strongly a prediction’s tokens attend to the region they claim. Prior attention-based detectors, which sum attention over the entire input modality and read a single response token, are weaker special cases; we show that summing only within the claimed region and aggregating across all prediction tokens recovers a stronger grounding signal. The same recipe applies almost trivially to other modalities and tasks: object detection in images and temporal localization in video and audio. Across multiple MLLM families and three modalities, MTLA improves hallucination AUROC by +7 to +38 over the best prior training-free baseline. Used as a confidence score for re-ranking, it nearly doubles the zero-shot COCO detection AP of an open-source 8B generalist (from 20.4 to 37.0), narrowing the gap to supervised detectors without any task-specific training.
[47] SpecTrack: Spectral Prompt Guided Adaptive Experts for Multispectral Object Tracking cs.CVPDF
Xingyu Tan, Yunrong Qin, Mengjie Hu
TL;DR: 本文提出了一种名为SpecTrack的多光谱目标跟踪方法,通过自适应容量分配来处理不同复杂度的搜索区域。该方法的核心是光谱自适应专家混合模块,结合光谱提示路由器来选择不同能力的专家,并在多个基准测试中实现了精度与效率的良好平衡。
Details
Motivation: 现有多光谱跟踪器通常对所有搜索区域使用固定容量的光谱-空间路径,忽略了不同帧和目标状态下跟踪难度的显著差异,导致在清晰区域计算冗余,而在模糊边界或光谱相似干扰物处推理能力不足。
Result: 在MUST、MSITrack和HOTC20基准测试上取得了有利的精度-效率权衡。精度导向的SpecTrack-L384在三个基准上分别达到了65.2%、51.9%和72.6%的AUC,达到SOTA或极具竞争力水平;平衡的SpecTrack-B224在MUST上以43.7 FPS达到62.4% AUC。在GOT-10k上的额外评估显示其在RGB域具有架构泛化能力,SpecTrack-L384达到79.3% AO。
Insight: 创新点在于将多光谱跟踪建模为搜索区域级别的自适应容量分配问题,并设计了光谱自适应专家混合模块和光谱提示路由器,根据区域复杂度动态激活不同能力的专家子集,同时共享全局专家以提供公共上下文,减少稀疏路由决策的碎片化。
Abstract: Multispectral image(MSI) and hyperspectral image(HSI) object tracking object tracking exploits recorded band-wise observations to improve target–background discrimination under similar RGB appearance, mixed pixels, illumination variation, occlusion, and clutter. However, existing trackers commonly process all search regions through a fixed capacity spectral–spatial path, ignoring that tracking difficulty varies substantially across frames and target states. Clear regions may require only lightweight local discrimination, whereas ambiguous boundaries and spectrally similar distractors often demand stronger contextual reasoning. To address this limitation, we propose SpecTrack, a spectral–spatial complexity-aware tracker that formulates MSI tracking as search-region-level adaptive capacity allocation. Its core component, the Spectral Adaptive Mixture-of-Experts (SAMoE) module, provides a capacity-ordered expert pool with progressively increasing latent rank, receptive field, and depth. Expert selection is guided by a Spectral Prompt Router, which fuses semantic context, spatial boundary cues, and a latent channel-variation cue computed after multispectral patch embedding to activate a sparse subset of SAMoE experts for each search region. In parallel, a Shared Global Expert supplies common latent spectral–spatial context to reduce fragmented sparse-routing decisions. Experiments on MUST, MSITrack, and HOTC20 demonstrate a favorable accuracy–efficiency trade-off. The accuracy-oriented SpecTrack-L384 achieves state-of-the-art or highly competitive AUCs of 65.2%, 51.9%, and 72.6% on the three benchmarks, while the balanced SpecTrack-B224 reaches 62.4% AUC at 43.7 FPS on MUST. An additional GOT-10k evaluation indicates RGB-domain architectural generalization, with SpecTrack-L384 achieving 79.3% AO.
[48] SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation cs.CVPDF
Shenbo Xie, Mingrui Cai, Xu Yang, Yifei Liu, Changxing Ding
TL;DR: 本文提出SparseCtrl-HOI框架,用于人-物交互视频生成,通过稀疏关键帧控制而非密集逐帧姿态序列来降低标注成本并提升运动多样性。该方法结合了时间控制旋转位置嵌入和基于多模态大语言模型的运动先验注入模块,以生成物理合理的过渡帧,并构建了SparseHOI-5K数据集进行验证。
Details
Motivation: 解决人-物交互视频生成中因依赖密集时序指导(如逐帧手-物姿态序列)导致的高标注成本和运动合成多样性受限的问题。
Result: 在构建的SparseHOI-5K数据集上进行了综合评估,结果表明该方法显著降低了标注开销,并合成了更优的直播电商视频。
Insight: 创新点包括引入稀疏时序控制框架、时间控制旋转位置嵌入机制以及利用多模态大语言模型提取高层运动先验的模块;客观来看,其将稀疏控制与高级语义先验结合,为降低生成任务对密集标注的依赖提供了新思路。
Abstract: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streaming e-commerce. A primary obstacle in this domain lies in the complexity of modeling fine-grained physical dynamics and the intricate spatial-temporal coordination between human hands and objects. Existing approaches to this problem typically rely on dense temporal guidance, e.g., frame-wise hand-object pose sequences, to strictly control the interaction process. However, such dense guidance incurs high annotation costs and affects motion synthesis diversity. To overcome these limitations, we introduce SparseCtrl-HOI, a novel sparse temporal control framework for HOI video generation. It requires only a few keyframes that capture interaction states at designated timestamps. Specifically, we employ a Time-Controlled Rotary Positional Embedding (TiRoPE) mechanism to temporally anchor these keyframes while preserving their spatial integrity. Subsequently, to govern the dynamics across intermediate frames, we propose a Motion Prior Injection Module that leverages Multimodal Large Language Models (MLLMs) to extract high-level motion priors. This empowers the model to hallucinate logically and physically plausible transitions. Furthermore, we build SparseHOI-5K, a high-quality and richly annotated dataset for HOI video generation with sparse temporal control. Comprehensive evaluations confirm that our method substantially reduces annotation overhead while synthesizing superior live-streaming e-commerce videos. Both our code and dataset are publicly available at https://mpi-lab.github.io/SparseCtrl-HOI.
[49] PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet cs.CV | cs.AIPDF
Xiaopei Wu, Chenshu Hou, Liang Peng, Dan Xu, Binbin Lin
TL;DR: 本文提出PVCap方法,通过PseudoCap数据增强和VoxelCapNet网络架构改进3D密集描述任务。PseudoCap通过实例级随机混合生成具有多样空间布局的伪帧和伪标签,增强模型描述空间关系的能力;VoxelCapNet则设计了基于体素特征的描述网络,为未来研究提供强基线。
Details
Motivation: 针对现有3D密集描述方法的两大局限:数据增强仅使用全局刚性变换导致空间布局多样性不足,以及网络架构设计简单导致语义信息提取不充分。
Result: 在ScanRefer和Nr3D基准测试中,CIDEr@0.5IoU指标分别超越当前最优方法11.41%和13.99%,达到新的SOTA水平。
Insight: 创新点包括实例级数据增强生成伪标签的PseudoCap框架,以及适配体素网络的VoxelCapNet架构;其数据增强策略和网络设计思路可推广至其他3D视觉-语言任务。
Abstract: 3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous methods, they suffer from two limitations. First, current research often employs global rigid transformations, such as rotation, to augment scenes without changing their spatial layouts. However, diverse spatial layouts are crucial for training a 3D dense captioning model to describe spatial relations between objects. Second, previous works mainly focus on the design of the caption generation pipeline while utilizing a simple network architecture for other components, i.e., backbone and detection head, which is crucial for extracting rich semantic information for captioning. In this paper, we propose PVCap to alleviate the aforementioned problems. Our PVCap consists of PseudoCap and VoxelCapNet. Specifically, PseudoCap employs a random mixing technique on instances within the dataset, generating numerous pseudo frames with diverse spatial layouts at the instance level. By utilizing a teacher-student framework, PseudoCap obtains pseudo caption labels for these pseudo frames. This data augmentation approach significantly increases the number of training samples and enhances the model’s ability to describe the environment effectively. Regarding VoxelCapNet, we introduce a robust caption network that utilizes voxel features and adapts the caption head to the voxel-based network architecture. Our VoxelCapNet can serve as a competitive baseline for future research on 3D dense captioning. Extensive experiments are conducted on two prevalent benchmarks, i.e., ScanRefer and Nr3D. Notably, our method surpasses current state-of-the-art by 11.41% and 13.99% in CIDEr@0.5IoU, respectively. Codes will be made publicly available.
[50] OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations cs.CVPDF
Jun Wei, Xinchang Liu, Yu Liu, Chuhua Yang, Shuhui Wang
TL;DR: OBBSeg提出了一种基于定向边界框(OBB)注释的中间监督范式,用于医学图像中不规则病灶的分割。该方法通过联合编码空间范围和方向,提供更紧凑的几何监督,以减少粗粒度框标注的模糊性,并引入可微分的Mask-to-OBB损失来缓解OBB固有的矩形偏差。在5种成像模态的13个数据集上的实验表明,OBBSeg不仅优于现有弱监督方法,而且达到了与全监督方法相当的性能。
Details
Motivation: 解决医学图像分割中像素级标注成本高昂的问题,利用弱监督作为替代方案,但现有弱监督方法(如轴对齐边界框)对细长或各向异性病灶的标注存在模糊性,因此提出使用定向边界框(OBB)作为中间监督来弥合全监督与弱监督之间的差距。
Result: 在13个数据集(涵盖5种成像模态)上的广泛实验表明,OBBSeg超越了现有弱监督方法,并达到了与全监督方法相当的性能水平,展示了其在医学图像分割中的高效性和可扩展潜力。
Insight: 创新点包括:1) 引入定向边界框(OBB)作为中间监督,更好地对齐不规则病灶的几何形状;2) 提出可微分的Mask-to-OBB损失,缓解OBB的矩形偏差;3) 结合提示驱动的语义引导模块(PAFE和DBFE)来增强前景表示并抑制背景干扰。从客观角度看,该方法通过几何约束和语义增强的联合优化,为弱监督医学图像分割提供了新的有效范式。
Abstract: Pixel-level annotation remains a major bottleneck in medical image segmentation, making weak supervision an attractive yet under-constrained alternative. We propose OBBSeg, an intermediate supervision paradigm guided by Oriented Bounding Boxes (OBBs) that bridges the gap between full and weak supervision. By jointly encoding spatial extent and orientation, OBBs provide compact geometric supervision that better aligns with elongated or anisotropic lesions, reducing the ambiguity of coarse box annotations. To mitigate the inherent rectangular bias of OBBs, we introduce a Mask-to-OBB loss, a differentiable formulation that enforces geometric consistency between predicted masks and OBB regions. Furthermore, we incorporate prompt-driven semantic guidance through two complementary modules-PAFE and DBFE-which enhance foreground representation and suppress background interference. Extensive experiments on 13 datasets across 5 imaging modalities show that OBBSeg not only outperforms existing weakly supervised methods but also achieves performance comparable to fully supervised approaches, demonstrating its potential for efficient and scalable medical image segmentation. The code is available at https://github.com/StarLxc3/OBBSeg.
[51] EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping cs.CV | cs.AIPDF
Innocent Onyenonachi, Peter J. Lawerance, Nadia Kanwal
TL;DR: 本文提出了一个名为EcoVision的AI驱动框架,用于利用无人机获取的高分辨率RGB图像进行盐沼植被监测和优势度制图。该框架采用模块化流程,结合了基于Transformer的语义分割、连通分量植被提取、使用ConvNeXt架构的细粒度物种分类以及基于网格的优势度评分,以2x2米的分辨率对两种重要的盐生禾草进行监测。
Details
Motivation: 旨在开发一个可扩展、高分辨率的盐沼监测系统,将AI驱动的像素级预测转化为生态学上可解释的指标,以替代或补充传统实地调查。
Result: 在分割任务上取得了可靠的物种掩码(平均IoU为0.56,像素级精度为0.96),对象级分类实现了极好的区分度(F1分数为0.99)。优势度估计与基于样方的实地调查结果高度吻合,平均绝对差异低于8%,并在实际调查条件下保持了精细的空间结构。
Insight: 创新点在于构建了一个从像素级语义分割到生态学优势度指标生成的完整、模块化AI工作流,并成功整合了Transformer和ConvNeXt等先进模型进行细粒度分类,为生态监测提供了可操作的高分辨率制图方案。
Abstract: High-resolution RGB imagery acquired from low-altitude UAV surveys was processed through a modular pipeline incorporating transformer-based semantic segmentation, connected-component vegetation extraction, fine-grained species classification using a ConvNeXt architecture, and grid-based dominance scoring at 2x2m resolution. The framework targeted two ecologically significant halophytic grasses, Spartina maritima and Puccinellia maritima, and was trained using a curated and manually annotated UAV imagery, along with biodiversity imagery sourced from publicly accessible datasets. In order to identify these plants from the imagery, our segmentation yielded reliable species masks (mean IoU = 0.56; pixel-level accuracy = 0.96), while object-level classification achieved very good discrimination (F1 = 0.99). Dominance estimates closely matched quadrat-based field surveys, with mean absolute differences below 8%, preserving fine-scale spatial structure under realistic survey conditions. The developed system, named EcoVision, establishes a practical foundation for scalable, high-resolution salt marsh monitoring, demonstrating how AI-driven workflows can translate pixel-level predictions into ecologically interpretable metrics.
[52] MobileWan: Closing the Quality Gap for Mobile Video Diffusion cs.CVPDF
Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv, Ioannis Lelekas, Noor Fathima
TL;DR: 本文提出了MobileWan,一个将服务器规模的5B参数视频扩散模型高效部署到移动设备上的方法。通过循环重构、结构化压缩和内存优化解码等技术,首次在商用移动设备上实现了5B参数规模的视频生成,在保持高质量的同时显著降低了延迟。
Details
Motivation: 现有移动视频扩散模型受限于参数规模(0.4-1.8B),导致生成质量不高。本文旨在证明高质量移动视频生成无需依赖小模型,而是可以通过优化技术将大型服务器模型部署到内存受限的移动硬件上。
Result: MobileWan在移动设备上以20秒端到端延迟生成5秒480x832分辨率、16 FPS的视频,在VBench基准测试中获得83.79分,达到了移动视频生成领域的新SOTA水平。
Insight: 创新点包括:将视频生成重构为分块自回归的循环蒸馏框架,结合因果线性注意力实现推理时RNN化以保持时序一致性;提出基于二元门控的可学习注意力头剪枝方法,通过噪声偏置稀疏目标和蒸馏微调进行端到端优化;结合采样步蒸馏和内存优化VAE解码,实现了大型模型在移动端的首次高效部署。
Abstract: Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coherence. In contrast, existing mobile video diffusion models remain limited to relatively small parameter budgets, typically 0.4-1.8B, restricting generation quality. In this work, we show that high-quality mobile video generation does not require small models. Instead, we demonstrate that a server-scale 5B-parameter video diffusion transformer can be deployed efficiently on memory-constrained mobile hardware through recurrent reformulation and structured compression. Starting from Wan2.2-5B, we rely on a recurrence distillation framework that converts video generation into a chunk-wise autoregressive process with constant-memory attention computation. Combined with causal linear attention, the model operates as an RNN at inference time while preserving temporal coherence across chunks. We further propose a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning. Together with sampling-step distillation and memory-optimized VAE decoding, MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device. Our system generates 5-second 480x832 videos at 16 FPS in 20 seconds end-to-end latency, achieving a VBench score of 83.79 and establishing a new state of the art in mobile video generation. Project page: https://qualcomm-ai-research.github.io/mobilewan
[53] Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability cs.CVPDF
Runfeng Qu, Pia K Bideau, Ole Hall, Julie Ouerfelli-Ethier, Klaus Obermayer
TL;DR: 本文重新审视了场景图生成(SGG)任务,从检测器条件可达性的角度系统分析了基于检测器和基于查询两种主流方法在预测行为上的差异,并发现它们具有互补性。基于此观察,作者提出了一种名为Dual-SGG的双查询方法,旨在整合两种推理机制以利用其互补优势。
Details
Motivation: 动机在于系统分析基于检测器和基于查询的SGG方法因其不同推理机制而产生的预测行为差异,并探索如何有效结合这两种方法以提升性能。
Result: 在Visual Genome、Open Images v6和GQA-200数据集上进行的广泛实验证明了所提Dual-SGG方法的有效性,但摘要中未明确提及具体的定量结果(如与SOTA的比较水平)。
Insight: 创新点在于从检测器条件可达性的新视角对SGG方法进行系统性分析,并设计了一种新颖的双查询架构来融合两种主流推理范式的互补行为,这为改进SGG模型提供了一种新的思路。
Abstract: Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning mechanisms. However, the discrepancy in their predictive behaviors, induced by these distinct mechanisms, has not been systematically analyzed. In this work, we design a controlled experimental setup to examine prediction discrepancies from the perspective of detector-conditioned reachability. The results suggest clear complementary clues. Motivated by this observation, we introduce a Dual-SGG method that consolidates both reasoning mechanisms via a dual-query design, thereby leveraging the complementary predictive behaviors of both detector-based and query-based methods. Extensive experiments on the Visual Genome, Open Images v6, and GQA-200 datasets demonstrate the effectiveness of the proposed method.
[54] Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition cs.CVPDF
Xinda Liu, Qinyu Zhang, Weiqing Min, Guohua Geng, Shuqiang Jiang
TL;DR: 本文提出结构化-压缩提示调优(SCPT)方法,用于增强视觉语言模型在细粒度图像识别中的性能。该方法通过语义关系编码(SRE)显式建模类间语义拓扑结构,并结合语义压缩损失(ScLoss)抑制冗余监督,从而提升语义对齐与细粒度判别能力。在14个细粒度基准测试上的实验表明,SCPT有效缓解了语义模糊性,并在少样本和基类到新类泛化设置中取得了最先进的性能。
Details
Motivation: 细粒度图像识别需要大量专业标注,而现有视觉语言模型(如CLIP)的零样本能力在捕捉细微视觉差异方面有限,且传统提示调优方法将类别标签视为孤立实体,忽略了丰富的语义关系,这限制了模型对层次依赖和类间相关性的建模能力。
Result: 在14个细粒度基准测试上进行的大量实验表明,SCPT方法在少样本和基类到新类泛化设置中均达到了最先进的性能水平。
Insight: 创新点在于将结构化语义关系(通过SRE编码类间拓扑)和语义压缩(通过ScLoss提取判别性成分)引入提示学习,以增强对复杂标签语义的理解,从而提升细粒度分类的判别能力。
Abstract: Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing reliance on extensive labeled data. However, their ability to capture subtle distinctions remains limited, leading to subpar recognition performance. While prompt tuning has proven effective for adapting VLMs, most existing methods treat class labels as isolated, discrete entities, overlooking the rich semantic relationships between them. This oversimplified assumption limits the model’s ability to capture hierarchical dependencies and inter-class correlations – both critical for distinguishing visually similar categories. The problem is especially acute in fine-grained classification, where accurate recognition depends on understanding complex label semantics. To address this, we propose Structured-Condensed Prompt Tuning (SCPT), which enhances semantic structure modeling in prompt learning. Specifically, we introduce Semantic Relation Encoding (SRE) to explicitly model inter-class semantic topology and encode structured label relationships. In parallel, we design a Semantic Condensation loss (ScLoss) to suppress redundant supervision and extract discriminative components from the global semantic space. Together, these components significantly improve semantic alignment and fine-grained discrimination. Extensive experiments on 14 fine-grained benchmarks show that SCPT effectively mitigates semantic ambiguity and achieves state-of-the-art performance in both few-shot and base-to-novel generalization settings.
[55] MoWorld: A Flash World Model cs.CVPDF
Team Moxin, Deyi Ji, Tianrun Chen, Xin Zhang, Jiale Yang
TL;DR: MoWorld是一个高效且实用的Flash世界模型,通过端到端框架实现从数据生成、预训练、蒸馏到高效推理的全流程优化,能够在无需高端GPU的情况下以高达50 FPS的帧率实现实时交互,并保持电影级视觉质量。
Details
Motivation: 为了解决世界模型在现实自主系统中对高帧率推理的需求,提升响应式感知、规划与控制的实用性,同时优化模型能力与部署成本,以实现大规模现实世界应用。
Result: 在综合评估中,MoWorld实现了领先性能,其平均推理成本仅为现有世界模型的30%-50%,并在神经处理单元(NPU)上达到高达50 FPS的实时交互速度。
Insight: 创新点包括:基于可扩展的3D原生数据引擎构建几何一致的训练数据,采用课程跨帧预训练策略稳定学习,通过高效去噪步蒸馏算法降低扩散训练成本,以及混合精度并行推理框架支持低成本实时部署。
Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive perception, planning, and control in real-world autonomous systems. To this end, we present MoWorld, a cost-effective yet high-performance Flash World Model with an end-to-end framework spanning data generation, pre-training, distillation, and efficient inference, enabling up to 50 FPS real-time interaction with cinematic visual quality without the need of high-end GPUs. To enable large-scale real-world deployment, MoWorld jointly optimizes model capability and cost throughout the entire development pipeline. Specifically, unlike existing approaches that primarily rely on large-scale video corpora, MoWorld is built upon a scalable 3D-native data engine accumulated from our large-scale 3D vision and generative modeling pipeline, enabling the efficient construction of geometrically consistent training data across diverse real-world and synthetic environments. Based on this foundation, a curriculum cross-frame pre-training strategy for stable and scalable World Model learning, an efficient denoising-step distillation algorithm to reduce diffusion training cost, and a mixed-precision parallel inference framework for low-cost real-time deployment. MoWorld is the first real-time interactive World Model built on the Neural Processing Unit (NPU) and can achieves up to 50 FPS in such the devices, enabling practical and efficient deployment at scale. Comprehensive evaluations demonstrate that MoWorld achieves leading performance; notably, its average inference cost is only 30%-50% of that of existing World Models, providing a practical foundation for large-scale real-world applications of World Models. We also demonstrate diverse applications of MoWorld.
[56] EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion cs.CVPDF
Onur Eker, Erkut Erdem, Aykut Erdem
TL;DR: 本文提出了EeveeDark,一个用于极低光视频增强的二进制神经网络框架。它通过事件流引导的传感器级融合,结合RAW数据的空间丰富性和事件流的时间精度,在保持细节的同时利用二进制量化降低计算开销。
Details
Motivation: 解决在资源受限环境下,极低光视频增强任务中难以平衡恢复质量与计算效率的挑战。
Result: 在合成和真实世界数据集上的实验表明,EeveeDark超越了先前的基于二进制神经网络的方法,并与全精度模型相比提供了更优的性能-效率权衡。
Insight: 创新点在于提出了一个结合传感器级RAW数据和事件流的二进制神经网络框架,并设计了模态特定的二进制编码器、轻量级融合块以及事件引导的跳跃门控机制,以实现动态时空细化。从客观角度看,其将事件相机的高时间分辨率与二进制网络的高效性相结合,为低光视频处理提供了一个新颖且实用的解决方案。
Abstract: Enhancing videos under extreme low-light conditions remains challenging due to the difficulty of balancing restoration quality and computational efficiency in resource-constrained settings. This paper introduces EeveeDark, a low-light video enhancement framework that combines the spatial richness of sensor-level RAW data with the temporal precision of event streams. Central to our model is a Binary Neural Network (BNN) architecture that reduces computational overhead by quantizing weights and activations while preserving detail. EeveeDark incorporates (i) modality-specific binary encoders for processing RAW frames and event data, (ii) a lightweight fusion block for integrating spatial and temporal cues, and (iii) an event-guided skip gating mechanism for dynamic spatiotemporal refinement. Experiments on synthetic and real-world datasets show that EeveeDark outperforms prior BNN-based methods and offers a favorable performance-efficiency trade-off compared to full-precision models. The project page is available at https://cyberiada.github.io/EeveeDark.
[57] VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection cs.CV | cs.AIPDF
Sharayu N. Deshmukh, Md Rashidunnabi, Nelton Tiago Gemo, Kurundkar G. D., Mahamune M. R.
TL;DR: 本文提出了VendorBench-100,一个用于深度伪造图像检测的统一跨范式基准测试。该基准使用一个精心构建的100张图像对抗语料库、统一的输出模式和评估框架,对36个代表性模型(包括商业API、零样本视觉语言模型和开源检测器)进行了评估。核心发现是模型在排名能力(ROC-AUC)与在默认阈值下的决策质量(MCC)之间存在系统性差异。
Details
Motivation: 当前深度伪造图像检测领域存在商业API、零样本视觉语言模型和开源检测器三种不同范式,但缺乏一个统一的评估协议来对它们进行直接、公平的比较。
Result: 在VendorBench-100基准上,商业API取得了最强的中位数性能,其次是视觉语言模型和开源检测器,但部分开源模型仍能与最好的视觉语言模型竞争。评估主要使用马修斯相关系数(MCC)进行排名,并报告ROC-AUC作为阈值无关的排名能力指标。
Insight: 论文的主要创新在于构建了一个强调现实世界挑战性场景(包含八类边缘案例)的小型但精炼的跨范式基准。其核心洞察是揭示了模型在排名能力(ROC-AUC)与操作点质量(MCC)之间普遍存在的不一致性,这比单一的排行榜排名更具指导意义。
Abstract: Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors. Despite their widespread use, these paradigms are rarely evaluated under a common protocol, making direct comparison difficult. We introduce VendorBench-100, a cross-paradigm benchmark that evaluates 36 representative models using a single adversarial 100-image corpus, a unified output schema, and a common evaluation framework. To ensure reliable assessment under the corpus’s intentional class imbalance, models are ranked primarily by the Matthews correlation coefficient (MCC), with ROC-AUC reported as a threshold-independent measure of ranking ability. Rather than maximizing dataset size, VendorBench-100 emphasizes challenging real-world scenarios through a curated taxonomy of eight edge-case families, including face swaps, text-to-video stills, AI photo edits, avatar compositing, opaque-provenance images, and compressed research frames. Our evaluation shows that commercial APIs achieve the strongest median performance, followed by vision LLMs and open-source detectors. However, individual open-source models remain competitive with the best vision LLMs. More importantly, we identify a consistent divergence between ranking ability (ROC-AUC) and operating-point quality (MCC), demonstrating that strong score discrimination does not necessarily produce reliable default-threshold decisions. This metric disagreement, rather than any single leaderboard ranking, is the central finding of the benchmark. We release the complete evaluation framework and benchmark results to support reproducible future research. The source code and data are available at: https://github.com/sharayu-20/vendorbench-100
[58] Straight-Path Flow Matching for Incomplete Multi-View Clustering cs.CVPDF
Yiteng Yuan, Junyan Wang, Zheyuan Liu, Hong Jia, Lei Fan
TL;DR: 本文提出了一种用于不完整多视图聚类的直线路径流匹配方法。该方法通过设计确定性的概率流路径来替代传统的扩散模型,直接在观测视图和缺失视图之间建立线性插值路径,并结合聚类级对齐和基于熵的对齐机制来增强跨视图聚类一致性。
Details
Motivation: 解决现有基于扩散模型的生成式不完整多视图聚类方法存在的问题:这些方法从与聚类无关的噪声初始化,依赖随机去噪动态,没有明确针对聚类目标进行设计。
Result: 在标准IMVC基准测试上的广泛实验表明,所提出的框架实现了新的最先进(SOTA)性能。
Insight: 创新点在于:1)用确定性的ODE流替代随机扩散轨迹,证明其在有限步长机制下能更好地保持聚类一致性;2)设计了直线路径流匹配视图补全机制;3)结合了聚类级和基于熵的对齐来强制跨视图聚类一致性。
Abstract: Incomplete Multi-View Clustering addresses the problem of clustering multi-modal data when certain views are missing. Recent end-to-end generative approaches leverage diffusion models to recover missing views via stochastic noise-to-data trajectories. While expressive, such mechanisms are not explicitly designed for clustering, as they initialize from cluster-agnostic noise and rely on stochastic denoising dynamics. In this work, we revisit probability path design in end-to-end generative IMVC. We introduce a flow-matching framework with a linear interpolation path between paired view representations, that replaces diffusion with probability flows between observed and missing views. We provide a formal analysis showing that deterministic ODE flows are inherently better aligned with clustering objectives than diffusion-based stochastic trajectories, especially in terms of transport mechanisms that respect class-conditional data distributions and maintain cluster consistency in finite-step regimes. Building upon this insight, we develop an end-to-end IMVC architecture that integrates straight-path flow-matching view completion with cluster-level and entropy-based alignment to enforce cross-view clustering consistency. Extensive experiments on standard IMVC benchmarks demonstrate that the proposed framework establishes new state-of-the-art performance.
[59] MAC-XA: Multi-view Anatomy-Correspondence Fusion for Coronary Stenosis Reporting from X-ray Angiography cs.CVPDF
Chen Jia, Baochang Zhang, Fatia Kusuma Dewi, Amir Yousefi, Heribert Schunkert
TL;DR: 本文提出MAC-XA方法,用于从冠状动脉X射线血管造影中生成结构化狭窄报告。该方法通过可控合成血管造影生成策略,学习跨视图解剖对应关系,在融合前将辅助视图特征与主视图坐标空间对齐,从而约束证据聚合到解剖一致区域。
Details
Motivation: 解决冠状动脉X射线血管造影多视图推理中的关键问题:由于血管三维拓扑结构导致投影依赖的分支重叠和透视缩短,单视图建模对病变定位和狭窄分级不完整且不稳定,而传统多视图融合缺乏可观察的跨视图对齐监督。
Result: 在合成数据和真实血管造影的零样本迁移实验中,该对齐约束设计相比单视图建模和传统多视图融合方法,提高了对应关系一致性和结构化狭窄报告性能。
Insight: 创新点在于将多视图狭窄报告重新表述为对齐约束的聚合问题,并引入可控合成数据生成来提供几何驱动的补丁级对应监督;通过解剖对应模块显式学习跨视图对应矩阵,实现基于解剖一致性的特征融合。
Abstract: Multi-view reasoning in coronary X-ray angiography is inherently a cross-projection geometric problem, yet automated report generation in this setting remains largely unexplored. The 3D vascular topology leads to projection-dependent branch overlap and foreshortening, rendering single-view modeling fundamentally incomplete and unstable for lesion localization and stenosis grading. Although multi-view fusion appears promising, learning anatomically consistent fusion from real angiograms is impeded by a critical limitation: cross-view alignment is unobservable and cannot be explicitly supervised. Consequently, conventional fusion relies on implicit correlations rather than verified anatomical correspondence. We address this by reformulating multi-view stenosis reporting as an alignment-constrained aggregation problem. A controllable synthetic angiography generation strategy is introduced to expose geometry-derived patch-level correspondence supervision unavailable in real data. An anatomy-correspondence module learns cross-view correspondence matrices that explicitly align auxiliary features within the main-view coordinate space prior to fusion, thereby constraining evidence aggregation to anatomically consistent regions. Experiments on synthetic data and zero-shot transfer to real angiograms show that this alignment-constrained design improves correspondence consistency and structured stenosis reporting compared to single-view modeling and conventional multi-view fusion methods. The code will be publicly available upon publication.
[60] AlayaWorld: Long-Horizon and Playable Video World Generation cs.CV | cs.HCPDF
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge
TL;DR: AlayaWorld是一个用于构建交互式生成世界的全栈开源框架,它通过视频世界模型自回归合成未来观测,支持用户在虚拟环境中自由导航和进行多样化交互(如战斗、施法等),旨在替代传统高成本、难定制的游戏世界制作流程。
Details
Motivation: 传统游戏世界构建依赖劳动密集型生产流程,成本高、定制难且部署后修改昂贵;视频世界模型提供了一种新范式,通过自回归合成观测来在线生成可玩世界,为游戏及具身智能等交互应用开辟新机会。
Result: 论文未在摘要中提及具体定量结果或基准测试,但发布了可复现的流水线、参考实现、评估工具和完整文档,为生成世界模型的未来研究和实时应用奠定了实践基础。
Insight: 创新点在于提出了一个模块化、可扩展的全栈框架,统一了从数据准备、模型架构、训练、推理加速到部署的完整开发流程,并支持开放式的实时交互,推动了生成世界模型从研究向实际应用的转化。
Abstract: Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than explicitly authoring every component of a virtual environment, these models autoregressively synthesize future observations conditioned on the current world state and user interactions, enabling playable worlds to be generated online. Trained on both gameplay recordings and real-world videos, they can capture diverse visual appearances and physical dynamics, opening new opportunities for interactive applications beyond gaming, including embodied intelligence. In this paper, we present \textbf{AlayaWorld}, a full-stack open-source framework for building interactive generative worlds. AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning. The framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-within a modular and extensible architecture. Alongside the framework, we release reproducible pipelines, reference implementations, evaluation tools, and comprehensive documentation, establishing a practical foundation for future research and real-time applications of generative world models.
[61] Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification cs.CV | cs.AIPDF
Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
TL;DR: 本文提出了一种基于token的双视图学习框架,用于乳腺钼靶图像的乳腺癌分类。该框架在冻结的视觉Transformer骨干网络中,通过引入专门的融合token,在多个网络深度实现CC和MLO视图之间的渐进式、结构化交互,从而有效整合互补信息。
Details
Motivation: 现有多视图学习方法通常依赖于特征级聚合或单阶段交叉注意力,这可能导致视图特定信息和共享表示纠缠,且交互仅限于有限的网络深度。本文旨在解决这些限制,实现更有效的双视图信息融合。
Result: 在VinDr-Mammo和CMMD数据集上的实验表明,该方法在线性探测、仅提示调优和传统融合基线方法上均取得了持续改进。在VinDr-Mammo BI-RADS分类任务中,取得了50.40%的F1分数和0.8090的AUC,在二分类设置中比双视图融合基线AUC提高了0.10。
Insight: 创新点在于将视图间交互重新定义为结构化的token级通信,使用专门的融合token作为跨视图依赖关系的中间载体,并在多个Transformer深度插入融合模块以实现分层传播。这避免了直接特征融合的局限性,并保留了视图特定的结构信息。
Abstract: Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.
[62] VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery cs.CVPDF
Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen
TL;DR: 本文提出VaseMuseum,一个用于古希腊陶器的轻量级模块化多模态智能数字博物馆框架,通过结合交互式虚拟博物馆与VaseAgent,整合多模态感知、3D感知推理、外部知识检索和推理时可靠性控制,以解决开放解释中细粒度视觉证据的可靠引用和证据不完整时的校准不确定性两大挑战。
Details
Motivation: 针对文化遗产领域(如古希腊陶器)中,现有视觉语言模型在开放解释时难以基于细粒度2D/3D视觉证据进行可靠的知识检索,以及在证据不完整、有噪声或模糊时容易产生自信但无依据的回答的问题。
Result: 在真实的数字博物馆模拟实验中,VaseMuseum相比支持检索的VLM基线,提高了引用有效性,减少了知识密集型查询中的幻觉,并在模糊情况下产生更中立的回答。
Insight: 创新点包括:1)结合权威网络和博物馆知识源进行证据检索,并在生成前通过源级控制选择多样且可验证的证据;2)通过响应级控制检查生成主张与证据池的一致性,鼓励在支持不足或冲突时给出中立、证据有界的答案;3)采用无需训练的GRPO风格选择机制,在不更新VLM主干的情况下偏好具有有效引用和校准置信度的响应。
Abstract: Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.
[63] FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration cs.CVPDF
Jin Jiang, Jia Wang, Panwen Hu, Weiran Zhao, Shengcai Liao
TL;DR: 本文提出FADRA,一种用于视频人脸修复(VFR)的频率感知扩散框架,通过迭代残差适应来平衡空间保真度和时间一致性。该方法利用预训练文本到视频扩散模型的时间一致性先验,引入轻量级LoRA适配器和低质量像素对齐特征融合模块进行任务适应,并设计了重复残差适应头(RRAH)进行逐步残差细化,同时采用频率感知损失在多个频段提供显式监督。
Details
Motivation: 现有视频人脸修复方法在复杂退化条件下难以平衡空间细节保真度和时间连贯性,因此需要一种鲁棒的方法来同时恢复高质量面部细节并保持时间一致性。
Result: 在广泛的实验中,FADRA在定量指标和视觉感知上均优于现有最先进方法,恢复了更好的面部结构并产生了更时间一致的视频。
Insight: 创新点包括:1)结合预训练视频扩散模型的时间先验与轻量级适配(LoRA和LQ特征融合)进行高效任务适应;2)提出LQ引导的重复残差适应头(RRAH),在流匹配步骤中反复利用退化观测进行逐步残差细化;3)引入频率感知损失,对感知重要的频段进行显式监督以提升感知质量并减少时间抖动。
Abstract: Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, existing methods still struggle to balance spatial fidelity and temporal coherence under complex degradations. To address this, we propose FADRA, a frequency-aware diffusion framework with iterative residual adaptation specifically tailored for robust VFR. We first leverage the strong temporal consistency of a pre-trained text-to-video diffusion model and introduce lightweight LoRA adapters together with a Low-Quality (LQ) Pixel-Alignment Feature Fusion module to efficiently adapt the frozen generative prior to the VFR task. To further adapt the frozen diffusion backbone to the downstream VFR task beyond LoRA-based adaptation, we introduce a Repeated Residual Adaptation Head (RRAH) for step-wise residual refinement after the diffusion backbone. To make this refinement explicitly guided by the degraded observation, RRAH further takes the LQ latent together with the current velocity prediction as input, allowing the model to repeatedly revisit LQ cues and predict residual updates at each flow-matching step. This LQ-guided repeated residual adaptation helps recover fine facial details while preserving the inherent temporal priors of the pre-trained model. Furthermore, to ensure the structural integrity of perceptually important details, we introduce a Frequency-Aware Loss that provides explicit supervision across multiple spectral bands, emphasizing visually sensitive frequency components that are crucial for perceptual quality and prone to temporal jittering. Extensive experiments demonstrate that FADRA recovers better facial structures and produces more temporally consistent videos than state-of-the-art methods, leading to clear gains in both quantitative metrics and visual perception.
[64] What Images Cannot Say: Language-Guided Olfactory Representation Learning cs.CV | cs.AI | cs.LGPDF
Eleftherios Tsonis, Xi Wang, Vicky Kalogeiton
TL;DR: 本文提出了SCENT,一种利用语言指导作为视觉与嗅觉之间语义桥梁的多模态框架。该框架通过视觉语言模型生成场景描述符,捕捉物体、环境上下文和视觉场景暗示的潜在气味线索,从而指导嗅觉表征学习。实验表明,SCENT在气味-图像和气味-文本检索任务上实现了最先进的性能,并能解耦复杂气味混合物。
Details
Motivation: 解决图像与嗅觉信号对齐的挑战,因为许多嗅觉线索源于环境上下文因素,这些因素在像素中不可直接可见。
Result: 在New York Smells数据集上,SCENT显著优于仅视觉基线,在气味-图像和气味-文本检索任务上达到SOTA水平。
Insight: 创新点包括使用语言作为视觉与嗅觉的语义桥梁,以及引入语言指导的潜在分解来分离物体特定气味与环境贡献,从而提升跨模态检索的可解释性。
Abstract: Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging because many olfactory cues arise from contextual environmental factors that are not directly visible in pixels. We introduce SCENT, a multimodal framework that uses language guidance as a semantic bridge between vision and olfaction. Our approach leverages Vision-Language Models (VLMs) to generate scene descriptors capturing objects, environmental context, and plausible ambient smell cues suggested by the visual scene. These descriptors provide semantic guidance for learning olfactory representations. We train a smell encoder that maps electronic-nose signals into a shared embedding space aligned with both visual and textual representations, and introduce a languageguided latent decomposition that separates object-specific odors from contextual environmental contributions. Experiments on the New York Smells dataset demonstrate that SCENT significantly improves crossmodal retrieval compared to vision-only baselines, achieving state-of-theart performance on smell-to-image and smell-to-text retrieval tasks. In addition, our framework produces interpretable olfactory representations that enable the disentanglement of complex smell mixtures. Our results reveal the importance of contextual semantic information for grounding olfactory perception in multimodal learning and pave the way for future research in this area.
[65] Temporal Modeling of Optically Variable Devices in Identity Documents cs.CVPDF
Glen Pouliquen, Joseph Chazalon, Guillaume Chiron, Oriol Ramos Terrades, Thierry Géraud
TL;DR: 该论文提出两种新颖的方法,用于在开放集场景下验证身份文件中透明光学可变设备(OVD,即“全息图”)的动态行为,旨在通过建模时间动态来提升远程身份文件验证的鲁棒性。这些方法采用自监督训练,无需攻击样本,并在公共数据集上超越了先前的最先进方法。
Details
Motivation: 现有远程身份文件验证系统在处理用户非受控条件下拍摄的视频时存在关键局限:要么孤立处理视频帧,忽略了OVD固有的动态特性,使系统易受交换攻击;要么仅关注一般全息图存在性,无法验证特定OVD类型。此外,逐帧视频标注的经济不可行性使得监督训练不切实际。
Result: 在公共数据集上,所提方法超越了先前的最先进(SOTA)方法,同时严格遵守工业约束。结果表明,建模时间动态对于在现实条件下防御复杂攻击至关重要。
Insight: 论文的创新点在于针对开放集场景(训练时攻击类型未知)设计了专门验证持有者肖像保护OVD动态行为的方法,并首次在自监督设置下无需攻击样本进行训练。从客观角度看,其将序列建模和异常检测应用于OVD验证,为安全特征分析提供了新的技术方向。
Abstract: Robust remote verification of identity documents relies on analyzing faint, transparent security features like Optically Variable Devices (OVDs), or “holograms”, within user-captured videos under uncontrolled conditions. Current systems, however, face critical limitations: existing methods often treat video frames in isolation, neglecting the intrinsic dynamic nature of OVDs and leaving systems vulnerable to swapping attacks, or focus on general holographic presence and lack the ability to verify specific OVD types. Moreover, the economic infeasibility of frame-by-frame video annotation makes supervised training impractical. In this work, we introduce two novel approaches for verifying the dynamic behavior of transparent OVDs protecting the holder’s portrait, specifically designed for open-set scenarios where attack types are unknown during training. We demonstrate that these approaches can be trained without any attack samples in a self-supervised setting, surpassing previous state-of-the-art methods on public datasets while adhering strictly to industrial constraints. Our results confirm that modeling temporal dynamics is essential for defeating sophisticated attacks under realistic conditions, and underscores the promise of sequence modeling and anomaly detection for OVD verification. Code is available at https://github.com/EPITAResearchLab/pouliquen.26.icdar.
[66] Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cs.CV | cs.AIPDF
Yoav Baron, Sara Dorfman, Roni Paiss, Daniel Cohen-Or, Or Patashnik
TL;DR: 本文研究了视觉语言模型(VLM)作为扩散模型图像编辑的条件编码器时,其空间定位能力下降的问题。作者提出了一个名为‘Analysis-by-Proxy’的框架,通过训练一个轻量级代理模型来分析VLM的中间表征,揭示了现有编辑流程在提取定位信号方面的根本性失败。
Details
Motivation: 动机在于解决VLM在作为条件编码器时,其强大的空间定位能力在复杂的多实体场景编辑中无法有效保持的性能差距问题。
Result: 研究发现,在单次前向传播的约束下,定位信号并未可靠地传播到通常用于条件提取的预定义层,而是隐藏在中间表征中,其位置随输入提示词而变化。
Insight: 创新点在于提出了‘Analysis-by-Proxy’分析框架,能够揭示VLM内部定位信息的具体编码位置,这为未来设计更合理的条件提取架构提供了新的思路和方向。
Abstract: Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarkable multimodal reasoning capabilities. While standalone VLMs demonstrate strong localization capabilities, editing pipelines frequently struggle to maintain this accuracy, particularly in complex, multi-entity scenes. In this work, we investigate this performance gap, hypothesizing that it stems from treating the VLM as a condition encoder. In this role, the model is restricted to a single forward pass, preventing the autoregressive generation process for which it was optimized, thereby failing to fully expose its capabilities. To investigate whether this spatial understanding persists when the VLM is used as a condition encoder, we introduce Analysis-by-Proxy. In this framework, we train a lightweight, interpretable proxy model on the VLM’s intermediate representations using an auxiliary localization task. By analyzing the VLM through this proxy, we uncover the specific VLM representations that encode localization information. Our findings expose a fundamental mismatch between how spatial knowledge is represented within a VLM condition encoder and how it is extracted by current editing pipelines. We reveal that under single-pass constraints, the localization signal does not reliably propagate to the predefined layer configurations commonly used for conditioning. Instead, this crucial signal remains hidden within intermediate representations, at locations that vary depending on the input prompt. Using our introduced Analysis-by-Proxy framework, we reveal the fundamental failures of existing condition extraction strategies in editing pipelines, opening the door to more principled design of conditioning architectures.
[67] HoloCount: A Holistic Visual Counting Benchmark for MLLMs cs.CVPDF
Jinhong Deng, Limeng Qiao, Guanglu Wan
TL;DR: 该论文提出了一个名为HoloCount的综合性视觉计数基准测试,用于评估多模态大语言模型(MLLMs)的计数能力。该基准通过一个三层分类法(语义计数、分析计数、鲁棒性测试)来系统评估模型在细粒度定位、空间推理以及对抗性场景下的表现。通过对20多个先进MLLMs的评估,揭示了模型在从感知任务转向复杂分析推理和不利场景时性能显著下降的关键差距。
Details
Motivation: 当前MLLMs在定性场景理解上取得了显著成功,但其定量精度(特别是计数能力)存在显著瓶颈,常伴有数值幻觉。现有的计数基准主要关注简化场景下的基本感知,未能捕捉在逻辑约束或对抗条件下出现的复杂失效模式。
Result: 通过对超过20个最先进的MLLMs进行详尽评估,发现即使顶级模型在任务从感知转向复杂分析推理和对抗场景时,性能也会显著下降。该基准为当前MLLM计数能力提供了系统性的评估图景。
Insight: 论文的创新点在于提出了一个层次化、诊断性强的综合性视觉计数基准(HoloCount),它超越了简单的感知计数,纳入了逻辑组合、空间推理和针对对抗场景(如高密度场景、语言偏见)的鲁棒性测试,为开发更可靠、更接地气的多模态系统提供了路线图。
Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offer a roadmap for developing more grounded and reliable multimodal systems. The dataset is available at https://mm-mvr.github.io/HoloCount/.
[68] Andha-Dhun: A First Look at Audio Descriptions in Hindi cs.CVPDF
Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta, Ganji Sreeram, Pailla Balakrishna Reddy
TL;DR: 本文首次系统研究了印地语音频描述(AD),提出了首个由人工撰写的印地语AD数据集Andha-Dhun,并探索了两种生成方法:从英文密集视频描述直接生成和翻译英文AD。通过困惑度和LLM作为评判指标评估了生成质量,并指出简单翻译会引入伪影并降低多样性,强调印地语AD需为印度视障观众进行内容适配而非严格忠实于源语言。
Details
Motivation: 随着印度电影认证中心(CBFC)的授权,需要将音频描述(AD)扩展到英语之外,但目前尚无针对任何印度语言生成AD的工作,本文旨在填补印地语AD研究的空白。
Result: 使用困惑度和LLM作为评判指标评估了两种生成方法的流畅性和质量;分析发现,与原始印地语AD相比,简单翻译会引入伪影并降低多样性,机器翻译无法适应文化参考,而人工翻译虽表现更好但仍不足。
Insight: 创新点在于构建了首个印地语AD数据集Andha-Dhun,并系统探索了生成与评估方法;客观分析表明,AD生成应优先考虑为目标受众(如印度视障观众)进行内容和文化适配,而非追求对源语言的严格保真,这对多语言可访问性研究具有借鉴意义。
Abstract: Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV shows, and with mandates from India’s Central Board of Film Certification (CBFC), there is a need to expand ADs beyond English. Yet, there is no work that generates ADs for any Indian language. To address this gap, we present the first systematic study of ADs in Hindi, contributing to aspects such as data, generation, and evaluation. We introduce Andha-Dhun, the first dataset of human-authored Hindi ADs collected from 8 full-length movies. We explore two approaches for generating ADs in Hindi: (i) directly from English dense video descriptions, and (ii) translating English ADs into Hindi. We evaluate these approaches using perplexity and LLM-as-a-judge metrics to assess fluency and quality respectively. We also analyze movies that have both English and Hindi human-authored ADs and find that naive translation introduces artifacts and narrows diversity compared to original Hindi ADs. Direct machine translation fails to adapt cultural references, while human-translated ADs do better but still fall short. Our findings emphasize that the purpose of Hindi ADs is accessibility for Indian BLV audiences, and that this requires adapting content for the audience more than strict fidelity to the source.
[69] Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation cs.CV | cs.AIPDF
Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos
TL;DR: 本文提出了PACR-Video,一个参数高效的多镜头长视频外推框架。它通过冻结文本到视频扩散变换器,并引入由学习的镜头角色提示令牌调节的低秩时间适配器,来保持循环实体、场景结构、视觉风格和因果进展,而无需对生成器进行全量微调。
Details
Motivation: 为了解决多镜头长视频生成中,在保持长期一致性(如实体、场景、风格和事件进展)的同时实现参数高效调优的问题。
Result: 在六个多镜头和长视频基准测试中,PACR-Video在分布质量、语义对齐、身份一致性、时间平滑度、运动稳定性、过渡连贯性和人类偏好方面,优于文本到视频、基于调优、记忆增强、流式和递归上下文基线方法。
Insight: 核心创新在于通过递归提示库存储紧凑的实体、位置、动作和风格提示,并根据预测的叙事依赖关系通过适配器门进行路由,结合镜头局部/故事全局的调优目标,实现了对稳定长视频外推的轻量级可控能力。
Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning. PACR-Video keeps a text-to-video diffusion transformer frozen and augments it with low-rank temporal adapters conditioned by learned shot-role prompt tokens. To maintain long-horizon coherence, it builds a recursive prompt bank that stores compact entity, location, action, and style prompts from previous shots, then routes them through adapter gates according to predicted narrative dependencies. A Shot-Local/Story-Global tuning objective combines next-shot reconstruction, cross-shot identity contrast, and prompt sparsity regularization, while an adapter composition schedule balances early-shot visual consistency with later-shot event progression and viewpoint change. Across six multi-shot and long-video benchmarks, PACR-Video outperforms text-to-video, tuning-based, memory-augmented, streaming, and recursive-context baselines on distributional quality, semantic alignment, identity consistency, temporal smoothness, motion stability, transition coherence, and human preference. These results show that compact prompt routing and lightweight temporal adaptation provide sufficient controllable capacity for stable long video extrapolation.
[70] EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage cs.CVPDF
Max Gonzalez Saez-Diez, Jihoon Chung, Adam D. Wolsky, Gregory Lanzalotto, Dean Knox
TL;DR: 本文介绍了EgoPolice数据集,这是一个从公开的警用随身摄像机视频中精心构建的真实第一人称警民互动数据集。该数据集以秒级粒度标注了对警务行为研究至关重要的动作标签,视频具有快速不规则相机运动、密集人际互动和罕见高风险事件等特点,为运动鲁棒和上下文感知的第一人称感知提供了挑战性基准。研究设置了分类和多项选择问答两项任务,并对开源和闭源模型进行了基准测试,发现即使是Gemini 2.5 Pro等最佳视频模型也难以准确预测’武器拔出’等高危动作。
Details
Motivation: 为解决大规模警用随身摄像机视频库中高效识别关键事件(如高风险警民互动)的难题,并为警务行为研究提供数据基础,作者构建了EgoPolice数据集。
Result: 在EgoPolice基准测试中,包括Gemini 2.5 Pro在内的现有最佳视频模型在预测’Weapon Out’等高危动作时仍表现不佳,凸显了该数据集的挑战性。
Insight: 创新点在于构建了首个专注于高风险警民互动的真实第一人称视频数据集,其秒级细粒度标注和包含快速运动、密集交互的场景特性,为开发运动鲁棒、上下文感知的模型提供了独特基准;从客观角度看,该数据集将计算机视觉任务与公共安全领域的实际需求紧密结合,推动了第一人称视频理解在复杂、高风险现实场景中的应用评估。
Abstract: We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera videos. We select police-civilian action labels that are critical for police behavioral research and annotate them at a second-by-second granularity. The videos feature rapid and irregular camera motion, dense human interactions, and rare high-stakes events, making the dataset a challenging benchmark for motion-robust and context-aware egocentric perception. We provide two different tasks, classification and multiple-choice question-answering, and benchmark both open-source and closed-source models. We find that even the best video models like Gemini 2.5 Pro still struggle to accurately predict high-risk actions such as “Weapon Out”. Beyond serving as a benchmark, EgoPolice provides a foundation for developing models capable of identifying events of interest in large-scale body-worn camera video repositories, enabling more efficient downstream human review.
[71] A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation cs.CVPDF
Linlin Zhang, Neema Jakisa Owor, Xiang Yu, Abby Watts, Yaw Adu-Gyamfi
TL;DR: 本文提出了一种新颖的交通标志状况评估框架,该框架整合了日间视觉性能和夜间逆反射性能的综合评估。方法采用三个微调的视觉语言模型(VLM)评估日间视觉性能的四个关键因素,并通过情感分析和CLIP评分转换为数值分数;夜间性能则使用LiDAR数据依据校准程序进行评估。最终,这些组件被整合成一个综合的标志状况指数(SCI)以指导维护。
Details
Motivation: 传统交通标志评估方法(如人工检查)存在主观、劳动密集、安全隐患以及设备昂贵等问题,且现有研究大多只关注日间或夜间单一方面的评估,缺乏对两者的综合考量。本研究旨在开发一个系统性的、成本效益高的框架,以克服这些局限性。
Result: 评估结果表明,在日间视觉性能评估中,LLaVA和Qwen模型的表现优于InternVL,在所有评估因素上的双向余弦相似度得分在0.67-0.76之间。在验证的462个交通标志中,该框架标记出68个因逆反射性能不足而需要立即更换的标志。
Insight: 主要创新点在于首次提出了一个整合日间视觉与夜间逆反射性能的综合评估框架,并创新性地应用微调VLM结合情感分析和CLIP评分进行日间性能的自动化、量化评估。这为交通基础设施的自动化、低成本维护提供了一个可行的技术方案。
Abstract: Traffic signs are crucial components of road safety, serving as visual tools under all lighting conditions. The Manual on Uniform Traffic Control Devices (MUTCD) specifies daytime visual factors such as legibility and color contrast, and nighttime retroreflectivity requirements. Traditional assessment methods rely on manual inspections, which the Federal Highway Administration (FHWA) notes are subjective, labor-intensive and pose safety concerns, while retroreflectometers are expensive and unaffordable for smaller agencies. Most existing studies focus on either daytime factors or nighttime retroreflectivity but rarely integrate both aspects comprehensively. This study develops a novel framework that systematically evaluates traffic signs through integrated daytime-nighttime assessment. The methodology employs three fine-tuned Vision Language Models (VLMs) for daytime visual performance assessment across four key factors: legibility, color, surface and shape integrity, and surrounding environment conditions. VLM predictions are converted to numerical scores through sentiment analysis and Contrastive Language-Image Pre-Training (CLIP) scoring, while nighttime performance is assessed using LiDAR-derived retroreflectivity following established calibration procedures. The framework integrates these components into a comprehensive Sign Condition Index (SCI) for maintenance guidance. Evaluation results demonstrated that LLaVA and Qwen outperformed InternVL, achieving bidirectional cosine similarity scores of 0.67-0.76 across all factors. Among 462 validated traffic signs, 68 were flagged by the proposed framework as requiring immediate replacement due to inadequate retroreflectivity performance. This research provides a cost-effective alternative to traditional manual inspections for comprehensive traffic sign condition assessment.
[72] Mitigating Domain Shift in Conditioned Floor Plan Generation: Synthetic Pre-training for Data-Efficient Adaptation cs.CVPDF
Matthieu Ospici, Arnaud Gueze, Luc Bourrat, Adrien Bernhardt
TL;DR: 本文研究了条件化户型图生成模型对领域偏移的鲁棒性问题,发现现有SOTA模型在跨数据集迁移时性能显著下降。为缓解此问题,作者提出了一种程序化方法生成大规模合成训练数据,该数据强制物理约束但牺牲建筑真实性。实验表明,合成数据预训练能显著提升零样本跨域性能,并为低数据场景下的微调提供高效初始化。
Details
Motivation: 解决户型图生成模型对领域偏移敏感的问题,因为实际应用中户型图因地区、文化和建筑规范差异而多样,而获取新标注数据成本高昂。
Result: 在三个公开数据集(RPLAN、MagicPlan和Swiss Dwellings)上评估,模型跨域迁移时性能下降高达一个数量级;合成数据预训练在MagicPlan上超越了域内训练,并在低数据微调场景中比真实数据初始化基线提升高达40%。
Insight: 创新点在于通过生成强制物理约束但非真实的合成数据来预训练模型,以增强领域泛化能力;客观分析表明,这种方法为数据高效的领域适应提供了新思路,尤其适用于标注稀缺的场景。
Abstract: Robustness to domain shift is a key requirement for floor plan generative models to be applicable beyond the single dataset they were trained on, as floor plans vary widely across regions due to distinct architectural cultures, spatial constraints, and construction practices, while acquiring new annotated datasets remains costly and domain-specific. Yet, no prior work has studied this robustness in the context of conditioned floor plan generation. In this paper, we evaluate state-of-the-art models from two fundamentally different generative paradigms across three public datasets (RPLAN, MagicPlan and Swiss Dwellings) and show that they are highly sensitive to domain shift, with up to an order of magnitude performance degradation when transferred across domains. To mitigate this with minimal target-domain supervision, we introduce a procedural method to generate a large-scale synthetic training dataset that enforces strict physical constraints (non-overlapping rooms, valid door placement, graph consistency) while intentionally sacrificing architectural realism through highly irregular spatial arrangements and aggressive geometric perturbation of room shapes. We show that pre-training on this synthetic data considerably improves zero-shot cross-domain performance, outperforming in-domain training on MagicPlan. Furthermore, it provides a highly effective initialization for fine-tuning, accelerating target domain adaptation and outperforming real-world initialization baselines by up to 40% in a low-data regime.
[73] AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models cs.CV | cs.AIPDF
Cong Su, Jiaju Han, Xuemeng Sun, Chengyin Hu, Qike Zhang
TL;DR: 本文提出了AirflowAttack,这是首个针对红外遥感视觉语言模型的对抗攻击方法,利用热气流湍流作为扰动先验。该方法通过轻量级生成器合成与输入无关的扰动,在多个CLIP骨干网络上实现了平均48.5%的攻击成功率,显著优于现有基线。研究还发现攻击会使某些模型将扰动误认为真实热证据,揭示了红外VLM生态系统中的关键漏洞。
Details
Motivation: 红外遥感视觉语言模型在安全关键场景中部署日益增多,但其对抗鲁棒性尚未得到充分研究。本文旨在探索如何利用热气流湍流这种物理扰动先验来攻击这类模型。
Result: 在五个不同的CLIP骨干网络上,该方法实现了平均48.5%的零样本场景分类攻击成功率,远超四个红外专用物理基线(27.7-37.0%)。在六个SOTA VLM上,场景分类准确率相对下降高达38.2%。
Insight: 创新点在于首次将热气流湍流作为对抗扰动先验武器化,并设计了输入无关的轻量级生成器。研究发现对抗扰动可能被模型误解释为真实热证据(如温度梯度和对流),这揭示了VLM在红外域中独特的脆弱性模式。
Abstract: Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior. A lightweight generator synthesizes a single input-agnostic perturbation regularized toward physically plausible airflow patterns. Optimized on one surrogate CLIP model, it attains a mean zero-shot scene-classification attack success rate (ASR, the fraction of samples whose top-1 class changes) of 48.5% across five diverse CLIP backbones, far exceeding four IR-specific physical baselines (27.7–37.0%). Applied to six state-of-the-art VLMs, it cuts scene-classification accuracy by up to 38.2% relative, yet paradoxically makes some models more confident in their IR analysis, confabulating the perturbation as genuine thermal evidence such as temperature gradients and convection. Ablations show the airflow prior raises physical plausibility at no measurable cost to attack success. Together with a benchmark spanning eleven models and four tasks, these findings expose critical vulnerabilities in the rapidly expanding IR VLM ecosystem.
[74] MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation cs.CVPDF
Jiaju Han, Ma Yaqi, Yahui Chai, Xuemeng Sun, Xin Li
TL;DR: 本文提出了MonoIR-RS,一个用于红外遥感视觉-语言学习的大规模数据集和基准。它通过红外感知的数据构建方法,结合CLIP风格的对比学习和视觉语言模型的指令微调,旨在解决红外遥感图像与语言理解任务中缺乏专门资源的问题。
Details
Motivation: 当前大多数遥感视觉-语言资源和模型主要关注可见光波段语义,导致红外视觉-语言理解研究不足。本文旨在填补这一空白,为红外遥感图像提供专门的语言对齐基准。
Result: 在AVIID基准测试中,合成的红外图像比灰度转换更接近真实热图像。对五个CLIP骨干和六个VLM骨干进行微调后,红外感知适应将CLIP平均召回率提升了高达12.8个百分点,并使VLM描述的红外线索覆盖率达到100%,同时将残留的RGB颜色泄漏降至接近零。
Insight: 创新点在于通过红外感知的数据构建和语言监督重写,将红外模态从RGB-IR双模态学习中分离出来,提供了一个可控、可复现的测试平台。这为红外遥感证据与语言对齐提供了新的方法,并显著提升了模型在红外任务上的性能。
Abstract: Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. We introduce MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions rewrite supervision around grayscale structure and infrared-style contrast instead of RGB appearance. We show that the synthesized infrared imagery is markedly closer to real thermal imagery than a grayscale conversion on the AVIID benchmark. We fine-tune five CLIP backbones and six VLM backbones, and calibrate them against zero-shot behavior: IR-aware adaptation lifts CLIP mean recall by up to 12.8 points and drives VLM captioning IR-cue coverage to 100% while reducing residual RGB-color leakage to near zero. By isolating the infrared modality from RGB-IR dual-modal learning, MonoIR-RS offers a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.
[75] CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cs.CVPDF
He Liang, Chenyang Ma, Yiming Zhang, Sangyun Shin, Andrew Markham
TL;DR: 本文提出了CAIRN,一种拓扑感知的大型多模态模型,用于多房间3D场景理解。该模型通过将Transformer注意力与场景层次结构对齐,并利用图神经网络丰富对象令牌的局部关系上下文,从而实现对包含多个互连房间的真实家庭环境的跨房间推理。
Details
Motivation: 现有的3D场景基础大语言模型主要关注简化的单房间3D场景中的问答任务,缺乏对包含多个互连房间和多样化对象类别的真实家庭环境进行推理的能力。
Result: 在作者提出的多房间3D场景理解基准CAIRN-MR(基于HM3D数据集构建)上,CAIRN在所有任务(包括定位、描述和四种问答任务)上都大幅超越了先前的3D-LLMs,同时在五个单房间基准测试中保持竞争力。
Insight: 核心创新点在于将场景的拓扑结构(对象级关系和房间级连通性)显式地整合到模型架构中,具体方法包括:使用图神经网络为对象令牌注入房间局部关系上下文、引入可学习的房间令牌进行房间级抽象、以及应用具有几何偏置的分层注意力掩码来根据场景拓扑路由信息。
Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We introduce CAIRN, a topology-aware 3D-LLM for multi-room 3D scene understanding. CAIRN aligns transformer attention with scene hierarchy, giving the model explicit awareness of object-level relations and room-level connectivity. It enriches object tokens with room-local relational context via a graph neural network, introduces learned room tokens for room-level abstraction, and applies a hierarchical attention mask with geometric bias to route information according to scene topology. CAIRN is developed on CAIRN-MR, a benchmark we introduce on HM3D for multi-room 3D scene understanding, covering grounding, captioning, and four question-answering tasks that progressively evaluate from intra-room perception to cross-room reasoning. Experiments show that CAIRN outperforms prior 3D-LLMs by a large margin across all CAIRN-MR tasks while remaining competitive on five single-room benchmarks.
[76] Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation cs.CVPDF
Songbur Wong, Xiaosong Jia, Junqi You, Bo Zhang, Pei Xu
TL;DR: 本文提出了’Point as Skeleton’框架,这是一个用于端到端自动驾驶评估的生成式传感器仿真框架。它通过自回归生成器,结合逐步更新的自车状态、交通参与者状态、场景地图以及点云骨架条件,来合成视觉观测。为了支持闭环交互,引入了Reset-and-Roll推理方法,并利用点云骨架来稳定自回归过程中的误差累积,从而实现了视觉保真度高且可交互的驾驶仿真。
Details
Motivation: 现有驾驶仿真方法(如CARLA和nuScenes)在闭环交互性和真实世界视觉保真度之间存在权衡,难以有效评估端到端自动驾驶系统。本文旨在解决这一挑战,开发一个既能支持闭环交互又具有高视觉真实感的仿真框架。
Result: 在nuScenes和nuPlan数据集上的实验表明,Point as Skeleton框架在闭环推演过程中显著提升了自回归生成的质量,证明了其在视觉保真的闭环驾驶仿真方面的潜力。
Insight: 主要创新点包括:1)将点云作为骨架,解耦前景与背景资产并投影为相机视图的绘制点和模板深度条件,为生成提供外观和几何线索以稳定误差累积;2)提出Reset-and-Roll方法,将滚动扩散推理适配到仿真中,防止未来条件潜在状态在仿真步骤间被固定,从而支持闭环推演;3)实现了一个基于nuPlan的渲染器级闭环生成接口,用于评估自车偏离原始日志时的生成效果。
Abstract: Evaluating end-to-end autonomous driving (E2E-AD) remains challenging, as existing driving simulation methods often trade off closed-loop interactivity (e.g., CARLA) and real-world visual fidelity (e.g., nuScenes). We present \textbf{\emph{Point as Skeleton}}, a generative sensor simulation framework for state-updated autoregressive driving video generation, in which an autoregressive generator synthesizes visual observations from step-wise updated ego states, actor states, scene maps, and point-cloud skeleton conditions. To support closed-loop rollout, we introduce Reset-and-Roll, which adapts rolling diffusion inference to simulation by preventing future-conditioned latent states from being committed across simulation steps. To stabilize error accumulation during step-wise autoregressive rollout, we introduce point-cloud skeletons that decouple foreground and background assets and project them into camera-view painted-point and template-depth conditions, providing appearance and geometric cues. We further implement a nuPlan-based renderer-level closed-loop generative interface for evaluating generation under ego deviations from the original log. Experiments on nuScenes and nuPlan show that \textit{Point as Skeleton} improves autoregressive generation quality during closed-loop rollout, demonstrating its potential for visually faithful closed-loop driving simulation. The code is available at https://github.com/krauwu/point-as-skeleton.
[77] ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation cs.CVPDF
Ruihang Zhang, Felix Taubner, Pooja Ravi, Kiriakos N. Kutulakos, David B. Lindell
TL;DR: 本文提出了ProxyPose,一种将单目视频中的6自由度姿态跟踪问题重新定义为视频到视频翻译任务的新方法。该方法仅需输入视频和第一帧中的一个标记像素,通过微调的视频扩散模型生成一个描绘彩色多面体(代理)的合成视频,该代理的运动与标记像素所在表面区域的局部刚体运动一致。由于代理的几何和外观已知,其完整的6自由度轨迹可通过现成的求解器进行经典姿态估计恢复。
Details
Motivation: 解决从单目视频中跟踪物体和表面的6自由度姿态这一长期存在的难题,现有方法需要3D模型、深度图、物体掩码或任务特定学习特征等额外输入,且在处理无纹理、透明、反光或可变形表面时存在困难。
Result: ProxyPose在6自由度姿态跟踪精度上达到了最先进水平(SOTA),且无需竞争方法所需的额外输入,其视频模型仅需在合成数据上进行微调。该方法还扩展到面部跟踪、相机姿态估计以及现有方法难以处理的野外场景。
Insight: 核心创新在于将姿态跟踪问题重构为视频到视频的翻译任务,利用大规模视频预训练模型来吸收处理挑战性材料、遮挡和变形等最困难方面,同时仅需像素级输入,无需对物体身份、边界或全局刚性做出假设。这种解耦方式将复杂的视觉理解问题转化为已知几何的代理姿态估计问题。
Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D models, depth maps, object masks, or task-specific learned features-and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input into a proxy video-a synthetic video depicting a colored polyhedron undergoing the same local rigid-body motion as the surface region at the marked pixel. Because the proxy’s geometry and appearance are known by construction, recovering its full 6-DoF trajectory reduces to classical pose estimation with off-the-shelf solvers. This formulation leverages large-scale video pre-training to absorb the hardest aspects of pose tracking-handling challenging materials, occlusions, and deformations-into the translation step, while operating at the pixel level with no assumptions about object identity, boundaries, or global rigidity. ProxyPose achieves state-of-the-art 6-DoF pose tracking accuracy without the additional inputs required by competing methods and after fine-tuning the video model only on synthetic data. We further demonstrate that ProxyPose extends to face tracking, camera pose estimation, and challenging in-the-wild scenes that are beyond the reach of existing approaches. Project page: https://ruihangzhang97.github.io/proxypose/.
[78] ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation cs.CV | cs.AI | cs.LGPDF
Tianjiao Yu, Xinzhuo Li, Yifan Shen, Onkar Susladkar, Yuanzhe Liu
TL;DR: ELSA3D是一个统一的3D基础模型,旨在通过弹性语义锚定技术,联合处理语言理解和3D几何生成。它通过引入锚定令牌和尺度感知的八叉树分词器,将文本语义与3D几何在不同抽象尺度上进行稀疏而精确的对齐,从而在单一骨干网络中实现3D生成与推理。
Details
Motivation: 现有统一3D模型将文本和3D令牌简单拼接为扁平序列,依赖自注意力机制,导致粗粒度结构线索和细粒度几何细节被混合成无差别的表示,无法有效进行跨模态交互。本文旨在解决文本与3D交互隐式且低效的问题。
Result: ELSA3D在图像到3D生成、文本到3D生成和3D描述任务上均达到了最先进的性能,超越了最强的统一基线模型。同时,相对于非弹性版本的相同模型,其FLOPs和推理延迟大约减少了一半。
Insight: 核心创新点是弹性语义锚定机制,通过稀疏的锚定令牌在不同几何尺度上选择性地路由语义线索并检索几何证据,实现了跨模态交互的稀疏化和精准化。轻量级路由器的设计使得计算和推理具有弹性,能够将跨模态能力集中在最需要对齐的地方,从而提升效率与性能。
Abstract: Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
[79] Vision as Unified Multimodal Generation cs.CVPDF
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi
TL;DR: 该论文提出将计算机视觉任务统一为多模态生成问题,通过一个统一的多模态模型在原生文本和图像生成空间中处理异构视觉任务,无需特定任务架构。SenseNova-Vision模型使用自然语言指令和可选视觉提示来指定任务、目标区域或视图,并生成文本、图像或混合输出以响应各种视觉任务。
Details
Motivation: 解决计算机视觉任务中异构架构和专用模型带来的碎片化问题,旨在通过统一的多模态生成框架整合多种视觉能力,为通用基础模型提供可扩展的集成路径。
Result: 实验表明,单个统一模型在结构化视觉理解、密集几何预测、分割和多视图视觉几何等任务上,能够匹配领先的任务专用系统性能,在多个基准测试中达到SOTA或相当水平。
Insight: 创新点在于将视觉任务统一表述为多模态生成,通过指令-响应范例和SenseNova-Vision语料库实现大规模训练,无需修改模型架构或添加任务特定预测头,支持语言定义的视觉变体任务组合。
Abstract: We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.
cs.GR [Back]
[80] SSA-3DGS: Unsupervised Removal of Screen-Space Artifacts for 3D Gaussian Splatting cs.GR | cs.CVPDF
Kristof Overdulve, Lode Jorissen, Nick Michiels
TL;DR: 本文提出了一种名为SSA-3DGS的无监督框架,用于解决3D高斯泼溅(3DGS)等新视角合成方法在处理带有屏幕空间伪影(如传感器缺陷、镜头污渍、水印等)的输入图像时,会将伪影错误地融入3D几何结构的问题。该方法通过联合优化3D场景和可学习的2D覆盖层,无需监督即可从3D场景几何中分离出静态伪影,从而恢复干净的3D场景并保留伪影本身。
Details
Motivation: 现实世界捕获的图像常包含固定在2D图像平面而非3D世界的屏幕空间伪影,这违反了3DGS等新视角合成方法对干净、多视角一致输入图像的假设,导致伪影被错误地烘焙为3D几何中的“漂浮物”或近相机伪影,从而降低新视角渲染质量。
Result: 在多种合成污染和自捕获的真实世界数据集上,SSA-3DGS相比在相同污染输入上训练的3DGS,将重建保真度提升了高达9 dB的PSNR,同时忠实地保留了污染伪影。
Insight: 论文的创新点在于提出了一种无监督的联合优化框架,利用跨视图的几何一致性,自动分离3D场景和2D伪影,无需人工标注或监督信号。从客观角度看,该方法为解决现实世界数据中普遍存在的屏幕空间污染问题提供了一种有效且通用的解决方案,提升了3DGS在非理想输入条件下的鲁棒性。
Abstract: Novel View Synthesis (NVS) methods, such as 3D Gaussian Splatting (3DGS), rely heavily on the assumption of clean, multi-view consistent, posed input images. Real-world captures can violate this assumption due to screen-space artifacts-static occlusions fixed to the 2D image plane rather than to the 3D world. Common examples include physical sensor defects, environmental obstructions (such as rain or mud on the lens enclosure), capture obstructions (such as a thumb over the camera sensor or a dashboard visible in dashcam footage), and digital overlays (such as watermarks or UI elements). When present, they are erroneously baked into the 3D geometry as “floaters” or near-camera artifacts, degrading the quality of novel-view rendering. In this work, we propose SSA-3DGS, an unsupervised framework that jointly optimizes a 3D scene and a learnable 2D overlay to recover a clean 3D scene and the corrupting artifacts. By exploiting geometric consensus across views, our method effectively disentangles static artifacts from the 3D scene geometry without supervision or manual input. Across diverse synthetic corruptions and a self-captured real-world dataset, SSA-3DGS improves reconstruction fidelity by up to 9 dB PSNR over 3DGS trained on the same corrupted inputs, while faithfully preserving the corrupting artifact.
cs.SE [Back]
[81] RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications cs.SE | cs.AI | cs.CLPDF
Evgeny Shilov
TL;DR: RuBench 1.0是一个针对代码代理的仓库级基准测试,包含25个从开源项目修复提交中提取的任务,其任务描述均以俄语原生撰写(非翻译),模拟真实用户请求风格。该基准旨在评估代码代理处理非英语任务规范的能力,并通过上游维护者的回归测试进行评判。
Details
Motivation: 现有仓库级代码代理基准测试的任务描述均为英语设计,无法衡量代理处理开发者以母语(如俄语)表述的真实维护请求的能力,因此需要创建原生非英语任务规范的基准。
Result: 在评估的部署产品配置中,最佳配置(Claude Code + Opus 4.8等)解决了78.7%的任务;但由于样本量较小(N=25),仅能统计分辨与最弱模型之间的差距。同时,审计发现产品在实际部署中可能静默替换模型(如20%任务被路由到备用模型),表明测量单元是产品而非单一模型。
Insight: 创新点在于首次构建了原生俄语任务规范的仓库级代码代理基准,强调真实语言使用场景;同时通过任务设计避免数据污染(所有提交晚于模型训练截止日期),并揭示了部署产品中模型替换的隐蔽行为,凸显了评估实际产品配置的重要性。
Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. Existing repository-level agentic benchmarks do not measure this setting: their task statements are English by design. We introduce RuBench 1.0, a benchmark of 25 tasks mined from recent fix commits in five live open-source repositories (aiohttp, aiogram, Laravel, NestJS, Fastify; Python, PHP, TypeScript, JavaScript), where each task is specified natively in Russian – written from scratch in the style of an actual customer request, not translated – and judged by the upstream maintainer’s regression tests, which we withhold from release. All 25 fix commits postdate the training-data cutoffs of every evaluated model, giving a contamination argument that holds task-by-task. We evaluate deployed product configurations (CLI agent + model + reasoning effort) – Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, and Codex CLI with GPT-5.5 – with three independent runs each, reporting pass@1 with task-level confidence intervals, paired comparisons, dollar cost, and token usage. The best configuration resolves 78.7% of tasks; at N=25 only the gaps to the weakest model are statistically resolvable, which we state explicitly. Auditing full trajectories of a fifth, hors-concours configuration (Claude Code + Fable 5, July 2, 2026 release), we caught the product silently substituting the model: on 5 of 25 tasks (20%) an official safeguard fallback re-routed routine HTTP-protocol fixes to Opus 4.8 – direct, reproducible evidence that the deployed product, not the model, is the unit actually measured. We release task statements, metadata, full agent trajectories, and diffs; grading oracles are withheld, with a SHA-256 manifest committed at publication time.
[82] UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation cs.SE | cs.AI | cs.CVPDF
Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng
TL;DR: 本文提出了UI2App基准测试,旨在评估视觉语言模型仅从UI截图推断交互行为的能力,重点关注可执行Web应用生成中的交互推理。该基准包含327张截图组成的45个状态一致截图集,并设计了一个端到端评估管道,从可执行性、导航可达性、视觉保真度和交互推理四个维度进行综合评估。
Details
Motivation: 现有文本驱动方法依赖复杂提示且布局表达有限,而图像驱动方法虽更贴近实际开发流程,但当前基准主要关注视觉保真度,缺乏对生成产物交互能力的系统评估。
Result: 在六个前沿视觉语言模型上的实验表明,视觉重建与交互实现能力存在显著不匹配:视觉保真度领先模型在交互推理指标(IIS)上仅得7.5分(满分100),排名第四,落后领先模型5.2倍;跨页面状态等高复杂度交互仍是普遍瓶颈,半数模型在该维度得分为零。
Insight: 创新点在于首次提出专注于交互推理的基准,其交互指标(IIS)通过功能正确性和状态管理复杂性评估推断交互,认可任何有效实现而非单一参考匹配;客观来看,该研究揭示了当前模型从静态截图推断完整交互行为仍面临关键挑战,为后续研究提供了系统性评估框架。
Abstract: Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence. Image-driven paradigms, which take UI screenshots as input, align more closely with real development workflows. However, current benchmarks focus primarily on visual fidelity and lack a systematic evaluation of the interaction capabilities in generated artifacts. To address this gap, we introduce UI2App, the first benchmark targeting interaction inference, the ability to recover application behavior from screenshots alone, without any textual or behavioral guidance. UI2App comprises 327 screenshots grouped into 45 state-coherent screenshot sets for runnable multi-route web applications. We design an end-to-end pipeline that evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses inferred interactions by functional correctness and state-management complexity, crediting any valid implementation rather than matching a single reference. Experiments on six frontier vision-language models reveal a marked capability mismatch between visual reconstruction and interaction realization: the visual-fidelity leader scores only 7.5 on IIS, ranking fourth and trailing the IIS leader by 5.2x. High-complexity interactions such as cross-page state remain a pervasive bottleneck, with half of the evaluated models scoring exactly zero on this dimension. Overall, the results indicate that inferring complete interaction behavior from static screenshots remains a key challenge for models.
cs.LG [Back]
[83] FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models cs.LG | cs.CVPDF
Bowen Xue, Zihan Min, Xingyang Li, Zhekai Zhang, Haocheng Xi
TL;DR: 本文提出了FourTune,一种用于扩散模型的高效后训练框架,采用端到端的W4A4G4(4-bit权重、激活和梯度)量化范式。它通过一个包含冻结数值稳定器的三分支混合流水线来稳定4-bit训练,并采用硬件高效的块级量化和定制融合内核来提升训练速度和减少内存占用。
Details
Motivation: 大型扩散模型的后训练面临巨大的内存占用和缓慢的训练速度挑战,现有的参数高效微调方法仅能部分解决这些问题。
Result: 在定制化、强化学习和蒸馏任务中,FourTune达到了全精度微调的质量。在FLUX.1-dev (12B)模型上,与BF16 LoRA相比,内存开销减少了2.25倍,端到端训练吞吐量提高了2.27倍。
Insight: 核心创新点在于提出了一个端到端的4-bit后训练框架,通过引入数值稳定器分支来隔离量化敏感异常值,从而在原生4-bit计算下实现稳定训练,并结合了硬件优化的量化与内核设计以提升效率。
Abstract: Diffusion models have become a dominant paradigm for high-quality generative modeling, while post-training is essential for adapting them to diverse downstream applications. However, post-training of large diffusion models is still challenging due to the prohibitive memory footprints and slow training speed, which existing parameter-efficient fine-tuning methods only partially address. To overcome these limitations, we propose FourTune, an efficient post-training framework for diffusion models based on an end-to-end W4A4G4 paradigm. FourTune introduces a triple-branch hybrid pipeline that augments the standard LoRA architecture with a frozen numerical stabilizer to isolate quantization-sensitive outliers, enabling stable training under native 4-bit computation. In addition, FourTune employs hardware-efficient block-wise quantization and customized fused kernels to support efficient quantized backpropagation and reduce memory bandwidth overhead. Across customization, reinforcement learning, and distillation tasks, FourTune matches the quality of full-precision fine-tuning. On FLUX.1-dev (12B), FourTune reduces memory overhead by 2.25$\times$ and increases end-to-end training throughput by 2.27$\times$ compared to BF16 LoRA.
eess.AS [Back]
[84] WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS eess.AS | cs.CL | cs.SDPDF
Sihang Nie, Jinxin Ji, Xiaofen Xing, Deyi Tuo, Chengbin Jin
TL;DR: 本文提出WordVoice框架,旨在解决基于大语言模型的文本转语音系统在细粒度控制方面的不足。通过构建包含五维词级标注的大规模双语数据集WordVoice-5A,并设计显式的声学规划机制与细粒度调制模块,实现了对时长、边界、能量、音高和音调等多个声学维度的精确解耦控制。
Details
Motivation: 当前基于LLM的TTS系统主要依赖隐式端到端生成,导致对词级声学属性的控制能力不足,难以满足有声读物、视频配音等需要精确风格干预和严格时间对齐的场景需求。
Result: 实验表明,WordVoice在保持竞争力的零样本合成稳定性的同时,实现了对多个声学维度的优越且解耦的控制性能。
Insight: 创新点包括构建大规模细粒度标注数据集WordVoice-5A,以及引入显式的边界标记机制进行声学规划,并结合细粒度声学调制模块来弥合离散标记与连续波形之间的分辨率差距,从而支持灵活的手动干预和多维属性对齐。
Abstract: While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate word-level acoustic attributes remains a critical bottleneck. This limitation is primarily amplified by the severe scarcity of fine-grained annotated datasets and the architectural challenge of integrating multi-dimensional control signals into discrete autoregressive generation. To address this, we propose a unified framework for highly precise word-level control. First, we construct WordVoice-5A, a massive 4.7k-hour bilingual dataset featuring five-dimensional word-level annotations (duration, boundary, energy, pitch and tone) developed through a rigorous linguistically-guided pipeline. Second, we introduce WordVoice to transform the implicit generation process into an explicit, highly controllable paradigm. Specifically, we introduce a bound-token mechanism within the LLM to formulate an explicit ``acoustic planning’’ process, enabling adaptive multi-task prosodic planning and flexible manual intervention. Furthermore, we augment the token-to-waveform stage with a fine-grained acoustic modulation module, bridging the resolution gap to strictly align word-level attributes between highly compressed discrete tokens and continuous waveforms. Extensive experiments demonstrate that WordVoice achieves superior, decoupled control over multiple acoustic dimensions while maintaining competitive zero-shot synthesis stability. The code and audio samples are publicly available at https://xxh333.github.io/wordvoice-demo/.
cs.RO [Back]
[85] GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory cs.RO | cs.CVPDF
Hu Zhu, Bohan Li, Xianda Guo, Hongsi Liu, Baorui Peng
TL;DR: 本文提出了HIOcc基准和GEM-Occ框架,旨在解决室内具身智能体长期跨空间语义占据建图的挑战。HIOcc统一了多个室内数据集,支持从局部预测到建筑级建图的评估。GEM-Occ则是一种高斯证据记忆框架,将局部视觉几何预测作为瞬时证据,融合到分层的持久化记忆中,以提升建图性能。
Details
Motivation: 现有室内语义占据研究主要关注单视图预测或房间级在线感知,缺乏对跨连接室内空间进行长期、大规模语义建图的探索。
Result: 在提出的HIOcc基准上进行实验,结果表明GEM-Occ在局部占据预测、在线地图稳定性、自由空间推理、重访一致性以及建筑级可扩展性方面,均优于先前的室内占据和基于高斯的建图基线方法。
Insight: 核心创新在于将局部几何预测视为需要融合的‘瞬时证据’,而非直接作为持久地图状态,并采用基于高斯和可见性/不确定性的因果更新机制。同时,提出了一个统一的多层次基准(HIOcc)来系统评估不同粒度的语义占据建图任务。
Abstract: Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantics. However, existing indoor occupancy benchmarks and methods mainly focus on single-view prediction or room-level online perception, leaving long-horizon semantic mapping across connected indoor spaces underexplored. We introduce HIOcc, a hierarchical indoor occupancy benchmark that unifies ScanNet, ScanNet++, and Matterport3D under a common sparse semantic occupancy format while preserving their native observation geometries, including perspective RGB-D frames and pano-centric observation groups. HIOcc supports three complementary evaluation regimes: local semantic occupancy prediction, room-level online occupancy mapping, and building-level mapping across connected panoramic environments. We further propose GEM-Occ, a Gaussian Evidence Memory framework for semantic occupancy mapping. Rather than using pointmaps as persistent map states, GEM-Occ treats local visual geometry predictions as transient evidence, converts them into semantic Gaussian occupancy evidence and free-space ray evidence, and fuses them into a persistent hierarchical memory through visibility- and uncertainty-aware causal updates. The memory is organized into local caches, room-level submaps, and a building-level graph, and can be queried at any time through Gaussian-to-occupancy splatting. Experiments on HIOcc show that GEM-Occ improves local occupancy prediction, online map stability, free-space reasoning, revisit consistency, and building-level scalability over prior indoor occupancy and Gaussian-based mapping baselines.
[86] FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning cs.RO | cs.AI | cs.CVPDF
Chuhao Zhou, Liquan Wang, Shuxin Cao, Xiangyu Chen, Yuxuan Hu
TL;DR: 论文提出FORGE方法,通过关键点轨迹推理实现功能性工具使用的泛化。该方法采用两阶段策略,将功能推理与动作执行解耦:首先从无动作数据中预测可泛化的关键点轨迹,然后利用少量演示将其映射到机器人动作。在七种工具的敲击功能基准测试中,FORGE在仿真和真实世界中对未见工具均优于现有方法,平均成功率提升超过2倍。
Details
Motivation: 解决机器人在使用特定工具训练后,无法将相同功能(如敲击)迁移到新工具上的问题,即功能性泛化差距。人类能基于视觉识别工具的共享功能意图,但机器人动作空间中的运动模式却因工具而异,需要弥合这一差距。
Result: 在七种工具的敲击功能基准测试中,FORGE在仿真和真实世界中对未见工具均优于最先进方法,平均成功率提升超过2倍,达到SOTA水平。
Insight: 创新点在于探索了包括可供性图像、人类视频提示和2D关键点轨迹在内的中间表示,发现关键点轨迹在功能表达性和动作可落地性之间取得最佳平衡,并据此设计了两阶段解耦策略。从客观角度看,该方法通过分离高层功能推理与低层动作执行,有效提升了工具使用的泛化能力。
Abstract: While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones – a gap we formalize as functional generalization. Such tools share a common functional intent that is visually recognizable, yet this perceptual similarity does not carry over to action space, where each tool demands an entirely different motor pattern. To bridge this gap, we explore intermediate representations including affordance images, human video prompts, and 2D keypoint trajectories, finding that keypoint trajectories best balance functional expressiveness and action groundability. Building on this, we propose FunctiOnal Reasoning and Grounded Execution (FORGE), a two-stage policy that decouples functional reasoning from action execution: predicting generalizable keypoint trajectories from action-free data, then grounding them into robot actions with limited demonstrations. On a seven-tool hitting-function benchmark, FORGE consistently outperforms state-of-the-art methods on unseen tools in both simulation and the real world, achieving over 2X improvement in average success rate.
[87] Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cs.RO | cs.CV | cs.LGPDF
Ryuji Oi, Hikari Otsuka, Kosuke Matsushima, Yuki Ichikawa, Masato Motomura
TL;DR: 本文提出了一种名为ActionCache的即插即用外部缓存方法,旨在解决基于流匹配的视觉-语言-动作(VLA)模型因迭代去噪过程导致推理延迟高的问题。该方法通过复用过去生成的中间动作,以目标动作附近的状态“热启动”新动作的生成,从而显著降低推理延迟。实验表明,该方法能保持高任务成功率,并对代表性VLA模型实现了高达11.75倍和34.43倍的推理加速。
Details
Motivation: 基于流匹配的VLA模型在机器人操作中表现出色,但其动作头中的迭代去噪过程是主要的计算瓶颈,阻碍了模型的实时部署。
Result: 在仿真和真实环境中的实验结果表明,ActionCache在低延迟状态下保持了高任务成功率,对代表性流匹配VLA模型π_{0.5}和GR00T-N1.6分别实现了高达11.75倍和34.43倍的推理加速。
Insight: 核心创新点是提出了一个无需训练、基于缓存的加速框架,通过存储和复用带有紧凑多模态键的中间动作,实现了跨任务和跨情景的相似上下文检索与动作生成预热,从而在不牺牲性能的前提下大幅提升推理速度。
Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, thereby drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes or even different tasks. Experimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving inference acceleration of up to $11.75\times$ and $34.43\times$ for representative flow-based VLA models, $π_{0.5}$ and GR00T-N1.6, respectively.
[88] WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation cs.RO | cs.CV | cs.GRPDF
Wongyun Yu, Youngwoon Kim, Minsu Cho
TL;DR: 论文提出了WristMimic,一个基于手腕引导的全身人形控制框架,用于将人类物体交互演示重定向到物理仿真中。该框架将无接触的身体运动与接触丰富的手部操作解耦,通过手腕的位姿目标引导身体和手腕,而手指则通过物体追踪和接触结果来学习抓取和操作行为。
Details
Motivation: 将人类物体交互演示重定向到基于物理的仿真中,不仅需要复现身体运动,还需要复现物体运动和接触,而仅依赖手部位置轨迹无法指定操作物体所需的接触力,直接跟踪这些轨迹会过度约束接触丰富的手指行为。
Result: 实验表明,WristMimic在性能上匹配或超越了使用完整手指位姿监督的方法,同时能够实现跨不同手部形态的手指不可知重定向。
Insight: 核心创新在于将手腕视为无接触身体运动与接触丰富手部操作之间的自然分界点,并提出了手腕特定的重置约束和奖励优先级机制来确保交互过程中手腕的可靠放置,从而实现了更灵活、更物理准确的手部操作学习。
Abstract: Retargeting human object interaction demonstrations to physics based simulation requires reproducing not only body motion but also the object motion and contacts that make manipulation succeed. However, position only hand trajectories do not specify the contact forces needed to manipulate objects, and directly tracking them can overconstrain contact rich finger behavior. We introduce WristMimic, a wrist guided whole body control framework that explicitly separates contact free body motion from contact rich hand manipulation. The contact free body and wrist are guided by kinematic pose targets, whereas the fingers are not directly supervised by human hand pose. Instead, they learn grasping and manipulation behaviors from object tracking and contact outcomes. Our key insight is that the wrist is the natural gate between these two regimes. It is largely free from contact and can be tracked kinematically, yet it determines the global hand configuration and places the fingers within reachable grasp affordances. To ensure reliable wrist placement during interaction, we introduce wrist specific reset constraints and reward prioritization. Experiments show that WristMimic matches or surpasses methods using full finger pose supervision while enabling finger agnostic retargeting across diverse hand embodiments.
[89] Learning to Throw Objects Safely in Multi-Obstacle Environments cs.RO | cs.CV | cs.LGPDF
Mohammadreza Kasaei, Klemen Voncina, Hamidreza Kasaei
TL;DR: 本文提出了一种用于多障碍物环境中安全投掷物体的机器人学习方法。通过引入势场状态表示,在固定尺寸网格上编码目标篮筐吸引力和障碍物排斥力,使强化学习策略能够泛化到任意数量和配置的障碍物场景。该方法在模拟中使用SAC等先进RL算法优化,并在真实机器人实验中实现了高达90%的成功率,展示了从模拟到现实的稳健迁移能力。
Details
Motivation: 现有机器人投掷方法(如TossingBot)假设无障碍环境,无法处理杂乱场景中的障碍物。本文旨在解决在随机放置障碍物的场景中,将物体安全投掷到目标篮筐内的问题。
Result: 在模拟实验中,势场表示相比显式状态编码获得了更高的成功率和更好的泛化能力;SAC算法在不同场景中表现最一致。真实机器人实验在未见过的可投掷物体和杂乱场景中实现了高达90%的成功率,证明了方法的有效性。
Insight: 创新点在于提出了紧凑的势场状态表示,统一编码目标吸引和障碍排斥,使策略能泛化到任意障碍配置;方法结合了示教初始化、模拟RL训练和sim-to-real迁移,为无结构环境中的安全投掷提供了实用解决方案。
Abstract: Robotic throwing enables fast and efficient object placement beyond the robot’s immediate workspace, but reliable throwing in cluttered environments remains underexplored. Existing approaches, such as TossingBot, learn throwing strategies from visual input but assume obstacle-free settings. In this paper, we address the problem of throwing objects into a target basket while avoiding obstacles placed randomly in the scene. We introduce a potential field state representation that compactly encodes both basket attraction and obstacle repulsion on a fixed-size grid, enabling reinforcement learning (RL) policies to generalize across arbitrary numbers and configurations of obstacles. The policy is initialized from kinesthetic demonstrations and optimized in simulation using three state-of-the-art RL algorithms (SAC, DDPG, TD3). Among these, SAC achieves the most consistent performance across scenarios. We compare the potential field representation against explicit state encodings and demonstrate that it achieves higher success rates and better scalability to unseen obstacle configurations. Real-robot experiments with unseen throwable objects confirm robust sim-to-real transfer, achieving up to $90%$ success in cluttered scenes. These results demonstrate that PFR provides a practical and robust representation for safe and efficient robotic throwing in unstructured environments. A video showcasing our experiments is available at: https://youtu.be/ZZnJf8ua2dE
[90] Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation cs.RO | cs.CVPDF
Jiaming Liu, Qingpo Wuwu, Nuowei Han, Hao Chen, Zhuoyang Liu
TL;DR: 本文提出了Lift3D-VLA,一个统一的视觉-语言-动作模型框架,旨在通过显式的3D点云推理和时序一致的动作生成来增强机器人操作能力。该方法基于改进的2D模型提升策略将3D点云与预训练的2D位置嵌入对齐,并引入了几何中心掩码自编码和分层时序动作建模,以同时捕获3D几何结构和物理动态。
Details
Motivation: 现有VLA模型在物理环境中的机器人操作缺乏几何理解和空间推理能力,且当前融合3D信息的方法受限于数据可用性和几何信息损失,无法在动态环境中联合捕获3D几何和时序结构化动作。
Result: 在22个模拟任务和8个真实世界操作任务中,Lift3D-VLA在MetaWorld和RLBench基准测试上分别比先前最佳VLA方法的平均成功率高出10.8%和11.1%,并在真实世界基准上超出最强基线4个百分点,同时展现出对分布外扰动的更强泛化能力。
Insight: 创新点在于将3D点云直接编码到VLA视觉编码器中以最小化空间信息损失,并通过双目标自监督框架(GC-MAE)联合重建当前点云并预测其未来几何演化,使2D编码器内化3D结构和物理动态;此外,分层时序动作建模利用LLM的多层协作预测动作块,实现了时序一致的预测。
Abstract: Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundamentally requires geometric understanding and spatial reasoning. While some VLA approaches attempt to incorporate 3D information, they are constrained by limited data availability and geometric information loss in current 3D encoding pipelines, and fail to jointly capture 3D geometry and temporally structured actions in dynamic environments. To address these limitations, we introduce Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and enables temporally coherent action generation. First, building upon our previous work Lift3D, an enhanced 2D model-lifting strategy is proposed to geometrically align 3D points with pretrained 2D positional embeddings. This design enables direct point-cloud encoding within the VLA vision encoder while minimizing spatial information loss. Based on explicit 3D inputs, we propose Geometry-Centric Masked Autoencoding (GC-MAE), a dual-objective self-supervised framework that reconstructs the current point cloud while predicting its future geometric evolution. This formulation allows the 2D vision encoder to internalize both 3D structure and physical dynamics. To fully exploit 3D representations, we further design layer-wise temporal action modeling, which leverages multiple layers of the LLM to collaboratively predict action chunks, enabling temporally consistent predictions. Across 22 simulated tasks and 8 real-world manipulation tasks, Lift3D-VLA achieves 10.8% and 11.1% higher mean success rates on MetaWorld and RLBench than the best-performing prior VLA methods, and outperforms the strongest real-world baseline by 4 percentage points, while exhibiting stronger generalization to out-of-distribution perturbations.
cs.AR [Back]
[91] BitFair: A 12nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision cs.AR | cs.CV | eess.IVPDF
Ang Li, Chang Gao
TL;DR: 本文提出BitFair,一种面向超低功耗XR视觉应用的12nm位串行CNN加速器,通过可学习的比特级提前终止和自适应比特排序实现软件-硬件协同设计,在满足亚毫秒延迟的同时显著提升能效。
Details
Motivation: 针对XR可穿戴设备在数瓦功耗和20毫秒延迟预算内实现始终在线感知的需求,现有位串行架构即使ReLU输出为零时仍处理所有比特,存在能效瓶颈。
Result: 在IBM DVS128 Gesture和N-MNIST数据集上分别达到96.5%和97.7%的准确率,相比已有XR视觉加速器能效提升4.0-22.1倍,准确率最高提升9.2%。
Insight: 创新点包括通过可学习的层间阈值实现动态比特级稀疏性利用,以及搜索层间比特顺序以优先处理信息量高的比特,在保证精度前提下最大化提前终止机会。
Abstract: Extended Reality (XR) wearables require always-on perception within tight power envelopes of a few watts and motion-to-photon latency budgets below 20 ms, leaving only a few milliseconds for neural-network inference. Bit-serial computing is attractive for such energy-efficient neural network acceleration, but many existing architectures still process all bits even when ReLU sets the final output to zero. This paper presents BitFair, a software-hardware co-designed bit-serial CNN accelerator with learnable bit-level early termination and adaptive bit ordering, working under the ultra-low-power and strict latency requirements of XR applications. BitFair exploits dynamic bit-level sparsity by learning per-layer thresholds that trigger early termination when partial sums reliably predict that the final ReLU output will be zero. Furthermore, it searches for layer-wise bit orders that prioritize informative bits, maximizing early termination without sacrificing accuracy. A GlobalFoundries 12nm FinFET implementation with a core area of 0.34 mm^2, 104 KB on-chip memory, and voltage scaling from 0.55 to 0.70 V achieves sub-millisecond latency, up to 117.0 BTOPS/W, and 0.07 pJ/SOP. On IBM DVS128 Gesture and N-MNIST, BitFair achieves 96.5% and 97.7% accuracy, respectively, while improving effective energy efficiency by 4.0-22.1x and accuracy by up to 9.2% over prior fabricated XR vision accelerators.
cs.IR [Back]
[92] CMDR: Contextual Multimodal Document Retrieval cs.IR | cs.AI | cs.CL | cs.CVPDF
Ryota Tanaka, Taku Hasegawa, Kyosuke Nishida
TL;DR: 本文提出了CMDR(上下文多模态文档检索)任务及对应的CMDR-Bench基准,旨在解决现有方法忽略文档跨页上下文信息的问题。作者提出了CMDR-Embed框架,通过联合编码多页内容并从中提取页面级嵌入来显式建模文档上下文,并设计了CMCL对比学习目标进行训练。实验表明该方法显著优于非上下文嵌入方法。
Details
Motivation: 现有多模态文档检索基准主要评估简单的词汇或语义匹配,且方法通常独立编码页面,忽略了解决跨页信息聚合查询所需的文档上下文信息。
Result: 实验证明CMDR-Embed在提出的CMDR-Bench基准上显著优于非上下文嵌入方法,突显了上下文感知多模态嵌入对推进文档检索的重要性。
Insight: 创新点在于提出了一个需要建模文档上下文的新检索任务与基准,并设计了能显式纳入文档上下文的联合编码框架及平衡上下文建模与页面区分性的对比学习目标。
Abstract: Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document. However, existing benchmarks primarily evaluate simple lexical or semantic matching, and most methods encode pages independently. Consequently, they overlook the contextual information in the document required to resolve queries that aggregate information across multiple pages. In this paper, we introduce CMDR and CMDR-Bench, a new multimodal document retrieval task and benchmark that require modeling document context. To address this challenge, we propose CMDR-Embed, a contextual multimodal embedding framework that explicitly incorporates document context by jointly encoding multiple pages and deriving page-level embeddings from a shared contextual representation. Furthermore, we introduce CMCL, a contextual multimodal contrastive learning objective that effectively trains CMDR-Embed by balancing contextual modeling with page-level discriminability. Experiments demonstrate that CMDR-Embed significantly outperforms non-contextual embeddings, highlighting the importance of context-aware multimodal embeddings for advancing document retrieval.
eess.IV [Back]
[93] TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring eess.IV | cs.CVPDF
Fadi Abdeladhim Zidi, Salah Eddine Bekhouche, Abdellah Zakaria Sellam, Gaby Maroun, Fadi Dornaika
TL;DR: 本文提出了一种名为TMF-RSE的三模态深度学习框架,用于从胸部影像中量化肺部疾病严重程度。该框架融合了二维胸部影像的外观特征、肺部分割掩码的结构特征以及视觉语言模型(VLM)的语义特征,并采用证据回归来同时提供严重程度预测和不确定性估计。
Details
Motivation: 从胸部影像中准确量化肺部疾病严重程度对于临床决策和资源分配至关重要。现有方法可能未能充分利用多模态信息的互补性,因此需要一种能有效融合外观、结构和语义信息的框架。
Result: 在Per-COVID-19 CT和RALO数据集上的实验表明,TMF-RSE超越了近期基于Transformer的基线模型。具体而言,在Per-COVID-19验证集上取得了4.02的MAE和0.9629的皮尔逊相关系数,在RALO地理范围任务上取得了0.339的MAE和0.973的皮尔逊相关系数。
Insight: 论文宣称的创新点在于提出了一个融合外观、结构和语义三种模态信息的互补融合机制,并引入了证据回归来量化预测的不确定性。从客观角度看,将视觉语言模型的语义信息与医学影像的特定结构先验结合,是一种有前景的多模态融合策略。
Abstract: Accurate quantification of lung disease severity from chest imaging is critical for clinical decision-making and resource allocation. We propose a tri-modal deep learning framework, TMF-RSE (Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty), that combines appearance features from two-dimensional chest inputs, structural features from lung segmentation masks, and semantic features from vision-language models (VLMs) for severity quantification. Our approach employs complementary fusion mechanisms that integrate semantic guidance, structural priors, and hierarchical interactions across modalities. The model employs evidential regression to provide both severity predictions and uncertainty estimates. Experiments on the Per-COVID-19 CT and RALO datasets show that TMF-RSE outperforms recent transformer-based baselines, achieving MAE of 4.02 and Pearson correlation of 0.9629 on Per-COVID-19 validation, and 0.339 MAE / 0.973 PC on RALO geographic extent.
cs.CR [Back]
[94] Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection cs.CR | cs.CV | cs.CYPDF
Junade Ali
TL;DR: 本文提出了一种基于溯因推理的概率性AI模型验证方法,用于法证合成媒体检测。该方法通过整合多种检测方法的输出,从事实矩阵中识别最可能的结论,从而显著降低误报风险,并首次对OpenAI的SynthID在合成图像上的表现进行了实证评估。
Details
Motivation: 针对AI模型基于概率行为的归纳推理在现实应用中可能不可行的问题,本文探索利用溯因推理来验证多个检测方法的结果,以在法证合成媒体检测中提高结论的可靠性。
Result: 该方法在合成媒体检测中能够不成比例地降低误报相对于真阳性召回的风险;同时,首次实证评估了OpenAI SynthID在合成图像上的表现,并评估了不同合成媒体检测方法的互补性。
Insight: 创新点在于将溯因推理引入概率性AI模型的验证过程,通过多方法结果佐证来提高法证检测的鲁棒性;客观来看,该方法为合成媒体检测提供了一种可解释性更强的融合策略,并填补了SynthID实证评估的空白。
Abstract: Artificial Intelligence (AI) models, at their core, apply general learnings from broad datasets to individual circumstances using probabilistic behaviour. This inductive approach stands in contrast to deductive reasoning approaches which seek to prove conclusions from their premises. However, research has shown that deductive reasoning with AI models is a challenging problem and in the real-world it may not always be feasible. An alternative way forward is to leverage abductive reasoning, seeking to corroborate the output of multiple approaches to identify the most likely conclusion from the factual matrix. We apply this to synthetic media detection in forensic settings, and find we are able to disproportionately lower the risk of false positives to true positive recall. We also provide the first empirical evaluation of OpenAI’s rollout of SynthID on synthetic images and evaluate how complementary different synthetic media detection approaches are.
cs.AI [Back]
[95] TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training cs.AI | cs.CLPDF
Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu
TL;DR: 本文提出了TurnOPD,一种用于长视野智能体高效在线策略蒸馏的轮次级预算策略。它通过自适应展开深度预算和渐进式轮次归一化损失预算,解决了传统智能体在线策略蒸馏中存在的资源浪费和深层决策轮次训练不足的问题。在ALFWorld、WebShop和Multi-Hop Search等基准测试中,TurnOPD在同等训练时间预算下取得了更优的验证准确率。
Details
Motivation: 在线策略蒸馏为语言智能体训练提供了一个有前景的框架,但其在长视野智能体任务中的应用探索不足。作者发现传统方法存在两个关键低效问题:全视野展开在尾部轮次上浪费资源,以及轨迹级KL目标导致深层决策轮次训练不足。
Result: 在ALFWorld、WebShop和Multi-Hop Search基准上,使用任务专用教师模型进行实验。结果表明,在相等的训练时间预算下,TurnOPD实现了更优的验证准确率,并将准确率-时间前沿推向了超越传统在线策略蒸馏的水平。
Insight: 核心创新在于将预算控制从轨迹级细化到轮次级,提出了自适应展开深度预算和渐进式轮次归一化损失预算两个控制器。这为长视野任务的高效蒸馏提供了新思路,即通过动态调整资源分配和损失权重来优化训练效率与效果。
Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student’s own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy–time frontier beyond vanilla OPD.
[96] PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cs.AI | cs.CLPDF
Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou
TL;DR: 本文提出了PolyWorkBench,一个用于评估多语言长视野LLM智能体在工作流任务中性能的基准测试。该基准包含五个领域的67个任务,要求智能体处理多语言输入、进行迭代推理、调用外部工具并生成结构化输出。研究还设计了一个结合结构评分、可执行验证和基于LLM语义评估的混合评估框架。
Details
Motivation: 现有基准大多隐含单语言假设,而现实应用常涉及统一工作流中的多语言输入输出,多语言性与智能体执行之间的交互关系尚未被充分探索。
Result: 实验结果表明,与单语言设置相比,最先进的LLM智能体在多语言工作流设置中性能显著下降。分析表明多语言性在推理和执行步骤中引入了复合效应。
Insight: 创新点在于构建了首个专注于多语言长视野工作流的基准,并提出了混合评估框架以同时捕捉功能正确性和跨复杂工作流的语言一致性。客观来看,该研究强调了在智能体评估中联合建模语言变异和程序决策的重要性。
Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs. To enable comprehensive evaluation, we propose a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment. This design allows us to capture both functional correctness and linguistic consistency across complex workflows. Empirical results show that state-of-the-art LLM agents suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts. Our analysis suggests that multilinguality introduces compounding effects across reasoning and execution steps, highlighting the importance of jointly modeling language variation and procedural decision-making in agent evaluation.
[97] Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory cs.AI | cs.CL | cs.MAPDF
Jihao Liu, Guoxiong Gao, Zeming Sun, Bin Wu, Shurui Liu
TL;DR: 本文提出了Danus,一个面向研究级数学推理的编排系统,其核心是使用共享事实图作为全局内存管理机制。该系统由一个执行规划和协调的主代理、多个并行执行证明搜索的工作代理,以及一个无状态验证器组成,旨在协调并行证明搜索并组织可靠的中间断言。
Details
Motivation: 当前基于LLM的数学推理代理在处理研究级问题时,难以有效协调并行证明搜索并保持中间断言的组织性和可靠性,因此需要一种有效的编排系统来解决这一挑战。
Result: 通过在代数几何、奇点理论和组合数学中的六个研究级案例研究进行评估,结果表明基于事实图的编排机制使Danus能够构建长而详细的数学证明,为长视野研究问题的数学推理代理扩展提供了有效途径。
Insight: 创新点在于引入共享事实图作为全局内存,结合主代理的规划协调、工作代理的并行搜索以及验证器的断言检查,实现了增量式构建长论证并保持证明状态的组织性,这为复杂数学问题的自动化推理提供了可扩展的架构思路。
Abstract: Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the resolution of open problems. However, scaling and orchestrating such agents effectively remains challenging, due to the difficulty of coordinating parallel proof search while keeping intermediate claims organized and reliable. In this paper, we propose Danus, an orchestration system for research-level mathematical reasoning centered on a shared fact graph as a global memory-management mechanism. Danus consists of a main agent that performs planning and coordination, multiple worker agents that carry out proof search in parallel, and a stateless verifier that checks proposed mathematical claims before they are admitted into the fact graph. Each verified fact is stored together with its proof and logical dependencies, allowing the system to build long arguments incrementally while keeping the shared proof state organized. The main agent periodically summarizes the evolving proof state, redirects workers across promising directions, and supports interaction with human mathematicians through progress reports. We evaluate Danus through six research-level case studies in algebraic geometry, singularity theory, and combinatorics, illustrating how the fact-graph memory mechanism enables Danus to construct long, detailed mathematical proofs. Our results suggest that fact-graph-based orchestration provides an effective route toward scaling mathematical reasoning agents for long-horizon research problems. Danus is open source at https://github.com/frenzymath/Danus.
[98] Rethinking Indic AI from a Lens of Cultural Heritage Preservation cs.AI | cs.CLPDF
Aparna Madva, Sharath Srivatsa, Srinath Srinivasa, Tulika Saha
TL;DR: 本文从文化遗产保护的角度重新审视印度次大陆的人工智能发展,探讨AI对印度语言和文化的影响。论文分析了印度语言的结构与社会语言学特征,回顾了印度自然语言处理(NLP)的历史演变,并讨论了新兴的印度基础模型如何解决资源与表征差距。最后,提出了基于解释学推理的’文化感知’研究新方向,旨在促进更稳健和包容的印度AI模型发展。
Details
Motivation: 研究动机在于探讨AI在印度次大陆的’双刃剑’效应:一方面AI可促进包容性,另一方面可能同质化世界观并边缘化少数语言与文化。论文旨在系统分析印度语言文化的独特性及其对AI基础模型构建带来的挑战。
Result: 论文未提供具体的定量实验结果,但通过纵向调查梳理了印度NLP领域的关键里程碑、方法转变和资源建设进展,并分析了当前印度基础模型在解决资源与表征差距方面的作用。
Insight: 创新点在于从文化遗产保护视角系统审视印度AI发展,并提出’文化感知’这一新研究方向,强调基于解释学推理构建能理解文化语境、确保跨低资源语言公平性能的AI模型,为包容性基础模型设计提供了新思路。
Abstract: As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization. AI is seen as a ‘’double-edged sword’’ where on the one hand, it can enable access and inclusion for a large population, on the other, it can homogenize worldviews and exclude underrepresented languages and worldviews. In this paper, we try to characterize this problem by addressing the extensive characteristic nature of Indian linguistics and the way they closely connect to cultural practices and worldview. We then perform a longitudinal survey of how Natural Language Processing (NLP) techniques have evolved in this space, tracing the historical development of Indic NLP, covering key milestones, methodological shifts, and resource creation efforts. In addition, the paper also examines the structural and sociolinguistic characteristics of Indian languages, such as rich morphology, complex scripts and grammar rules, diglossia, and large dialectal variation, and explains how these create unique challenges for building AI foundation models. We then discuss the growing role of Indic foundation models and analyze how these models address these long-standing resource and representation gaps. Finally, we propose a research direction called ‘Culture Sensing’, which re-imagines AI based on hermeneutic reasoning. Culture Sensing aims to address open problems such as ensuring equitable performance across low-resource languages and producing outputs that are culturally meaningful. By bringing together past work, current techniques, and emerging trends, this paper outlines research directions that can guide the next phase of Indic NLP and contribute to the development of more robust and inclusive Indic foundation models.
[99] Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment cs.AI | cs.CVPDF
Han-Jun Ko, Jr-Jen Chen, Haobo Yuan, Hsin-Ying Lee, Tiancheng Shen
TL;DR: 本文提出VAORA(视觉动作结果推理对齐)方法,通过设计两种互补的奖励机制来解决视觉语言模型在交互式物理推理中的泛化问题:视觉对齐奖励将推理锚定于视觉上下文,视觉-动作对齐奖励将推理与模型动作引发的视觉结果关联。该方法在PHYRE和Virtual Tool基准测试中验证了其在新任务和未见环境下的有效性。
Details
Motivation: 针对视觉语言模型在交互式物理推理中难以泛化到未见任务和环境的问题,主要解决模型推理与物理现实矛盾(幻觉链式推理)以及推理与动作之间的错位这两个关键失败模式。
Result: 在PHYRE和Virtual Tool基准测试中,VAORA在新任务和未见环境设置下均表现出色,证实了通过该方法可以诱导出具有基础性和泛化性的物理智能。
Insight: 创新点在于提出两种互补的奖励设计来直接抑制幻觉推理并缩小推理与行为间的差距,同时通过预训练领域专家代理估计成功概率来提供平滑密集的奖励,以提升训练稳定性。从客观角度看,该方法将物理推理与任务泛化通过视觉动作结果的对齐机制进行桥接,为提升模型在物理交互场景中的鲁棒性提供了新思路。
Abstract: Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model’s reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model’s action. Together, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.