Table of Contents
- cs.CL [Total: 17]
- cs.CV [Total: 55]
- cs.AI [Total: 2]
- cs.LG [Total: 2]
- cs.CY [Total: 1]
- eess.IV [Total: 2]
- q-bio.NC [Total: 1]
- cs.GR [Total: 1]
- cs.RO [Total: 4]
cs.CL [Back]
[1] AgentGUI: An Interface for Observing and Steering Long-Running AI Agents cs.CL | cs.AI | cs.HCPDF
Xuan Zhao, Jiwoong Sohn, Qinyue Zheng, Michael Moor
TL;DR: 本文介绍了AgentGUI,一个用于观察和引导长时间运行AI智能体的本地图形用户界面。该界面旨在解决人类监督滞后于智能体自主能力发展的问题,提供丰富的轨迹可视化、手动与自动引导功能,并能集成协调开源与前沿智能体框架。用户研究表明,使用AgentGUI能显著缩短从智能体轨迹中识别关键元素的时间(快38%),其自动漂移预防功能在小型本地智能体上最高可将任务完成率提升34个百分点。
Details
Motivation: 随着AI智能体处理复杂、长时间任务的能力日益增强,由于缺乏以人为中心的交互界面,人类监督能力已系统性地滞后。本文旨在通过开发一个用户友好的本地GUI来解决这一问题,实现对多个并发、长时间运行智能体会话的无缝观察和引导。
Result: 一项受控用户研究显示,使用AgentGUI从智能体轨迹中识别关键元素的时间有统计学上的显著减少(快38%,p=0.023)。在一项初步实验中,其自动漂移预防功能在0.8B到9B参数规模的模型阶梯上,将小型本地智能体的任务完成率最高提升了34个百分点(每个模型N=50次运行)。
Insight: 论文宣称的创新点在于提供了一个集成的、本地托管的GUI,用于对长时间运行的AI智能体进行观察和引导,并支持多种框架的集成与协调。从客观角度看,其将轨迹可视化、手动/自动干预以及跨框架协调整合到一个统一界面中的设计,为解决智能体可观察性和可控性问题提供了实用的工具创新。
Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI’s automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B–9B model ladder (N=50 runs per model). AgentGUI is publicly available through its project website (https://agent-gui-project.github.io) and open-source repository (https://github.com/eth-medical-ai-lab/agent-gui), along with a demo video (https://youtube.com/watch?v=GSDyxN1gTF0).
[2] Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs cs.CLPDF
Farhan Farsi, Shayan Bali, Mohammad Heydari Rad, Negar Heidary, Donya Rooein
TL;DR: 该论文通过构建Symphony-Bias多模态数据集,评估了十种不同架构和规模的多模态大语言模型在文本、视觉和音频三种模态下对22种乐器的性别关联偏见。研究发现,92%的乐器层面结果与社会科学研究一致,其中竖琴和鼓的性别关联在所有模型和模态中尤为一致,且文本模态对性别刻板印象的强化作用最强,音频最弱。
Details
Motivation: 鉴于大语言模型日益融入日常生活并用于信息检索,研究其可能延续社会偏见和强化刻板印象的风险至关重要。本文从乐器文化性别类型化的社会科学研究出发,旨在探索多模态大语言模型中的性别偏见。
Result: 在Symphony-Bias数据集上的评估显示,模型对乐器的性别关联与社会科学发现高度一致(92%),其中竖琴和鼓的性别关联在所有模型和模态中表现最一致。研究进一步发现,与刻板印象的吻合度在音频模态中最弱,视觉模态较强,文本模态最强。
Insight: 论文的创新点在于构建了首个跨文本、视觉和音频的并行多模态偏见数据集Symphony-Bias,并系统评估了多模态模型在不同模态下性别偏见的差异,揭示了模态特异性表征会不同程度地放大对乐器的性别关联,为理解和缓解多模态AI偏见提供了新视角。
Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments. Building on social-science research on the cultural gender-typing of instruments, we introduce Symphony-Bias, a parallel multimodal dataset spanning text, vision, and audio. We evaluate ten multimodal models with diverse architectures and scales across 22 musical instruments, analyzing how they associate each instrument with three gender categories: {male, female, non-binary}, across three modalities: {text, vision, audio}. Our results show that 92% of instrument-level outcomes align with prior social-science findings, with the harp and drums showing particularly consistent gendered associations across all evaluated models and modalities. We further find that alignment with social stereotypes is weakest in audio, stronger in vision, and strongest in text, suggesting that modality-specific representations can differentially amplify gendered associations with musical instruments.\footnote{The Symphony-Bias dataset will be publicly released upon acceptance of the paper.}
[3] Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text cs.CL | cs.AIPDF
Aman Kumar, Lasitha Vidyaratne, Dipanjan D Ghosh, Arnab Chakrabarti, Ahmed K Farahat
TL;DR: 本文研究了财务披露文本中的细粒度不一致性分类问题,将冲突类型分为数值、时间、指代、事实和逻辑等类别。研究在SBID-FD合成基准上系统比较了冻结嵌入分类器、微调编码器、证据增强分类器、提示大语言模型和LoRA适配生成模型等多种方法。
Details
Motivation: 财务披露文本包含多种可能冲突的声明,仅检测冲突是不够的,审查工作流还需要确定其具体类型,因为不同类型的不一致性需要不同的证据和下游检查。
Result: 在SBID-FD基准的5940个实例上,微调的3亿参数编码器达到61.9%准确率,略高于LoRA适配的Qwen3.5-9B模型(61.5%)和GPT-5.4(61.3%)。提供黄金证据跨度可将微调编码器提升至65.3%。
Insight: 研究表明,紧凑的监督编码器在特定任务上可以达到与大语言模型相当的性能,体现了效率优势;同时,证据定位质量是性能提升的关键瓶颈,特别是对指代不一致性影响显著,而事实和逻辑不一致性即使有相关证据仍难以处理。
Abstract: Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks. We study this problem as fine-grained inconsistency classification. Using a fixed 5,940-instance snapshot of SBID-FD, a synthetic financial-disclosure benchmark with 11 inconsistency labels and paired reference evidence spans, we compare frozen embedding classifiers, fine-tuned encoders, evidence-augmented classifiers, prompted large language models, and LoRA-adapted generative models under a shared evaluation protocol. A fine-tuned 300M encoder reaches 61.9% accuracy, compared with 61.5% for a LoRA-adapted Qwen3.5-9B model and 61.3% for GPT-5.4. Because these systems differ in architecture, supervision, training objective, and input format, we interpret this as a practical efficiency result for compact supervised encoders rather than a controlled conclusion about model scale. Supplying gold evidence spans improves the fine-tuned encoder to 65.3%, whereas automatically predicted spans recover a meaningful but incomplete share of that gain, indicating that localization quality remains a bottleneck. Class-level analyses show that Referential inconsistencies are especially sensitive to localization quality, while Factual and Logical inconsistencies remain difficult even when the relevant evidence is provided. Together, the oracle, distractor, and per-class analyses separate localization errors from residual type-discrimination errors, indicating that progress requires both stronger evidence extraction and better reasoning over closely related inconsistency categories.
[4] Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification cs.CL | cs.LG | q-bio.QMPDF
Linyu Li, Zhi Jin, Yichi Zhang, Dongming Jin, Yuanpeng He
TL;DR: 本文提出EC-Reason-Bench,一个无需训练的诊断性基准,用于探究通用大语言模型在酶分类任务上表现不佳的原因。研究发现,外部知识是关键,必须在推理前提供;在闭卷设置下,推理策略的效果取决于模型的弃答倾向;在开卷设置中,最佳LLM的聚合分数与简单投票检索邻居的EC号相当,但推理主要作用是仲裁冲突证据而非提供知识;准确度遵循同源性可用性定律。
Details
Motivation: 解决通用大语言模型在酶功能预测(一种层次化、知识密集型的蛋白质功能分类)中存在的异常现象:虽然能正确预测粗粒度的一级分类,但在预测完整EC号(二至四级)时准确率几乎为零,而专用模型和工具仍保持可用。旨在诊断为何通用LLM在此任务上得分接近零,以及在不更新模型权重的情况下能恢复多少性能。
Result: 在多个强推理LLM上的实验表明:在开卷访问外部知识后,性能显著提升,缩小了模型间的差距;闭卷设置中,级联和思维链推理的效果取决于模型的弃答倾向;最佳LLM设置与简单投票检索邻居EC号的聚合分数相当,但平均分数掩盖了在对抗性证据集上的巨大增益和多功能酶上的巨大损失。
Insight: 创新点在于将酶分类能力分解为四个正交杠杆(输出结构、外部知识、推理结构、推理鲁棒性)进行独立测量,并提出了一个无需训练的诊断评估协议。核心洞察是:外部知识对酶分类至关重要且必须置于推理之前;推理在开卷设置中主要充当冲突证据的仲裁者而非知识来源;单一数字的排行榜无法全面反映模型性能,需考虑同源性可用性等深层因素。
Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model’s tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.
[5] ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models cs.CL | cs.AIPDF
Ruxi Gu, Zhenliang Zhang, Wei Wang
TL;DR: 本文提出了ForgetBench基准,用于系统评估大型语言模型在持续知识编辑下的遗忘动态。该基准包含概念问答和场景问答两种范式,通过序列化编辑框架构建时序知识流,并引入统一评估框架量化长期保留动态。实验表明现有方法难以平衡长期保留与泛化质量,凸显了未来LLM需要更鲁棒的记忆机制。
Details
Motivation: 现有评估范式主要关注单步推理或静态知识编辑,无法捕捉持续模型修改中知识保留与退化的时序动态,因此需要系统评估LLMs在重复更新下的知识遗忘行为。
Result: 在多种模型和编辑方法上的广泛实验表明,现有方法无法平衡长期保留与泛化质量,具体结果通过ForgetBench的时序衰减、保留强度和跨实例稳定性等指标量化。
Insight: 创新点在于提出首个系统评估LLM持续知识编辑下遗忘动态的基准,通过概念与场景双范式解耦事实保留与关系知识保存,并引入建模知识演化的统一评估框架,为未来鲁棒记忆机制设计提供量化分析基础。
Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification. In this work, we propose ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing. ForgetBench introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation. Building upon a sequential editing framework, we construct temporally ordered knowledge streams and evaluate model behavior across multiple editing stages. To quantitatively analyze long-term retention dynamics, we further introduce a unified evaluation framework that models knowledge evolution over time, enabling the measurement of temporal decay, retention strength, and cross-instance stability. Extensive experiments across diverse models and editing methods demonstrate that existing approaches fail to strike a balance between long-term retention and generalization quality. Our findings highlight the need for more robust memory mechanisms that can effectively acquire, update, and preserve knowledge over time in future LLMs. Code will be released upon acceptance.
[6] CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG cs.CL | cs.IRPDF
Lang Zhou, Yingjian Chen, Shuxuan Li, Kun-Yu Lin, Zhilin Zhao
TL;DR: 本文提出了CMT-RAG框架,用于解决多轮多跳检索增强生成(RAG)任务中的推理和依赖追踪问题。该框架通过将对话上下文表示为子问题级别的推理轨迹,并构建会话级有向无环图(DAG)来存储持久记忆,从而提升后续查询中相关推理和证据的恢复效率。
Details
Motivation: 现有RAG系统通常将对话记忆表示为原始对话历史、重写查询或非结构化摘要,难以恢复后续查询所需的具体先前推理步骤和证据,因此需要一种能对齐检索的对话记忆表示方法。
Result: 在提出的MuMu-QA基准和语料级RAG基准上的实验表明,CMT-RAG在答案准确性上持续优于五类RAG基线方法,达到了SOTA水平。
Insight: 创新点在于将对话记忆与检索对齐,通过结构化推理轨迹(包含面向检索的子问题和依赖关系)来组织会话级DAG记忆,这为多轮多跳RAG提供了可解释且高效的记忆管理机制。
Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history, rewritten queries, or unstructured summaries, making it difficult to recover the specific prior reasoning steps and evidence required for follow-up queries. Our key insight is to align conversational memory with retrieval by representing dialogue context as sub-question-level reasoning traces. Building on this insight, we introduce MuMu-QA, a benchmark for multi-turn multi-hop RAG with explicit cross-turn sub-question dependency annotations, and CMT-RAG, a complementary memory framework for this setting. At each turn, CMT-RAG employs a state-space trace generator, whose recurrent state serves as runtime memory, to incorporate recent conversational context and decompose the current query into structured trace drafts containing retrieval-oriented sub-questions and dependencies on earlier traces. It then grounds these drafts with retrieved evidence and stores them as persistent memory traces in a session-level DAG, enabling future turns to efficiently recover relevant prior reasoning and evidence. Experiments on MuMu-QA and corpus-level RAG benchmarks show that CMT-RAG consistently outperforms five categories of RAG baselines in answer accuracy.
[7] Mergeable Model-Side Aggregation States for Long-Context Language Models cs.CL | cs.AIPDF
Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
TL;DR: 本文针对长上下文语言模型在集合聚合任务中性能随上下文增长而下降的问题,提出了一种模型侧聚合接口。该方法通过维护基于哈希的超对数草图状态,在冻结的语言模型旁处理聚合信息,支持跨上下文段合并和直接读取,避免了额外的生成-执行-返回循环。实验表明,该方法在百万级记录的基数估计中平均相对误差为1.6%,并在聚合推理任务中显著提升了模型性能。
Details
Motivation: 解决长上下文语言模型在非加性、基于集合的聚合任务中性能不可靠的问题,这类任务广泛存在于日志、程序输出、表格和多轮对话中。
Result: 在包含100万条记录的基数估计实验中,平均相对误差为1.6%;在聚合推理任务上,相比直接全上下文推理,在Qwen和Gemma模型上分别提升了63.2和56.3个百分点;在Oolong-Synth子集上,Qwen达到91.1%准确率,Gemma达到99.3%。
Insight: 创新点在于引入了模型侧聚合接口,将超对数草图状态与冻结的语言模型解耦,实现了紧凑、可合并的聚合状态维护,避免了传统方法中的额外推理循环,为长上下文中的集合操作提供了高效解决方案。
Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped statistics, which widely exist in logs, program outputs, tables, and multi-turn conversations. To provide the aggregation state required by these tasks, we introduce a model-side aggregation interface that maintains compact Hash-based HyperLogLog (HLL) sketch states alongside a frozen language model. While the model processes the context, an extractor maps each relevant record to a canonical identity. The identity is then hashed and updates the HLL state. These states can be merged across context segments and/or read out directly for downstream reasoning, avoiding an additional generate-execute-return cycle. We validate the proposed approach by setting the HLL state size as 2 KiB (2,048 registers), which does not increase with context length or set cardinality. In a distinct-count experiment involving one million records, the mean relative error was 1.6%. In a separate merge test, states built from as many as 256 segments produced exactly the same readout as a single pass over the same stream. On 3,969 aggregate-then-reason tasks from 174 source windows, the fixed-budget interface reached 99.2% accuracy on Gemma 4 (31B, BF16), compared with 100.0% under exact aggregation; the paired gap was 0.8 percentage points (95% window-cluster CI: 0.5-1.3 points). On a matched set of 174 items, our method improved over direct full-context reasoning by 63.2 points on Qwen and 56.3 points on Gemma. The corresponding gains over chain-of-thought (CoT) reasoning were 60.9 and 63.2 points, respectively. On a fixed 1,200-task Oolong-Synth subset, our method reached 91.1% on Qwen and 99.3% on Gemma. Code is available at https://github.com/songdc98/sketchops.
[8] Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation cs.CLPDF
Xuan Feng, Guihong Liu, Tianlong Gu, Shuai Zhao, Xuemin Wang
TL;DR: 本文提出专家引导的互蒸馏(EGMD)方法,以解决多模态假新闻检测器因依赖不可靠证据(如领域特定捷径和语义不一致的图文对)而导致的跨领域泛化能力差的问题。该方法通过输入级校准、表示级专家引导教师对齐领域统计,以及决策级原型锚定领域特定学生进行互学习和双通道蒸馏,从而学习在整个预测流程中信任何种证据。
Details
Motivation: 动机在于解决多模态假新闻检测器因数据不平衡和语义不一致的图文对导致的跨领域泛化能力差的问题,这些因素使模型依赖不可靠的领域特定捷径证据。
Result: 在两种语言的四个数据集上,EGMD实现了最先进的准确率,同时将领域偏差降低了高达57.3%;作者还构建了Weibo_Balanced基准,以隔离不平衡对泛化的影响。
Insight: 创新点包括输入级校准编码图文对一致性作为共享增益、专家引导教师对齐领域统计并集中领域特定模式,以及原型锚定学生通过互学习和双通道蒸馏继承教师特征几何和校准预测,从而减少局部领域先验。
Abstract: Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and semantically inconsistent text-image pairs that make cross-modal evidence unreliable. We propose Expert-Guided Mutual Distillation (EGMD), which learns what evidence to trust across the prediction pipeline. At the input level, input-level calibration encodes pair-level coherence as a shared gain before fusion. At the representation level, an expert-guided teacher aligns domain statistics and encourages domain-specific patterns to concentrate in specialized experts. At the decision level, prototype-anchored domain-specific students use mutual learning and dual-channel distillation to inherit the teacher’s feature geometry and calibrated predictions while discouraging local domain priors. We further construct Weibo_Balanced, a domain-balanced benchmark that isolates the effect of imbalance on generalization. Across four datasets in two languages, EGMD achieves state-of-the-art accuracy while reducing domain bias by up to 57.3%.
[9] Contrastive ESA: Human Evaluation of Multiple Translations at Once cs.CL | cs.HCPDF
Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley, Martin Popel
TL;DR: 本文提出了一种名为对比性错误跨度标注(cESA)的新协议,用于机器翻译的人工评估。该协议通过同时展示同一源输入的多个翻译版本,让标注者在共享上下文中标记错误跨度并给出绝对质量评分,从而减少标注噪声和成本。
Details
Motivation: 当前机器翻译的人工评估通常孤立地评估单个输出,存在标注噪声高、成本高的问题,cESA旨在通过对比多译文来提升评估的一致性和效率。
Result: 在英日翻译的12个模型大规模人工评估中,cESA相比标准逐点评估减少了标注时间和噪声,并能直接生成绝对质量判断,无需事后校正即可进行可解释的非参数模型排名。
Insight: 创新点在于将对比评估与错误跨度标注结合,允许标注者跨多译文共享上下文进行判断,从而获得更一致、高效的绝对质量评分,这为机器翻译评估提供了新的实用协议。
Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.
[10] Constitutional Midtraining: Content Presence Drives Alignment Gains cs.CL | cs.AI | cs.CY | cs.LGPDF
Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard
TL;DR: 本文提出了一种名为’宪法中期训练’的方法,通过在模型训练中期(而非训练后)注入基于原则和价值观的内容(即’宪法’内容),来增强语言模型的持久对齐能力。研究通过大规模实验(120B规模)对比了仅重放的控制组与四种不同宪法中期训练条件,评估了模型在压力下的对齐、价值冲突解决、勒索等任务上的表现,发现宪法中期训练能带来广泛且持久的对齐增益,且不损害模型的一般能力。
Details
Motivation: 动机在于探索训练中期的干预(与训练后对齐隔离)是否能产生比训练后对齐更持久、更深入的对齐效果,因为训练后对齐通常被认为是浅层的,容易在微调中被削弱。
Result: 在自建和已有基准测试(包括压力对齐、价值冲突、勒索等)上,宪法中期训练的模型在泛化性和持久性上优于控制组,尤其在勒索任务上优势显著(SFT后优势仍存,降低了17.5个百分点)。但在需要主动抵抗上下文压力或冲突的场景中,优势在SFT后会减弱。在能力测试(MMLU、ARC-Easy、piqa、GSM8K)上,宪法中期训练平均未造成性能损失。
Insight: 创新点在于提出了’宪法中期训练’这一新范式,证明了在训练中期注入价值观内容能带来持久、低成本的对齐增益,且内容的存在比其结构(如课程顺序、审议推理)更重要,为以SFT为中心的流程提供了一个有效的补充方案。
Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic’s Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
[11] Metis: Memory Foundation Model cs.CL | cs.LGPDF
Zeyu Zhang, Ziliang Guo, Yihang Sun, Xichong Zhang, Xixuan Hao
TL;DR: 本文提出了首个记忆基础模型原型Metis,旨在为现有基础模型赋予原生记忆能力。该模型通过内部持久且动态演化的记忆状态,以及自主存储和利用信息的原生记忆过程,实现了无需外部模块的端到端记忆功能。
Details
Motivation: 当前AI智能体的记忆功能主要依赖外部模块实现,基础模型本身的原生记忆能力尚未被充分探索,本文旨在填补这一空白。
Result: 通过大量实验,论文证明了Metis具备原生记忆能力,并对其优势、局限和行为进行了详细分析。
Insight: 创新点在于从架构层面将记忆状态内化于模型主干,并通过记忆注意力机制进行访问;同时,通过构建大规模记忆专用训练数据和多重优化目标,在中训练阶段习得原生记忆过程,实现了无需梯度更新的在线记忆维护。
Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inference time, all learned model weights remain frozen, while the native memory states are autonomously transformed through standard forward computation. Through extensive experiments, we show that Metis exhibits native memory capabilities and further provide a detailed analysis of its strengths, limitations, and behaviors. To facilitate future research on memory foundation models, we release our project and model checkpoints.
[12] SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cs.CLPDF
Jianze Wang, Kunwang Zheng, Ying Liu, Yu Cao, Qilong Zhang
TL;DR: SERPO是一种用于开放式测试时强化学习的方法,它通过共同演化响应证据、查询特定评分标准和策略参数,构建了一个自我演化的闭环系统,以在没有外部奖励模型的情况下优化语言模型。
Details
Motivation: 现有测试时强化学习方法依赖于答案投票,无法自然扩展到开放式生成任务,因为有效响应无法映射到共享的规范答案,因此需要从模型自身输出中构建可靠的奖励信号。
Result: 在两个模型配置、两个领域内基准和四个领域外基准上,SERPO将HealthBench和ResearchQA的性能分别提升了20.63和20.31分,六个基准的宏观平均提升了8.06分,并支持领域外迁移和跨基准持续演化。
Insight: SERPO的创新点在于用三向演化闭环(响应演化、评分标准演化和策略演化)替代答案投票,通过好坏响应归档和概率标准评分生成奖励信号,实现了开放式生成任务的自我优化和跨基准适应。
Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model’s own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.
[13] Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion cs.CLPDF
Jinlan Liu, Zhiying Tu, Yongchao Xing, Yicheng Liu, Bolin Zhang
TL;DR: 本文提出了一种名为DuPLeR的双路径大语言模型推理框架,用于解决多模态少样本知识图谱补全任务。该框架通过结合多模态LLM先验与事实支持结构来构建校准的关系图,并在细化的关系拓扑上进行双层次结构推理,同时利用查询相关的多模态信号来调节信息传递以增强实体表示。
Details
Motivation: 现实世界知识图谱中新实体和关系的出现使得归纳式知识图谱补全在少样本和零样本场景下变得困难,而现有的多模态信息和LLM先验虽然能丰富稀疏的关系上下文,但也可能引入噪声或幻觉证据,因此需要一种能有效整合并校准这些信息的框架。
Result: 在两个多模态知识图谱基准的八个归纳变体上的实验表明,DuPLeR在数据稀缺的知识图谱补全场景中实现了鲁棒的性能,具体表现为在少样本设置下达到了先进的水平。
Insight: 创新点在于提出了双路径推理机制,即通过校准的关系图进行结构推理,并结合查询相关的多模态信号来调节和补充实体表示,从而有效利用LLM先验和多模态信息,同时减轻噪声和幻觉的影响,为少样本知识图谱补全提供了新的解决方案。
Abstract: Knowledge graph completion (KGC) aims to infer missing facts in knowledge graphs (KGs), thereby improving their completeness and supporting downstream intelligent applications. However, emerging entities and relations in real-world deployments make inductive KGC difficult, especially under few-shot and zero-shot settings. Multimodal information and Large Language Model (LLM)-derived priors can enrich sparse relational contexts, but they may also introduce noisy or hallucinated evidence. To address these issues, we propose DuPLeR, a \textbf{Du}al-\textbf{P}ath \textbf{L}LM \textbf{R}easoning framework for multimodal few-shot KGC. DuPLeR builds a calibrated relation graph by combining multimodal LLM-derived type priors with factual support structures, and performs dual-level structural reasoning over the refined relation topology. Moreover, a dual-pathway multimodal enhancement module regulates message passing with query-relevant multimodal signals and supplements entity representations after graph propagation. Experiments on eight inductive variants of two multimodal KG (MMKG) benchmarks show that DuPLeR achieves robust performance in data-scarce KGC scenarios.
[14] Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? cs.CLPDF
Arnav Hiray, Agam Shah, Caleb Lu, Meghaj Tarte, Harsit Mittal
TL;DR: 论文介绍了首个基于真实信用卡协议构建的金融素养数值推理基准CreditCardQA,包含1,800个问题,包括反映消费者自然提问方式的第一人称变体。评估了多种大语言模型和推理模型在思维链和程序链提示下的表现,发现程序链能提升性能,尤其对基础推理较弱的模型有效,并缩小了开源与闭源系统间的差距。错误分析表明,失败更多源于金融规则误用、条件遗漏和合同条款误解,而非算术错误,且涉及比较、条件逻辑和货币约束的问题尤其困难,错误常出现在可能影响低收入或财务脆弱人群的边缘案例中。
Details
Motivation: 解决现有语言模型在真实金融场景下的数值推理能力评估不足的问题,特别是针对消费者日常遇到的信用卡费用、利息和支付等复杂合同条款的理解与计算。
Result: 在CreditCardQA基准上,程序链提示相比思维链带来一致的性能提升,尤其改善了基础推理较弱模型的表現,并缩小了开源与闭源系统之间的差距;错误分析显示模型在比较、条件逻辑和货币约束任务上表现较差。
Insight: 创新点在于构建了首个基于真实金融协议的数值推理基准,强调第一人称自然提问和边缘案例;客观分析认为,研究揭示了语言模型在应用领域特定规则和复杂条件推理上的局限性,程序链提示可作为提升模型结构化推理的有效策略,且基准设计关注了社会影响(如对脆弱人群的潜在风险)。
Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting. Overall, PoT yields consistent performance gains, particularly for models with weaker baseline reasoning, and narrows gaps between open- and closed-source systems. Through error analysis, we show that failures arise less from arithmetic and more from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. We further analyze question difficulty and find that comparisons, conditional logic, and monetary constraints are especially challenging. We also find that errors often arise in edge cases such as late-payment penalties or small-balance scenarios that are more likely to affect lower-income or financially vulnerable individuals.
[15] TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning cs.CLPDF
Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen
TL;DR: 本文提出了TREK(旅行推理与评估套件),这是一个用于评估LLM智能体在复杂旅行规划任务中合成可行行程的基准测试。该基准包含800个多约束任务,并提供了一个包含21万条记录的合成知识库和基于RESTful API的工具沙箱。通过完全确定性的规则评估器(而非LLM评判),对15个LLM智能体在九个约束维度上进行了评估,发现即使最强的GPT-5.6模型也仅在46.2%的可解任务中生成完全可行的计划。
Details
Motivation: 现有智能体基准测试通常单独评估行程的各个属性,并使用软性指标或LLM评判标准,无法确保生成的计划可执行、可复现或可审计。TREK旨在解决这一问题,为合成同时满足多重约束(如约束正确性、无幻觉、时空可执行性、预算有效性及满足用户未声明的个性化需求)的可行行程提供一个严谨的评估框架。
Result: 在TREK基准上评估了15个LLM智能体,最强模型GPT-5.6仅在46.2%的可解任务中生成完全可行的计划,中位数成功率仅为6.6%,最低为0.0%。满足旅行者未声明的个性化需求是普遍瓶颈,即使在最先进的模型上也未解决。评估使用完全确定性的规则评估器,确保了结果的可复现性和可审计性。
Insight: 创新点在于构建了一个集成了合成知识库、工具沙箱和确定性规则评估器的端到端基准测试,首次实现了对行程规划智能体在多重硬约束下的联合、可验证评估。其方法论强调可复现性和可审计性,通过提供人工验证的黄金参考和确定性评分器,清晰区分了智能体能力与评估严格性之间的差距。
Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler’s unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers’ unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.
[16] DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search cs.CL | cs.IRPDF
Raphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, Amélie Chatelain
TL;DR: 本文提出了一种完全开放的端到端检索模型训练方法,通过重建和整理6.65亿英文对比预训练对和188万监督微调对,训练了DenseOn(单向量稠密模型)和LateOn(ColBERT风格延迟交互模型),并在BEIR基准上取得了该规模类别的新SOTA结果。随后通过翻译训练扩展到多语言检索,训练了mDenseOn和mLateOn模型,发现延迟交互模型在未见语言上泛化能力更强。
Details
Motivation: 当前最先进的检索模型依赖封闭训练数据,导致可复现性差距,本文旨在提供完全开放的训练方案,并研究英文监督如何通过翻译训练迁移到多语言检索。
Result: DenseOn和LateOn在BEIR基准上分别达到56.20和57.22的平均nDCG@10,为该规模类别设定了新的SOTA结果;多语言模型在翻译训练支持的语言上表现良好,且延迟交互模型对未见语言和脚本的泛化能力更强。
Insight: 创新点包括完全开放的训练数据与代码发布、翻译训练作为多语言泛化策略的探索,以及发现延迟交互模型通过词级匹配能更好地泛化到未见语言,这为多语言检索模型设计提供了新思路。
Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts. This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe. We publicly release the models, datasets, and training code.
[17] Mental World Modeling cs.CLPDF
Hao Fei, Yiran Zhao
TL;DR: 该论文提出了’心理世界建模’(MWM)理论框架,将隐藏的心理状态(如信念、意图、情感)作为世界模型的核心组成部分,以更准确地预测人类行为。论文还实例化了MENTIS这一无需训练、可完全检查的基线方法,并在一个涵盖文本、图像和音视频故事的情境决策数据集上验证了显式建模心理状态对预测人类决策的重要性。
Details
Motivation: 现有世界模型仅关注物理场景(是什么、在哪里、如何演变),而人类行为由隐藏的心理状态驱动,忽略这些状态会导致在看似正确的场景中预测出错误的行动。
Result: 在涵盖文本、图像和音视频故事的手动构建、质量可控的情境决策数据集上,对8个基于现代LLM的世界模型进行的实验表明,显式建模心理状态对于预测人类决策至关重要。
Insight: 核心创新在于将心理变量(信念、愿望、意图等)作为世界模型的一等公民,与物理状态耦合,并提出了一个无需训练、模块化且可解释的基线实现MENTIS,为从模拟物理场景到模拟其中行动者心智的世界建模指明了新方向。
Abstract: World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.
cs.CV [Back]
[18] DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving cs.CVPDF
Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li, Mattia Segu, Lei Ke
TL;DR: 本文提出DVPSFormer,一种高效的在线深度感知视频全景分割架构,用于自动驾驶场景的4D场景理解。该方法通过显式场景离散化机制,利用分割查询表示前景和背景区域,并结合离散到连续深度头单次解码度量深度,同时引入在线多数投票机制利用时序一致性优化实例跟踪中的分类。
Details
Motivation: 解决现有深度感知视频全景分割方法依赖计算昂贵的多阶段流程或离线跟踪,无法满足实时决策需求的问题,旨在实现高效统一的在线4D场景理解。
Result: 在Cityscapes-DVPS和SemKITTI-DVPS基准测试上达到了新的最先进水平。
Insight: 创新点在于显式场景离散化机制和在线多数投票机制,前者将语义与几何学习紧密耦合并显著降低延迟,后者利用时序一致性提升跟踪鲁棒性,为在线机器人感知提供了简洁高效的解决方案。
Abstract: Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.
[19] Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cs.CVPDF
Tianci Wu, Siqi Cao, Guangming Zhu, Jiang Lu, Siyuan Wang
TL;DR: 本文提出了一种名为KDA(Knowledge-guided Disentanglement with Atomic Actions)的方法,用于解决复杂场景中动作识别面临的挑战。该方法利用大语言模型将动作标签分解为原子动作,通过知识注入模块和知识解耦模块增强视频特征表示,并引入知识解耦损失以实现更精确的语义解耦,从而提升动作识别的性能。
Details
Motivation: 复杂场景中的动作识别通常涉及多个并发细粒度动作,现有基于整体表示的方法难以捕捉细微交互和细粒度语义,而基于提示或视觉/结构线索的解耦方法又缺乏明确的语义指导或粒度较粗。
Result: 大量实验表明,KDA提高了特征可区分性,并在多标签动作识别基准测试中取得了最先进的性能。
Insight: 创新点在于利用LLMs提供的细粒度原子动作语义作为显式指导,通过知识注入与解耦模块实现动作表示的增强与精确解耦;所提出的模块具有良好的通用性,可轻松集成到其他方法中。
Abstract: Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely on holistic representations, which are insufficient for capturing subtle interactions and fine-grained semantics. While recent prompt-based approaches introduce disentanglement, they lack explicit semantic guidance, and methods based solely on visual or structured cues remain coarse-grained. In this paper, we propose Knowledge-guided Disentanglement with Atomic Actions (KDA), which leverages fine-grained semantic knowledge to enhance action representations and enable more precise disentanglement. Specifically, we use Large Language Models (LLMs) to decompose action labels into atomic actions, providing explicit spatial-temporal semantics. A Knowledge Injection Module (KIM) first integrates atomic action knowledge into video features. Based on this enhanced representation, a Knowledge Disentanglement Module (KDM) further disentangles atomic action knowledge to produce more precise semantic guidance for action disentanglement. A Knowledge Disentanglement Loss (KD Loss) is introduced to encourage clearer disentanglement of knowledge components within KDM. Extensive experiments demonstrate that KDA improves feature discriminability and achieves state-of-the-art performance on multi-label action recognition benchmarks. Moreover, KIM and KDM can be readily integrated into other methods, demonstrating strong generality.
[20] A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment cs.CV | cs.AI | cs.LGPDF
Hua Qian, Manisha Kotha, Tuan Tran, Jennifer Shin, Haining Zheng
TL;DR: 本研究开发了一种混合计算机视觉方法,通过图像量化暴露皮肤面积以进行皮肤暴露评估。该方法首先使用Mask R-CNN从170张室内涂装图像中识别人体并去除背景干扰,然后采用基于颜色的算法分割暴露皮肤。
Details
Motivation: 解决传统皮肤暴露评估方法依赖人工估算、效率低且难以规模化的问题,旨在从图像中自动提取半定量暴露信息。
Result: 生成的暴露皮肤与身体像素比例与人工估算结果的一致性约为80%,证明了方法的有效性。
Insight: 创新点在于结合实例分割(Mask R-CNN)与颜色分割的混合深度学习框架,为图像中的暴露评估提供了可扩展的解决方案,并可扩展至身体部位识别、个人防护装备检测和视频分析等应用。
Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background interference; a color-based algorithm then segmented exposed skin. The resulting exposed-skin-to-body pixel ratios showed approximately 80% agreement with human estimates. The approach demonstrates a scalable way to extract semi-quantitative exposure information from images, with future extensions to body-part recognition, PPE detection, and video-based exposure analysis.
[21] TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions cs.CV | cs.AIPDF
Xinran Liu, Shouqian Shi, Yutong Chen, Ge Wang, Xin-Wei Yao
TL;DR: TraceCLIP是一个无需训练的方法,旨在从CLIP模型的内部机制中恢复局部语义信息。它通过分析CLS注意力输出中的特定patch贡献,重建出具有判别性的局部语义特征,并将其用于密集视觉语言理解任务,如零样本语义分割。
Details
Motivation: CLIP的全局图像-文本对齐目标未能直接约束局部视觉-语言对应关系,现有方法要么需要额外监督或任务特定适配,要么无法有效定位CLIP内部最易访问局部语义的位置。
Result: 在八个零样本语义分割基准测试中,TraceCLIP无需额外训练、外部视觉基础模型或区域级监督,在两个主干网络和不同背景设置下,平均mIoU比之前最强的无需训练方法提升了1.3到4.5个百分点。
Insight: 核心创新在于提出了一种无需训练的分析框架,通过追踪和隔离写入CLS注意力输出的patch特定贡献来恢复潜在的局部语义证据,并进一步利用贡献特征构建语义-测地拓扑门来校准最终层的patch亲和力。这表明在全局对齐表征的内部构造中,空间局部化的语义可能仍然是可访问的。
Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.
[22] Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography cs.CVPDF
Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci
TL;DR: 本文提出了Rad-JEPA 3D,一个用于3D计算机断层扫描(CT)的自监督预训练联合嵌入预测框架。该方法通过一个融合Mamba状态空间分支和分组查询注意力分支的混合H-Mamba编码器,从掩码视图预测完整扫描的潜在特征,并结合隐状态正交正则化(HSOR)来提升表示质量。在约12万份CT扫描上预训练后,该模型在封闭式视觉问答和空间推理任务上取得了最先进或极具竞争力的结果。
Details
Motivation: 解决现有3D医学图像编码器在自监督预训练中,难以有效保留下游任务(如器官解缠、异常检测、空间理解)所依赖的粗粒度空间和几何结构的问题。
Result: 在约12万份CT扫描上预训练后,Rad-JEPA 3D仅用40亿参数,就在封闭式视觉问答(VQA)任务上取得了与最先进模型(SOTA)相当的结果,并在Spatial-Med基准测试中获得了最佳的平均空间推理分数。
Insight: 核心创新点在于提出了混合H-Mamba编码器,将擅长建模切片间连续性的Mamba分支与擅长捕捉跨平面空间上下文的分组查询注意力分支相结合,并通过轻量级路由机制融合。此外,提出的隐状态正交正则化(HSOR)通过层间对齐师生隐状态来减少特征冗余,从而生成更一致和更具判别性的体积表示。
Abstract: Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet existing volumetric encoders often fail to preserve the coarse spatial and geometric structure that downstream reasoning depends on, limiting their performance on organ disentanglement, abnormality detection, and spatial understanding when paired with language models. We introduce Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view. At its core is a hybrid H-Mamba encoder that fuses a Mamba state-space branch, which models inter-slice continuity through sequential scanning, with a grouped-query attention branch, which captures cross-plane spatial context, combined through a lightweight per-token router. To improve the quality of intermediate representations, we further propose Hidden States Orthogonal Regularization (HSOR), which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder. This layer-wise regularization produces more consistent and discriminative volumetric representations, leading to improved performance on organ recognition and spatial reasoning tasks. Pretrained on approximately 120,000 CT scans, Rad-JEPA 3D attains state-of-the-art results despite its compact size: with only 4.0B total parameters, it achieves competitive results with state-of-the-art on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark. Ablation studies confirm that the hybrid block and HSOR contribute complementary gains, and that the induced spatial structure can substitute for raw language-model scale on volumetric reasoning tasks.
[23] WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models cs.CVPDF
Jiamin Xu, Cong Wang, Zheng Dong, Chi Wang, Renshu Gu
TL;DR: 本文提出WildShadowRemover框架,通过LoRA微调预训练视频扩散模型,结合细节注入模块和阴影掩码引导的频率分解调制模块,以保留高频纹理并抑制阴影伪影,同时利用Depth Anything的单目深度先验提供几何感知指导,用于解决复杂光照和多样阴影外观下的野外视频阴影去除问题。
Details
Motivation: 解决野外视频阴影去除的挑战,包括复杂光照、多样阴影外观和有限训练数据,该任务在无约束真实场景中尚未被充分探索。
Result: 在构建的大规模配对视频阴影去除数据集WildShadow上进行广泛实验,结果表明该方法在阴影去除质量和时间一致性上优于现有方法,在具有挑战性的野外场景中产生时间一致的无阴影视频,并展现出强大的泛化能力。
Insight: 创新点包括:通过LoRA微调适配预训练视频扩散模型以保留生成先验;引入细节注入模块和阴影掩码引导的频率分解调制模块来选择性恢复纹理;利用单目深度先验提供几何感知指导;构建了大规模合成数据集WildShadow作为基准。
Abstract: Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model’s powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.
[24] Spline-Based Boundary Representations for Sparse View Reconstruction and Simulation Using Isogeometric Analysis cs.CV | cs.CEPDF
Davor Dobrota, Vsevolod Skorokhodov, Chenghao Xu, Olga Fink, Malcolm Mielle
TL;DR: 本文提出FORGE-SIM方法,直接从稀疏的带姿态RGB图像中重建出多片B样条边界表示模型。该方法通过优化样条表示本身,生成紧凑、平滑且封闭的几何体,这些几何体天然兼容计算机辅助设计和仿真工作流。此外,该方法还能将观测数据(如热状态和语义信息)投影到重建模型上,以支持即时仿真应用。
Details
Motivation: 解决基于图像的重建方法虽然能恢复视觉细节,但其生成的模型表示通常不适合数值仿真(如缺乏封闭性、平滑性)的问题,旨在弥合计算机视觉与数值分析之间的鸿沟。
Result: 实验表明,所获得的模型质量足够高,能够支持热仿真和模态分析,证明了其在仿真应用中的有效性。
Insight: 主要创新点在于将基于图像的重建与仿真就绪的建模统一在一个优化框架内,直接重建出与CAD和仿真工作流原生兼容的B样条边界表示,并支持观测数据的同基投影,从而为仿真驱动的设计、检测和数字孪生应用提供了新工作流。
Abstract: Image-based reconstruction aims to recover three-dimensional geometry from images. Recent advances have enabled the recovery of visually detailed models, yet their representations are not well-suited for numerical simulation. Simulation frameworks typically require explicit, watertight, and smooth geometries to ensure numerical robustness and accuracy, properties that surfaces extracted from image-based reconstructions lack. We propose FORGE-SIM, a method to directly reconstruct a multi-patch B-spline boundary representation from sparse posed RGB images without manual intervention. By optimizing the spline representation itself, our approach produces compact, smooth, and watertight geometries that are natively compatible with both Computer Aided Design and simulation workflows. Additionally, we introduce a strategy to project observation-derived fields, such as a thermal state and semantic information, onto the reconstructed models in the same spline basis, enabling immediate use in simulation. We demonstrate that the obtained models are of sufficiently high quality to enable thermal simulation and modal analysis. By unifying image-based reconstruction and simulation-ready modeling within a single optimization framework, this work removes a long-standing barrier between computer vision and numerical analysis. We anticipate that it will enable new workflows for simulation-driven design, inspection, and digital twin applications.
[25] LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models cs.CVPDF
Bowen Chen, Shreshth Saini, Balu Adsumilli, Alan C. Bovik
TL;DR: 本文提出了LumaGuide,一种无需训练的扩散模型框架,通过可微分能量引导在采样过程中调整输出分布以匹配目标特征分布,并将其应用于高动态范围(HDR)图像生成,通过控制感知均匀PQ空间中的亮度分布实现HDR一致的行为,同时保持语义保真度。
Details
Motivation: 预训练扩散模型受限于训练数据的统计偏差,难以生成高动态范围(HDR)内容,因此需要一种无需重新训练的方法来引导模型生成HDR图像。
Result: 实验结果表明,通过对齐亮度直方图足以诱导HDR一致的行为,包括连贯的高光和保留的阴影细节,同时保持语义保真度。
Insight: 创新点在于提出了一种无需训练的分布塑形框架,通过可微分能量引导在采样时直接控制输出分布,实现了灵活的HDR生成,并可扩展至视频生成与时间一致性约束,为可控生成提供了新思路。
Abstract: Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models. Instead of modifying model parameters, LumaGuide steers the sampling process to match target feature distributions via differentiable energy-based guidance. We instantiate this framework for HDR generation by controlling luminance distributions in perceptually uniform PQ space. Our results show that aligning luminance histograms is sufficient to induce HDR-consistent behavior, including coherent highlights and preserved shadow detail, while maintaining semantic fidelity. Beyond HDR, LumaGuide enables flexible specification of target distributions through data-driven presets, reference images, or text-driven predictors, and extends naturally to video generation with temporal consistency constraints. More broadly, our work demonstrates that controllable generation can be achieved by directly shaping output distributions at sampling time, without retraining diffusion models.
[26] Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment cs.CV | cs.LG | eess.IVPDF
Takeshi Nishikawa
TL;DR: 本文研究了一种用于风力涡轮机防撞的轻量级猛禽物种实时分类方法。通过使用DINOv2-L作为教师模型,对三种轻量级学生模型进行知识蒸馏,并利用视频帧提取技术扩充数据集以解决相似物种间的混淆问题。最终,模型集成在边缘设备上实现了高效的实时推理。
Details
Motivation: 旨在开发一种轻量级模型,用于实时识别猛禽物种,以减少风力涡轮机与鸟类的碰撞。核心挑战在于区分外观相似的物种,并满足边缘设备对计算效率和精度的双重需求。
Result: 在采用视频和源图像级别分组划分的数据集上,三个学生模型的集成宏召回率达到0.935 +/- 0.004(在传统图像级别划分上为0.955),参数量约为教师的八分之一。在独立测试集上,白尾海雕的召回率最高提升了38.6个百分点,且被误判为虎头海雕的错误率从61%降至15%。在NVIDIA Jetson Orin Nano上部署的EfficientNet-B0模型,FP16推理速度达到3.19毫秒/图像。
Insight: 主要创新点在于通过视频帧提取有效扩充稀有物种数据集以缓解类别不平衡和混淆问题,并验证了知识蒸馏在保持教师模型大部分性能的同时实现模型轻量化的有效性。研究还表明,数据集扩充和教师模型微调是性能提升的关键,而非单纯更换更先进的教师模型或仅使用蒸馏技术。
Abstract: We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parameters) as a teacher, we distilled three lightweight students (MobileNetV4, ViT-Small, and EfficientNet-B0). To reduce confusion between closely related species, we expanded the dataset to 12,519 images, including an increase in Steller’s Sea Eagle images from 463 to 2,050 via video-frame extraction. Under a group split that separates samples at the video- and source-image level to mitigate source leakage at that granularity, the three-student ensemble achieved a macro recall of 0.935 +/- 0.004 over five distillation seeds (0.955 on a conventional image-level split, retaining 97.5% of the teacher’s macro recall) with roughly one-eighth as many parameters. On a subset of 1,258 images disjoint from the former training images, White-tailed Eagle recall improved by up to 38.6 percentage points, while the rate at which it was misclassified as the Steller’s Sea Eagle decreased from 61% to 15% of errors. TensorRT FP16 deployment of EfficientNet-B0 on an NVIDIA Jetson Orin Nano achieved 3.19 ms/image including host-device transfer (313 images/s), with 99.95% argmax agreement with FP32. In five-seed controlled comparisons, neither distillation (versus CE-only) nor the change from a DINOv2-L to a DINOv3-L teacher yielded a clear ensemble-level improvement; the primary gains stem from the dataset expansion and teacher re-fine-tuning.
[27] Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models cs.CVPDF
Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu
TL;DR: 该论文研究了多模态大语言模型(MLLMs)在处理视觉中心任务时的失败情况,特别是当视觉证据与预训练知识冲突时。作者通过图像重建和提出的WhatIfVis基准测试,分析了视觉信息的可用性和多模态上下文敏感性。研究发现,MLLMs能够编码粗粒度视觉证据,但在利用这些证据时存在不可靠的控制问题,通过监督微调和激活修补可以改善这种可控性。
Details
Motivation: 解决MLLMs在视觉中心任务中表现不佳的问题,尤其是当视觉输入与模型预训练知识相矛盾时,探究其失败原因在于视觉编码还是后感知利用。
Result: 在提出的WhatIfVis基准测试(涵盖时空、颜色、数量、大小和重量五个维度)上,未经微调的原始模型表现出不稳定的视觉上下文敏感性;监督微调(SFT)提高了可控性并具有跨领域泛化能力;激活修补进一步定位了所有六个模型中特定架构深度的视觉与先验权衡。
Insight: 创新点在于通过图像重建和WhatIfVis基准分离诊断视觉信息的可用性与多模态上下文敏感性,并发现视觉与先验的权衡可通过学习到的向量进行控制,这为改善MLLMs的视觉证据可靠性提供了新方向。
Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.
[28] MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer cs.CVPDF
Nazanin Amini, Kevin Desai
TL;DR: 本文提出了MoSAIC,一个用于局部参考条件运动风格迁移的潜在扩散框架。该框架通过解剖学区域分解内容和参考特征,通过单独的路径保留根轨迹,并将用户选择的参考路由到指定的身体部位。其核心贡献是对齐干预监督,通过受控的局部变换构建同步参考和反事实目标,使训练中请求的区域响应和需要保留的运动都变得直接可观察。
Details
Motivation: 编辑角色运动通常需要将一个手势或步态从一个或多个参考运动中迁移过来,同时保留源动作、时间、根轨迹和未选中的身体区域。然而,现有的运动数据集很少为任意的局部内容-参考组合提供配对目标,且自重建训练可能导致扩散模型在利用路由参考不足的情况下重现内容运动。
Result: 在一个包含128个运动和896个路由条件的冻结评估中,相对于全身路由,部分掩码路由将保留区域误差从70.64毫米降低到66.45毫米,并将匹配噪声的脱靶泄漏从18.08毫米降低到9.88毫米,同时保持了正面的选中区域响应。一项匹配预算的延续研究进一步表明,保留对齐干预监督使选中目标响应相对增加了8.8%,请求路由影响集中度提高了2.0个百分点。
Insight: 论文的创新点在于提出了对齐干预监督方法,通过构建同步参考和反事实目标来解决训练数据缺乏配对目标和模型利用参考不足的问题。从客观角度看,其将运动特征按解剖学区域分解、独立处理根轨迹以及用户指定部分路由的设计,为局部、可控的运动编辑提供了更精细的响应-保留权衡方案。
Abstract: Editing character motion often requires transferring a gesture or gait from one or more reference motions while preserving the source action, timing, root trajectory, and unselected body regions. Existing motion datasets, however, rarely provide paired targets for arbitrary part-local content–reference combinations, and self-reconstruction training may allow a diffusion model to reproduce the content motion while underusing the routed reference. We present MoSAIC, a latent diffusion framework for part-local reference-conditioned motion style transfer. MoSAIC factorizes content and reference features by anatomical region, preserves the root trajectory through a separate conditioning pathway, and routes user-selected references to individual body parts. Its central contribution is aligned intervention supervision, which constructs synchronized references and counterfactual targets through controlled local transformations, making both the requested regional response and the motion to be preserved directly observable during training. In a frozen evaluation comprising 128 motions and 896 routed conditions, part-masked routing reduces preserved-region error from 70.64 to 66.45mm and matched-noise off-target leakage from 18.08 to 9.88mm relative to whole-body routing, while retaining a positive selected-region response. A matched-budget continuation study further shows that retaining aligned intervention supervision produces an 8.8% relative increase in selected-target response and a 2.0-percentage-point increase in requested-route influence concentration. These results demonstrate that MoSAIC improves the response–preservation trade-off required for selective and controllable part-local motion editing.
[29] Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation cs.CVPDF
Pengyu Jie, Wanquan Liu, Rui He, Pengcheng Li, Weiping Wen
TL;DR: 该论文提出了一种用于未配准白光成像(WLI)和窄带成像(NBI)内窥镜病灶分割的可靠性感知复数域融合框架。该框架首先建立拓扑正则化的特征对应关系并评估其可靠性,然后基于此可靠性,在可学习的复数表示中,有选择性地融合WLI和NBI特征,其中WLI主要提供外观相关的幅度响应,NBI提供结构敏感的相位响应。
Details
Motivation: 内窥镜检查中,WLI和NBI能提供互补的病灶视图,但由于视角变化、组织形变和手持顺序采集,其成对观测图像通常存在空间未配准问题,直接融合容易混合非对应区域,甚至可能损害病灶边界的分割效果。
Result: 在成对的WLI/NBI内窥镜数据集上的实验表明,所提出的可靠性感知配准基础和复数域融合方法持续提升了病灶分割性能。角色互换和模块消融研究进一步验证了模态角色设计和可靠性引导的跨模态交互的必要性。
Insight: 创新点在于明确建模了WLI和NBI在复数域中的不同角色(幅度 vs. 相位),并引入可靠性感知机制来抑制局部不匹配区域的不可靠跨模态交互,这为处理未配准多模态图像融合提供了一种新思路。
Abstract: White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries. To address this problem, we propose a reliability-aware complex-domain fusion framework for paired-but-unregistered WLI/NBI lesion segmentation. The framework first establishes topology-regularized feature correspondence and further estimates where the cross-modal correspondence is reliable. Guided by this reliability, the model selectively fuses WLI and NBI features in a learnable complex representation. In this representation, WLI-derived cues mainly provide appearance-related magnitude responses, while NBI-derived cues provide structure-sensitive phase responses. Unlike conventional real-valued or symmetric multimodal fusion, the proposed method explicitly models the different roles of WLI and NBI and suppresses unreliable cross-modal interaction in locally mismatched regions. Experiments on paired WLI/NBI endoscopic datasets show that the proposed reliability-aware registration grounding and complex-domain fusion consistently improve lesion segmentation performance. Role-reversal and module ablation studies further validate the necessity of both the modality-role design and reliability-guided cross-modal interaction.
[30] Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering cs.CVPDF
Yu Wang, Sharon Li
TL;DR: 本文通过提出跨分支语义引导框架,探究统一多模态模型的理解与生成分支是否共享可迁移的语义空间。研究发现,从理解分支提取的语义方向可有效引导生成分支实现可控图像合成,但反向引导效果有限,揭示了架构统一并不保证语义对齐。
Details
Motivation: 旨在探究统一多模态模型中理解与生成能力是否共享统一且可迁移的语义空间,解决因异构表示和训练目标差异导致的直接比较困难。
Result: 在图像合成任务中,理解分支的引导向量能提升语义忠实度,而生成分支向理解分支的引导效果有限,表明存在表征不匹配。
Insight: 创新点在于提出基于干预的跨分支语义引导框架,用于探测多模态表征;客观分析揭示了理解分支捕获对象级语义而生成分支编码低层外观特征的不对称性,为模型可解释性提供了新工具。
Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.
[31] FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing cs.CV | cs.AIPDF
Hongyang Wang, Yichen Shi, Hongrui Li, Yiru Huo, Jun Feng
TL;DR: 本文提出FAS-R1,一个面向推理的统一多任务多模态大语言模型框架,用于人脸反欺诈任务,涵盖真实性分类、攻击类型识别和欺诈区域定位。该方法采用两阶段训练策略,包括使用高质量长链思维数据集进行监督微调,以及针对FAS任务的GRPO后训练,并通过退化模拟增强和难度感知GRPO提升模型对困难攻击样本的泛化能力。
Details
Motivation: 现有的人脸反欺诈模型多为判别式且以标签为中心,而基于MLLM的方法虽然能提供结构化输出,但仍主要依赖监督微调,导致生成的推理依据模板化且对困难攻击的优化不足。
Result: 主要3B参数的FAS-R1模型在域内达到98.75%的真实性准确率、93.33%的攻击类型准确率和96.30%/94.73%的AP@40/AP@50;在跨域真实性泛化和答案-推理质量评估中均优于对比系统,且不同基础模型的实验显示出良好的扩展性。
Insight: 创新点包括:1)构建高质量长链思维数据集FAS-R1-23K用于冷启动监督微调;2)提出退化模拟增强和难度感知GRPO,分别提升视觉质量变化下的稳定推理和缓解简单样本主导问题,尤其针对化妆、面具等模糊攻击;3)统一的多任务推理框架实现了端到端的语义理解和证据定位。
Abstract: Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but still rely mainly on supervised fine-tuning, often producing template-like rationales and weak optimization for difficult attacks. We propose FAS-R1, a two-stage reasoning-oriented MLLM framework for unified FAS prediction, covering authenticity classification, attack-type recognition and spoof-region localization. FAS-R1 first uses FAS-R1-23K, a high-quality long-CoT dataset, for cold-start supervised fine-tuning, and then performs FAS-specific GRPO post-training. Degradation-Simulated Augmentation (DSA) encourages stable spoof-cue reasoning across visual-quality shifts, while Difficulty-Aware GRPO (DA-GRPO) mitigates easy-sample dominance that may leave difficult task–attack groups under-optimized, especially for subtle or ambiguous attacks such as makeup and mask attacks. The main 3B FAS-R1 model achieves 98.75% authenticity accuracy, 93.33% attack-type accuracy, and 96.30/94.73% AP@40/AP@50 in-domain. It also outperforms the compared systems in cross-domain authenticity generalization and answer-and-rationale quality. Experiments with different base models further show favorable scaling behavior. The code will be released soon.
[32] EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding cs.CVPDF
Yuyun Chen, Tianao Li, TianQuan Feng, Cen Chen, Huiping Zhuang
TL;DR: 本文提出了EgoSafe-Bench,一个专门用于评估第一人称视角下视觉安全理解中因果推理能力的基准测试。该基准包含12,000个评估样本,通过一个分层的推理评估协议来强制模型进行逻辑一致的推理,而非依赖表面关联。对多个先进大视觉语言模型的评估揭示了它们在描述性任务上表现良好,但在因果推理和逻辑闭环方面存在显著缺陷。
Details
Motivation: 现有的大视觉语言模型在标准基准测试中表现出色,但在动态、部分可观察的第一人称体验中,难以区分表面关联与真正的因果逻辑推理。现有的评估主要基于第三人称监控视频和二元分类指标,无法暴露这种认知差距。
Result: 对Qwen3-VL、Gemini、VideoLLaMA 3等最先进的大视觉语言模型进行了广泛评估,结果显示模型在描述性评分上很高,但在因果推理和逻辑闭环方面表现出明显的脆弱性,存在显著的感知-推理脱钩现象。
Insight: 论文的创新点在于提出了一个专门针对第一人称安全场景的因果推理基准测试EgoSafe-Bench及其分层的推理评估协议,该协议强制模型进行从特征锚定到盲点推断和意图推理的严格推理轨迹,从而系统地评估和促进逻辑鲁棒的视频理解系统的发展。
Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions.Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.
[33] Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking cs.CVPDF
Xiaozhen Qiao, Da Zhang, Yubin Guo, Junyu Gao, Zhiyuan Zhao
TL;DR: 本文提出了一种名为SATATrack的语义感知时序自适应框架,用于解决无人机反无人机跟踪任务中的挑战。该框架通过语义感知上下文传播模块,利用目标语言描述引导时序上下文传播以保持目标身份,并通过时序感知分布对齐模块在线对齐特征分布以应对快速外观变化。
Details
Motivation: 无人机反无人机跟踪任务中,观测无人机和目标无人机同时运动导致视角快速变化、运动模糊、尺度变化和视觉相似干扰物等问题,使得传统基于固定视觉表示的方法难以进行可靠的外观匹配。
Result: SATATrack在UAV-Anti-UAV基准测试上取得了最先进的性能,同时在反无人机和无人机目标跟踪任务中保持竞争力。
Insight: 创新点在于利用稳定的目标语言描述作为语义锚点来引导时序状态传播,并结合在线特征分布对齐来减少视频特定的测试时偏移,而无需更新模型参数,从而增强了在快速变化场景下的跟踪鲁棒性。
Abstract: UAV Anti-UAV tracking is an emerging low-altitude security task for localizing an adversarial UAV using the onboard camera of a moving observer UAV. It differs from conventional UAV tracking and ground-based Anti-UAV tracking because both the camera platform and the target move simultaneously. This dual-dynamic setting induces rapid viewpoint changes, motion blur, scale variation, and visually similar distractors, making reliable appearance matching difficult. Under such rapidly changing conditions, fixed visual representations are often insufficient because target appearance becomes unreliable and feature distributions may deviate from the training domain. The target language description remains stable across frames and can therefore serve as a semantic anchor for temporal state propagation, while online feature-distribution alignment can reduce video-specific test-time shifts. In this paper, we propose \emph{SATATrack}, a Semantic-Aware Temporal Adaptation framework for UAV Anti-UAV tracking. SATATrack introduces Semantic-Aware Context Propagation (SACP), which uses the target description to guide temporal context propagation across backbone stages and preserve target identity under rapid appearance changes. An auxiliary contrastive regularizer is used during training to discourage responses to semantically similar background regions. During inference, Temporal-Aware Distribution Alignment (TADA) aligns feature distributions online without updating model parameters, combining recent-frame estimates with training-time statistics for stability. SATATrack achieves state-of-the-art performance on the UAV-Anti-UAV benchmark while remaining competitive in Anti-UAV and UAV object tracking tasks. The code will be available at https://github.com/XiaozhenQiao/SATATrack.
[34] CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling cs.CVPDF
Yuyang Huang, Yabo Chen, Wenrui Dai, Ziyang Zheng, Haibin Huang
TL;DR: 本文提出了CineWeaver,一个无需训练的统一框架,用于解决电影级长视频生成的挑战。该框架通过操纵预训练视频扩散模型中的位置编码和注意力模式来打破时间连续性,从而实现清晰的多镜头转换,并引入了镜头路由参考条件机制和锚点记忆机制,以实现细粒度控制和长视频生成。
Details
Motivation: 现有文本到视频扩散模型难以同时满足多镜头生成、对角色和场景的细粒度可控性以及长视频生成的要求,通常需要针对特定需求进行定制和重新训练。本文旨在无需重新训练的情况下,用一个统一框架同时解决这些挑战。
Result: 实验结果表明,CineWeaver能够生成长时间、高质量的电影视频,具有一致的身份、稳定的全局外观和清晰的镜头转换。
Insight: 核心创新点在于揭示了预训练视频扩散模型存在偏向时间连续性的结构偏差,并据此提出了无需训练的推理时干预方法。具体技术包括通过位置编码和注意力操作实现镜头转换、镜头路由参考条件机制实现细粒度控制,以及锚点记忆机制保持长视频的全局一致性。
Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.
[35] From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation cs.CVPDF
Shaoxuan Wu, Xiao Zhang, Xiaodi Zhao, Yunzhi Tian, Yilin Tang
TL;DR: 本文提出了一种名为TrailNet的轨迹引导不确定性感知网络,用于弱监督医学图像分割。该方法通过联合利用眼动追踪数据中的注视点和轨迹,从空间语义建模扩展到时间上下文建模,以增强目标感知。此外,网络采用多尺度不确定性解码器来减少噪声带来的监督不确定性,并引入循环蒸馏策略实现无需眼动数据的推理。
Details
Motivation: 医学图像分割严重依赖费时费力的像素级标注,而眼动追踪提供了一种可集成到临床工作流程中的低成本解决方案。然而,有效建模时间轨迹具有挑战性,且探索性注视引入的噪声限制了分割性能。
Result: 在两个公开数据集上的实验结果表明,TrailNet优于现有最先进方法,分别达到了81.25%和81.85%的Dice分数。
Insight: 创新点在于提出了轨迹引导的时空编码器来建模时间上下文并与空间语义互补交互,以及利用类别互斥约束的多尺度不确定性解码器来产生确定性预测。从客观角度看,将眼动轨迹的时间信息与空间注视点结合,并设计专门的噪声处理机制,是提升弱监督分割性能的有效途径。
Abstract: Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution that can be naturally integrated into clinical workflows. Recorded by eye trackers, gaze conveys the spatial regions of clinicians’ attention through fixations and the temporal context of clinicians’ progressive visual perception from trajectories. Nevertheless, effective modeling of temporal trajectories remains challenging, and noise in gaze caused by exploratory fixations greatly limits segmentation performance. To overcome these limitations, we propose the Trajectory-guided Uncertainty-aware Network (TrailNet), which exploits gaze-supervised medical image segmentation from spatial semantics modeling to temporal context by jointly leveraging fixations and trajectories. Specifically, the proposed trajectory-guided spatio-temporal encoder models temporal context and establishes complementary interactions with image spatial semantics to strengthen target perception. Furthermore, the multi-scale uncertainty decoder leverages category mutual-exclusivity constraints to produce deterministic predictions and mitigate supervision uncertainty induced by noise. To enable gaze-free inference, we further introduce a cycle distillation strategy that transfers feature-level knowledge via teacher-student networks. Experimental results on two public datasets demonstrate that TrailNet outperforms state-of-the-art methods, achieving Dice scores of 81.25% and 81.85%, respectively.
[36] TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models cs.CVPDF
Jinkun Zhao, Kui Zhang, Wenjun Wu
TL;DR: 本文提出了一种名为TPCD的色调压力对比解码方法,用于缓解视觉语言模型在高压提示下做出无根据承诺的问题。通过在安全中性指令和高压力指令下产生的logits进行对比解码,并结合任务先验/分歧门控机制,有效降低了攻击成功率,同时保持了较高的正面准确率。
Details
Motivation: 解决视觉语言模型在高压提示下容易做出无根据承诺(如识别无法阅读的文本、报告不确定时间或确认不存在物体)的问题,探索利用压力诱导的分布作为对比解码的负分支来缓解这一偏差。
Result: 在包含800个示例的tone-matters基准测试中,LLaVA-1.5-7B模型在高压下的攻击成功率达到66.75%,使用安全中性化后降至9.88%,而完整TPCD进一步降至0.50%。结合任务先验/分歧门控后,攻击成功率降至1.63%,同时正面准确率保持在54.44%。在GLM-4.6V和Llama-3.2-Vision模型上的留出测试也验证了简单门控机制的有效性。
Insight: 创新点在于将压力诱导的分布作为对比解码的负分支,提出TPCD方法;同时引入任务先验/分歧门控机制来平衡攻击成功率和正面准确率。从客观角度看,该方法利用模型自身在不同指令下的响应差异来检测和缓解承诺偏差,为视觉语言模型的可靠性提升提供了新思路。
Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced distribution itself can serve as a contrastive-decoding negative branch. Tone-pressure contrastive decoding (TPCD) subtracts logits produced under a high-pressure instruction from logits produced under a safe neutral instruction. On the 800-example tone-matters benchmark, LLaVA-1.5-7B under pressure reaches 66.75% attack success rate (ASR); safe neutralization reduces ASR to 9.88%; full TPCD reaches 0.50% but collapses positives to 15.56%. A benchmark-specific task-prior/disagreement gate preserves measured positive accuracy (54.44%) while lowering ASR to 1.63% on LLaVA. Treating this LLaVA analysis as the design split, full $n=800$ negative and $n=780$ matched-positive held-out runs on GLM-4.6V and Llama-3.2-Vision show that simple gates can improve over safe neutralization, with sensitivity analyses bounding the weak time-positive subtask. A category-prior-free answer-disagreement router reduces held-out aggregate ASR to 6.93%, improving over both safe neutralization (10.98%) and branch disagreement (9.67%) while matching branch disagreement’s 79.94% positive accuracy, although it remains post-hoc and surface-form based. We conclude that pressure is a useful probe of commitment bias and a viable mitigation signal, but the current gates are not yet independently validated grounding-aware detectors.
[37] MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models cs.CVPDF
Yitao Zhu, Mengjun Liu, Yingji Fu, Haowen Pang, Anqi Qiu
TL;DR: 本文提出MedARC,一种无需训练的自适应冗余压缩框架,用于处理3D医学视觉-语言模型中的视觉令牌序列过长问题。该方法通过整合视觉编码器的自注意力、视觉-文本相似性和局部特征偏离度三种互补线索来评估令牌重要性,并采用显著性感知的合并策略来压缩冗余令牌。在CT-RATE和MR-RATE数据集上的实验表明,MedARC能有效减少视觉令牌开销和推理时间,同时保持或提升诊断性能。
Details
Motivation: 3D医学图像在视觉-语言模型中会产生过长的视觉令牌序列,且存在大量空间和切片间冗余。现有令牌压缩方法通常采用均匀缩减或依赖单一重要性信号,容易误删与查询临床相关或结构独特的区域,因此需要一种更精细的自适应压缩方法来解决这一局限。
Result: 在CT-RATE和MR-RATE基准测试中,MedARC显著减少了视觉令牌开销和推理时间,同时诊断性能得以保持甚至提升。其多线索评分成本被处理更少令牌带来的节省所抵消,对于更大语言模型预期收益更明显。
Insight: 创新点在于提出了一种无需训练的多线索令牌重要性评估框架,整合了模型内在视觉焦点、查询相关性和结构独特性三种互补信号,并采用合并而非丢弃的压缩策略。从客观角度看,这种自适应冗余压缩方法为处理高维医学图像提供了一种可解释且高效的令牌管理方案,尤其适用于计算资源受限的临床部署场景。
Abstract: Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression methods typically apply uniform reduction or rely on a single importance signal, increasing the risk of removing regions that are clinically relevant to the query or structurally distinctive. To address this limitation, we propose MedARC, a unified, training-free framework for Adaptive Redundancy Compression of visual tokens in 3D medical VLMs. MedARC estimates token importance by integrating three complementary cues: self-attention from the VLM vision encoder, which reflects the model’s intrinsic visual focus; similarity between projected visual tokens and text embeddings, which identifies query-relevant regions; and deviations of local visual foundation model features from the volume-level feature center, which highlight structurally distinctive anatomy. The resulting importance distribution guides a saliency-aware merging strategy that preserves informative tokens while consolidating redundant ones rather than simply discarding them. Experiments on CT-RATE and MR-RATE show that MedARC reduces visual-token overhead and inference time while preserving or improving diagnostic performance. Its multi-cue scoring cost is outweighed by the savings from processing fewer tokens, with greater benefits expected for larger language models.
[38] Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification cs.CVPDF
Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Hamed Damirchi, Stephen Gould, Damith Ranasinghe
TL;DR: 该论文研究了视觉模型中从输入到最终表示的计算路径(即表征轨迹)对分布外(OOD)检测和图像分类的价值。不同于孤立分析中间层,作者追踪样本在深度维度上的连续变换,分离出类别一致性传输与输入特异性创新等成分。研究发现,这种轨迹在不同架构和数据集中表现出强烈的样本连续性与架构特异性模式,并能有效提升OOD检测性能和图像分类准确率。
Details
Motivation: 动机在于探究视觉模型在逐层计算过程中,最终表示所丢弃的中间信息是否包含有价值的证据,以改进OOD检测以及在干净和偏移数据上的图像分类性能。
Result: 在平衡的OpenOOD基准测试中,提出的仅使用分布内(ID)数据的“转换惊喜度”评分在131/152个非饱和比较中降低了FPR95,尤其在视觉破坏性强和语义距离远的分布偏移上提升最大。此外,冻结的更新探针在71/72个干净模型-数据集案例中提升了分类性能,偏移数据上的增益则因架构和损坏类型而异。
Insight: 创新点在于将模型的中间表示视为一个连续的、保留样本身份的轨迹进行分析,并从中分离出不同的运动成分(如类别一致性传输与输入特异性创新)。这为模型可靠性评估提供了一个与模型组织和数据偏移类型共同决定的、广泛有用的补充信号。
Abstract: Vision models do not form a representation at once; each block revises it. We ask whether the resulting computation path contains evidence that the final representation discards, and whether that evidence improves OOD detection and image classification on clean and shifted data. Unlike approaches that treat intermediate layers as separate snapshots, we retain sample identity across depth and study the transformations connecting successive states. We separate class-coherent transport from input-specific innovation, and coordinate movement from relational reorganization. Across supervised, self-supervised, vision–language, hierarchical, and convolutional encoders, these paths show strong sample-specific continuity and architecture-specific depth profiles that recur across datasets. They are also practically useful. An ID-only transition-surprise score complements strong final-state detectors, reducing FPR95 in 131/152 non-saturated comparisons on a balanced OpenOOD grid; gains are largest for visually disruptive and semantically far shifts, and remain positive on near-OOD for most detectors. Frozen update probes improve 71/72 clean model–dataset cases, while shifted-data gains vary with architecture and corruption type. Computation paths therefore provide a broadly useful reliability signal whose value is determined jointly by model organization and the shift encountered.
[39] Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer cs.CVPDF
Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Stephen Gould, Damith Ranasinghe
TL;DR: 本文研究了零样本分布外(OOD)检测器在不同部署场景下的可移植性问题,发现现有基准排名无法可靠地预测检测器性能,因为排名会因数据集和视觉语言模型(VLM)的不同而发生逆转。作者通过分析视觉语言logits中的互补证据通道(如绝对匹配水平和相对/空间锐度),解释了这种不可移植性的原因,并提出了一个不依赖OOD样本或外部语料库的检测器无关包装器——互补证据守卫(CEG),以融合多种证据来提升检测性能。
Details
Motivation: 当前选择零样本OOD检测器通常依赖于基准排名,但这一假设忽略了检测器在不同部署领域(如不同分布内数据集和VLM)下的性能可能发生逆转,导致实际应用中的可靠性问题。
Result: 在17个分布内数据集、3个VLM和7个零样本OOD检测器的实验中,CEG无需OOD样本或学习融合,显著降低了检测器敏感性,将GL-MCM的FPR95从38.1降至28.8,MCM从42.6降至30.5,优于使用熵、logit方差或随机噪声的对照方法。
Insight: 创新点在于揭示了零样本OOD检测器性能不可移植的根本原因——互补证据通道(如水平和锐度)无法相互恢复,并据此设计了一个简单有效的非补偿性融合框架CEG,仅依赖经验分布内百分位数即可提升检测鲁棒性,为实际部署提供了更可靠的解决方案。
Abstract: Selecting a zero-shot out-of-distribution (OOD) detector for a new deployment is typically based on benchmark rankings, implicitly assuming that the highest-ranked detector will transfer across domains. We show that this assumption does not hold. Through a controlled portability audit across seventeen in-distribution datasets, three vision-language models, and seven representative zero-shot OOD detectors, we find that detector rankings reverse across deployments, every detector exceeds $80%$ FPR95 on at least one domain, and the preferred detector depends on both the in-distribution data and the underlying VLM. We trace these reversals to complementary evidence channels in vision-language logits. Corpus-free detectors rely on different combinations of absolute match level and relative or spatial sharpness, while WordNet-based methods additionally depend on external semantic coverage. A simple proposition shows that level and sharpness cannot generally be recovered from one another, explaining why no single detector transfers reliably across deployments. Motivated by this diagnosis, we introduce the Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles. Controls replacing these channels with entropy, logit variance, or random noise do not reproduce the gains. Without OOD samples, auxiliary corpora, or learned fusion, CEG reduces detector sensitivity and improves GL-MCM from $38.1$ to $28.8$ and MCM from $42.6$ to $30.5$ family-balanced FPR95.
[40] SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM cs.CVPDF
Jingxuan Su, Shenglin Wang, Tiesong Zhao, Ge Li, Wei Gao
TL;DR: 本文提出了SpatialQ,一个用于评估3D高斯溅射(3DGS)场景质量的多模态框架。它通过一个增强的VGGT编码器学习3D感知的质量表示,并利用基于Qwen的多模态大语言模型进行基于多种输入(原始图像、深度图、点云渲染、相机参数)的推理,以解决现有图像质量评估方法对2D感知线索的依赖以及通用MLLM在稳定质量回归方面的不足。
Details
Motivation: 3DGS已成为新颖视图合成和3D场景重建的有效表示,产生了对可靠质量评估的需求。现有图像质量评估方法局限于2D感知线索,而通用多模态大语言模型并非为稳定质量回归设计,可能产生不可靠的判断。
Result: 摘要中未提及具体的定量实验结果或基准测试比较。
Insight: 创新点在于提出了一个专门针对3DGS场景的质量评估框架,其核心是结合了3D感知的质量表示学习(通过融合多视图特征和几何线索如深度、点云结构)以及基于多种模态输入(图像、深度、点云、相机参数)的、由MLLM驱动的接地推理机制,旨在超越仅基于外观的特征,捕捉场景级的空间结构和跨视图一致性等关键质量因素。
Abstract: 3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.
[41] JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation cs.CVPDF
Ionuţ Grigore, Călin-Adrian Popa
TL;DR: 本文提出JEPADepth,一种自监督单目深度估计框架,在标准光度重建损失基础上,引入受I-JEPA启发的掩码预测损失作为补充训练目标。该方法利用预训练DINOv3 ViT编码器的表示空间计算掩码预测损失,推理时无需额外计算开销。在KITTI基准上,JEPA目标持续提升基于DINOv3的光度基线性能,零样本迁移到Make3D和Cityscapes时达到最优或接近最优水平。
Details
Motivation: 解决自监督单目深度估计过度依赖耦合深度、姿态和外观假设的光度重建损失问题,通过引入表示空间的掩码预测任务提供互补学习信号。
Result: 在KITTI基准上,JEPA目标持续提升DINOv3基线的性能,与基于Transformer的SOTA方法竞争力相当,并优于强CNN基线;零样本迁移到Make3D和Cityscapes时,在多项指标上达到最佳或接近最佳性能。
Insight: 创新性地将I-JEPA的掩码预测范式引入深度估计,在表示空间进行结构化掩码预测以增强特征学习;采用推理时丢弃预测器和目标编码器的设计,实现零部署开销的性能提升;证明了预训练视觉Transformer表示与掩码预测任务对几何任务的有效性。
Abstract: Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Predictive Architectures (I-JEPA) for self-supervised depth learning. Our method augments a standard photometric pipeline with a masked prediction loss computed in the representation space of a pretrained DINOv3 Vision Transformer encoder. A predictor infers target-region embeddings from visible context-region embeddings under structured masking, and is discarded along with the target encoder at inference time, adding no deployment cost. On KITTI, adding the JEPA objective consistently improves performance over the same DINOv3-based photometric baseline, without changing the inference-time architecture. Compared to prior monocular self-supervised methods, JEPADepth is competitive with state-of-the-art transformer-based approaches and outperforms strong CNN-based baselines on the standard benchmark. In zero-shot transfer (trained on KITTI and evaluated without fine-tuning), JEPADepth achieves the best or near-best performance among the compared methods on both Make3D and Cityscapes across multiple metrics.
[42] Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution cs.CV | cs.AIPDF
Mingkuan Feng, Zhengqi Wen, Jianhua Tao
TL;DR: 本文提出了一种名为解耦视觉处理(DVP)的高效训练框架,用于多模态大语言模型(MLLM)的视觉指令微调。该方法通过将预训练LLM的上层解码器层替换为一个轻量级、专门处理视觉token的单Transformer块,从而在训练时仅需更新极少参数,显著降低了计算成本。
Details
Motivation: 动机在于,MLLM在深层网络中视觉与文本token的表征需求差异显著,对整个模型进行全参数微调计算代价高昂且通常不必要,因此需要一种参数高效的适配方法。
Result: 在LLaVA-1.5框架上的实验表明,DVP在MME、POPE和ChartQA基准测试上取得了具有竞争力的性能,同时仅训练了总参数的一小部分。
Insight: 创新点在于提出了一种解耦的视觉处理路径,将视觉与文本token在共享处理后的深层网络中进行分离处理,并通过仅更新一个新增的轻量级Transformer块来实现高效学习,这为MLLM的参数高效微调提供了新思路。
Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.
[43] Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach cs.CVPDF
Yinghao Hou, Jiahe Fan, Yuanhao Pu, Zongyuan Chen, Hong Xie
TL;DR: 本文提出了一种名为跨尺度定向参数注入(CDPI)的简单线性探针方法,用于分析异构多模态大语言模型(MLLM)在免训练融合过程中的知识转移机制。研究发现,跨尺度知识转移具有选择性,主要增益集中在高级推理能力上,而感知能力基本保持不变,这揭示了融合的本质是语言侧推理在特定低干扰区间的选择性转移,而非广泛的能力继承。
Details
Motivation: 现有研究对异构MLLM免训练融合中知识转移机制的理解不足,尤其是在评估任务范围扩大时,不清楚不同能力(如感知与推理)能否跨尺度转移。本文旨在探究跨尺度异构融合中知识转移的具体模式与选择性。
Result: 在四个Qwen3-VL模型对和十二个多模态基准测试上的实验表明,性能增益主要集中在推理(尤其是高级推理)任务上,而感知性能则接近原始目标模型水平。组件消融实验进一步表明,高级推理增益主要源于语言模型,比率分析发现正向选择性转移主要发生在小比率区间。
Insight: 创新点在于提出CDPI这一简单线性探针来量化分析知识转移,并通过理论分析和大量实验揭示了知识转移的选择性模式:融合主要实现语言侧推理能力在特定低干扰区间的定向转移,而非全面的能力继承,这为理解模型融合机制提供了新视角。
Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvements in aggregate performance do not reveal what a smaller model actually inherits. Existing studies are largely designed and evaluated on limited task sets or aggregate metrics; as evaluation expands to broader task collections, whether different capabilities can transfer across scales remains poorly understood. To investigate this question, we introduce Cross-Scale Directional Parameter Injection (CDPI), a simple linear probe to analyze cross-scale knowledge transfer during heterogeneous fusion. A local theoretical analysis indicates that knowledge transfer selectivity is determined at first order by capability-dependent responses to a shared injection direction, while second-order curvature effects constrain the effective transfer regime. Across four Qwen3-VL model pairs and twelve multimodal benchmarks, our experiments reveal a consistent pattern of selectivity: gains concentrate on reasoning, particularly high-level reasoning, whereas perception performance remains close to that of the original target model. Component-wise ablations further show that high-level reasoning gains arise primarily from the language model, while ratio analysis finds that positive selective transfer occurs mainly in the small-ratio regime. These findings recast cross-scale heterogeneous MLLM fusion as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance.
[44] Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation cs.CVPDF
Yongxin Su, Linjie Hou, Feng Wang, Jialin Tang, Zhijun Li
TL;DR: 本文提出Genie Sim PanoWorld,一种从单张360度全景图生成可自由导航的高保真3D场景的两阶段前馈流水线。该方法首先生成轨迹可控的全景视频,然后将其提升为3D高斯场景,支持实时自由视点漫游,并可直接用于具身AI应用。
Details
Motivation: 解决从单张全景图重建高保真、可自由导航3D场景的问题,避免现有方法存在的缺乏度量轨迹控制、大范围相机运动下遮挡处理困难、以及需要昂贵多GPU服务器或逐场景优化等限制。
Result: 在全景视频生成和下游3D重建任务上均优于几何条件基线方法,能够零样本泛化到未见过的室内场景,生成结果支持实时自由视点漫游。
Insight: 创新点在于通过显式的轨迹可控全景视频桥接生成与重建;采用NavMesh规划的SE(3)漫游轨迹注入潜在视频扩散模型,结合长短轨迹混合训练和基于捷径模型的自一致性目标,仅需4步无分类器引导去噪即可生成高保真视频;随后通过前馈全景重建器将视频提升为3D高斯场景,可直接用作仿真就绪资产。
Abstract: We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned $\mathrm{SE}(3)$ roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long–short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.
[45] Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time cs.CVPDF
Xinyi Wang, Yuyang Huang, Yalin Su, Pengcheng Luan, Tao Zhang
TL;DR: 本文提出了一种名为AnchorSteer的训练免费框架,旨在提升文本到图像扩散模型在推理时对复杂组合提示的忠实度。该框架通过语义锚定(Semantic Anchoring)和反思引导(Reflective Steering)两个协同组件,分别对初始噪声和去噪轨迹进行细粒度控制,以更好地利用预训练先验并实现生成过程中的自我纠正。
Details
Motivation: 当前文本到图像扩散模型在视觉质量上表现出色,但在处理复杂组合提示时经常出现对齐不精确的问题。现有无需训练的方法要么在改进初始噪声时效率低下或语义注入不可靠,要么在改进去噪轨迹时缺乏及时诊断和纠正语义错误的明确机制。
Result: 在GenEval和T2I-CompBench++基准上的大量实验表明,AnchorSteer在保持高视觉质量的同时,在文本-图像对齐方面持续优于现有基线方法。
Insight: 创新点在于提出了一个双组件框架,结合了基于CLIP先验提取和新型潜在先验分数蒸馏采样(LP-SDS)的语义锚定来优化初始化,以及一个基于VLM诊断的主动“思考-擦除-润色”循环(反思引导)来动态修正去噪轨迹,从而弥合了CLIP与扩散模型先验之间的领域差距并实现了细粒度的生成控制。
Abstract: While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositional prompts. An effective strategy is to improve the inference process of diffusion models, thereby better leveraging their pretrained priors to address misalignment. Existing training-free methods can be divided into two categories. The first category focuses on improving the randomly sampled initial noise, either performing costly search over noise pools or manipulating sampled noise without ensuring reliable semantic injection. The second category focuses on improving the denoising trajectory, lacking explicit mechanisms to timely diagnose and correct semantic errors. we propose \textbf{AnchorSteer}, a training-free framework that exerts fine-grained control over \textbf{both initialization} and \textbf{the denoising trajectory}. AnchorSteer consists of two synergistic components: \textbf{Semantic Anchoring} replaces uninformative Gaussian noise with text-aligned initializations via CLIP-based prior extraction and a novel Latent-Prior Score Distillation Sampling (LP-SDS) objective. Specifically, LP-SDS distills CLIP visual priors into the knowledge distribution of diffusion models, mitigating the domain gap between CLIP-based priors and diffusion-based priors. \textbf{Reflective Steering} transforms passive denoising with an active Think–Erase–Retouch loop that enables mid-generation self-correction. It leverages VLM-based diagnosis to detect semantic deviations and performs targeted latent refinement to suppress erroneous content and recover missing attributes. Extensive experiments on GenEval and T2I-CompBench++ demonstrate that AnchorSteer consistently outperforms existing baselines in text–image alignment while preserving high visual quality.
[46] Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation cs.CVPDF
Xiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu
TL;DR: Visko Orbis 1.0是一个用于实时交互式长视频生成的’直播模型’。它允许用户在生成过程中随时更改提示词,并实时看到更新效果,支持长文本到视频、图像到视频和视频续写,并能保持多小时的生成质量。
Details
Motivation: 解决现有视频生成系统无法在生成过程中进行实时、交互式控制的问题,旨在实现用户可随时干预、提示词可切换的长视频实时生成。
Result: 在长视频竞技场比较中,Visko Orbis 1.0在整体偏好度和时间稳定性方面获得了最高评分,达到了最先进的实时交互式视频生成系统的水平。
Insight: 其核心创新在于结合了有界多尺度记忆机制来跨视频片段保持主体、场景和风格一致性,以及基于蒸馏的分块流式生成器和流式视频超分架构,实现了实时4K视频生成。
Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers real-time 4K video generation at 24 FPS using an optimized GPU serving engine. In long-form Arena comparisons, Visko Orbis 1.0 obtains the highest overall-preference and temporal-stability ratings among state-of-the-art real-time interactive video-generation systems.
[47] TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models cs.CVPDF
Taewon Kang, Matthias Zwicker
TL;DR: 本文提出了一种名为TPD(Temporal Prior Decoupling)的训练免费框架,用于解决文本到视频扩散模型中存在的时序先验抑制问题。该方法通过构建时序反事实并定义被抑制的信号方向,在扩散采样过程中恢复后期事件信号,从而在不破坏早期片段连贯性的情况下,在视频后期帧中实现被抑制的事件。
Details
Motivation: 解决文本到视频扩散模型在生成包含时序变化(如早期场景持续而新事件在后期出现)的复杂提示时,经常无法在对应帧中实现后期事件的问题,即识别出的时序先验抑制现象。
Result: 实验表明,TPD显著改善了后期概念的实现,同时保持了时序连贯性和视觉保真度,并且这种针对性抑制问题在不同文本到视频骨干模型中普遍存在。
Insight: 创新点在于提出了一种训练免费、骨干模型无关的框架,通过定义并恢复被抑制的信号方向(而非像先前减法投影方法那样移除它),并引入基于扩散时间步和视频帧联合解析的帧选择性下界约束,来保证被抑制信号的贡献,从而有效解决时序先验抑制问题。
Abstract: Text-to-video diffusion models generate temporally coherent content from natural language, yet when a prompt describes an early scene that persists while a new event emerges on top of it—such as “a tall sandcastle standing on a beach where a wave rushes in and washes it away”—generation frequently fails to realize the late-segment event in the corresponding frames. We identify this failure as Temporal Prior Suppression (TPS): the dominant prior of the early segment captures the cross-attention trajectory across the temporal axis and suppresses the guidance signal needed for late-segment realization, a competing tendency existing guidance mechanisms do not model. We introduce Temporal Prior Decoupling (TPD), a training-free framework that restores suppressed late-segment signals during diffusion sampling. TPD constructs a temporal counterfactual by conditioning on the early segment alone, and defines the discrepancy between the full-prompt and counterfactual trajectories as a suppressed signal direction. Rather than removing this direction as in prior subtractive projection methods, TPD restores it through a frame-selective lower-bound constraint resolved jointly over diffusion timestep and video frame, realizing the suppressed event in the late frames without disrupting early-segment coherence: where prior work enforces upper-bound feasibility to remove unwanted semantics, TPD enforces lower-bound feasibility to guarantee suppressed-signal contribution. TPD runs entirely within standard diffusion sampling without retraining, and is defined purely in classifier-free guidance space, making it backbone-agnostic by construction. Experiments show that TPD significantly improves late-concept realization while preserving temporal coherence and visual fidelity, and that the targeted suppression recurs across distinct text-to-video backbones.
[48] Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives cs.CV | cs.AI | cs.MMPDF
Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng, Dayong Ye
TL;DR: 本文提出了一种名为Dualin的双阶段反演方法,用于文本到图像扩散模型的逆向工程。该方法通过联合恢复目标图像的语义提示词和潜在噪声,解决了现有提示词反演方法在稳定性和视觉保真度方面的不足。
Details
Motivation: 现有提示词反演方法存在显著局限:基于梯度的方法不稳定且难以解释,常导致生成图像存在严重伪影;而无梯度方法虽能产生人类可读的提示词,但因缺乏细粒度细节对齐而无法保持视觉保真度。作者认为这些局限源于将提示词反演视为逆向工程的充分条件,而忽略了编码结构信息的潜在噪声的关键作用。
Result: 在多个数据集上的广泛实验表明,Dualin能同时生成高质量的倒置提示词,并在图像保真度方面达到了最先进的水平。
Insight: 核心创新在于将逆向工程视为一个双变量问题,联合优化提示词和潜在噪声。具体地,第一阶段利用视觉语言模型CLIP和大语言模型反演出忠实、可解释的硬提示词;第二阶段通过无条件DDIM反演精确重建目标图像的潜在噪声,保证了结构信息层面的一致性。理论证明,反演出的噪声支持无需重新优化的灵活图像编辑,为精确可控的图像编辑建立了稳健基础。
Abstract: Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion methods suffer from significant limitations: (1) gradient-based methods are unstable and uninterpretable, often resulting in generated images with severe artifacts; (2) gradient-free methods yield human-readable prompts but still fail to preserve visual fidelity due to the lack of fine-grained detail alignment. We contend that the limitations stem from treating prompt inversion as a sufficient condition for reverse engineering, ignoring the critical role of the latent noise that encodes structural information. Consequently, we propose Dualin (Dual inversion), a two-stage method that jointly recovers both the semantic prompt and latent noise of the target image. In the first stage, we integrate vision-language model, CLIP and large language model to invert a faithful, human-interpretable hard prompt. In the second stage, unconditional DDIM inversion reconstructs the exact latent noise of the target image, guaranteeing the consistency at the structural information level. Theoretically, we prove that the inverted noise enables flexible image editing without re-optimization. Extensive experiments on diverse datasets demonstrate that Dualin simultaneously generates high-quality inverted prompts and achieves state-of-the-art image fidelity. Additionally, Dualin can establish a robust foundation for the precise and controllable image editing.
[49] Multimodal fusion of visual and morphometric features for avian bone classification cs.CV | cs.AIPDF
Nevio Dubbini, Lisa Yeomans, Marco Pavia, Ramazan Parmaksiz, Ayse Atas Hooglugt
TL;DR: 本研究提出了一种多模态融合框架,结合基于卷积神经网络的图像分析和骨测量数据,用于鸟类骨骼的分类。该框架在骨骼元素识别任务上达到86%的测试准确率,在科级分类任务上达到51%的top-1准确率,证明了视觉与形态测量信息在深度学习框架中融合的可行性。
Details
Motivation: 人工智能在考古学中的应用潜力巨大,但在动物考古学中,尤其是鸟类骨骼遗骸的识别方面仍很有限,本研究旨在解决这一问题。
Result: 在超过10,000张图像的数据集上,骨骼元素分类的测试准确率达到86%;科级分类更具挑战性,top-1准确率为51%,top-3准确率为75%,表明正确分类常出现在最可能预测中。
Insight: 创新点在于提出了一个特征级的多模态架构,融合了预训练EfficientNet_V2_S提取的视觉特征与标准化的形态测量数据,并采用BiRefNet和SAM2的两阶段管道进行自动图像分割,为AI辅助的动物考古识别建立了方法学基线。
Abstract: Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, particularly for the identification of avian skeletal remains. This study presents a proof-of-concept multimodal framework that integrates convolutional neural network-based image analysis with osteometric measurements for the classification of bird bones. Using a dataset of more than 10,000 images from multiple museum and research collections, two classification tasks were investigated: skeletal element identification and family-level taxonomic classification. Prior to classification, images were automatically segmented using a two-stage pipeline combining BiRefNet and SAM2. Visual features extracted with a pre-trained EfficientNet_V2_S backbone were fused with standardized morphometric data through a feature-level multimodal architecture. The model achieved 86% accuracy on the test set for bone-type classification, demonstrating reliable recognition of skeletal elements. Family-level classification proved more challenging, reaching 51% top-1 accuracy but 75% top-3 accuracy, indicating that correct taxa were frequently included among the most probable predictions. These results demonstrate the feasibility of combining visual and morphometric information within a unified deep-learning framework and establish a methodological baseline for future AI-assisted zooarchaeological identification. The approach contributes to ongoing efforts to develop scalable, interpretable, and archaeologically meaningful tools for the study of avian remains.
[50] Long-Tailed 3D Point Cloud Dataset Distillation cs.CVPDF
Jiahao You, Xu Han, Jinfeng Xu, Xianzhi Li
TL;DR: 本文提出了首个针对长尾3D点云数据集蒸馏的研究,通过自适应合成预算分配和3D长尾分布匹配两个核心模块,在保持数据集训练效用的同时,显式地处理了类别分布不平衡问题。该方法在ShapeNet55基准测试上比现有最优方法提升了7.0个百分点的分类准确率。
Details
Motivation: 现有3D点云数据集蒸馏方法主要关注几何和表示挑战,忽略了点云数据集中普遍存在的长尾类别分布不平衡问题,导致训练和测试集均遵循长尾分布,影响了模型性能。
Result: 在ShapeNet55基准测试上,该方法将分类准确率提升了7.0个百分点,显著优于现有最先进方法,证明了其在长尾3D点云数据集蒸馏中的有效性。
Insight: 创新点在于首次将长尾分布问题引入3D点云数据集蒸馏,通过自适应合成预算分配和3D长尾分布匹配(包括全局-局部特征对齐和先验感知监督)来优化合成点云,既保持了全局类别分布和类内结构多样性,又通过类别依赖的专家监督确保尾类样本可识别且头类模式多样。
Abstract: Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving their training utility, enabling efficient 3D point cloud training. Current point cloud dataset distillation methods only tackle geometric and representation challenges while ignoring the distributional imbalance prevalent in point cloud datasets where both training and test splits follow long-tailed class distributions. To our knowledge, we present the first study on long-tailed point cloud dataset distillation. Rather than focusing primarily on geometric and representation properties or simply constructing a class-balanced synthetic set, our framework explicitly accounts for long-tailed class distributions via two core modules. First, we design Adaptive Synthetic Budgeting to allocate class-wise synthetic budgets according to class quantity and the expected benefit of additional synthetic samples. Given the allocated budgets, we further design 3D Long-Tailed Distribution Matching to optimize synthetic point clouds through Global-Local Feature Alignment and Prior-Aware Supervision. The former preserves both global class distributions and diverse intra-class structures, while the latter provides class-dependent expert supervision to keep tail-class samples recognizable while maintaining diverse head-class patterns. Extensive experiments demonstrate the effectiveness of our method, lifting classification accuracy by 7.0 points on ShapeNet55 against state-of-the-art methods.
[51] See2Think: Do Multimodal Models Really Use Intermediate Visual States? cs.CV | cs.AIPDF
Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen
TL;DR: 本文提出了See2Think评估框架,包括See2ThinkBench基准和Visual Action-of-Thought评估方法,用于系统评估多模态大语言模型是否真正利用中间视觉状态进行推理。研究发现,模型的视觉推理能力高度依赖于模型本身和环境设置,且忠实渲染是主要瓶颈。
Details
Motivation: 现有基准测试在任务覆盖范围上有限,且评估过于关注最终答案,缺乏对中间视觉状态(如草图、注释、工具生成图像)如何被生成、渲染和使用的诊断,因此无法确定多模态模型是否真正依赖这些视觉状态进行推理。
Result: 在包含12个任务类别、1200个开放性问题的新基准See2ThinkBench上评估了代表性的专有和开源多模态模型。结果显示,没有一种推理设置能在所有任务中持续占优;在任务相关的干扰反馈下,模型表现出对视觉状态的行为依赖,准确率在受控干预下下降超过10个百分点。
Insight: 创新点在于提出了一个统一的评估框架,通过记录文本思考、视觉动作、渲染状态和后续推理来诊断模型对中间视觉状态的利用。客观分析表明,该研究揭示了模型视觉推理的瓶颈(忠实渲染)以及模型行为对视觉状态的依赖性,为理解多模态模型的内部推理机制提供了新工具和洞见。
Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
[52] DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation cs.CVPDF
Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, Junyao Gao
TL;DR: 本文提出DistillAlign方法,用于协调自回归视频蒸馏中的模式覆盖与模式寻求。现有方法通常将初始化与分布匹配蒸馏阶段解耦,并仅依赖视觉评分评估学生模型,导致分布对齐不足。通过引入分布评估协议和联合蒸馏目标,该方法在生成质量、覆盖率和多样性上均取得提升。
Details
Motivation: 现有自回归视频蒸馏方法将初始化与分布匹配蒸馏阶段解耦,分别追求不同目标分布,且仅通过视觉评分评估中间学生模型,忽略了分布对齐的重要性,导致次优精炼结果。
Result: 实验表明,该方法在生成质量、覆盖率和多样性上均优于基线;即使使用Wan-1.3B DMD教师模型,其性能也超过使用Wan-14B精炼的基线,突显了分布对齐的关键作用。
Insight: 创新点在于从分布视角重新审视自回归视频蒸馏,提出分布评估协议以揭示视觉评分隐藏的差异,并设计联合蒸馏目标结合模式寻求与模式覆盖约束,确保分布对齐,从而提升蒸馏效果。
Abstract: Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages – which then pursue different target distributions – and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD’s reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD’s mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.
[53] Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory cs.CVPDF
Yanbo Ding, Zhizhi Guo, Quanyue Song, Yishan He, Zhixiang He
TL;DR: 本文提出了Ripple,一种用于实时流式音视频生成的系统,通过跨模态循环记忆机制解决现有方法延迟高、成本高且难以生成长视频的问题。Ripple结合固定长度滑动窗口注意力与模态特定记忆状态来高效处理流式推理并保持长期上下文,并通过跨模态记忆交互增强音视频同步。采用三阶段训练策略,最终在480P分辨率下达到约28 FPS,比教师模型快得多,并能实现连贯的长视频生成。
Details
Motivation: 现有音视频生成模型虽然质量高,但延迟大、成本高,且难以支持实时应用和长视频生成,因此需要开发一种低延迟、高效的流式生成系统。
Result: 在短视频和长视频基准测试中,Ripple在性能上优于现有的离线和在线联合音视频生成方法,在480P分辨率下实现约28 FPS的实时生成速度。
Insight: 创新点包括跨模态循环记忆机制以平衡效率与长期上下文,以及三阶段训练策略(包括因果注意力适应、端到端蒸馏和在线强化后训练),这些方法可借鉴于其他流式多模态生成任务。
Abstract: Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video generation methods have been proposed, they remain costly and fail to support long-form generation. To address this, we propose \textbf{Ripple}, a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism. To enable efficient streaming inference while preserving long-term context, Ripple combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context. Cross-modal memory interaction is further introduced to enhance audio-visual synchronization. To learn this memory-augmented model effectively, we devise a three-stage training recipe: (1) adapting a bidirectional audio-video teacher to block-wise causal attention with simulated memory, (2) optimizing the memory construction and interaction pipeline through end-to-end distillation, and (3) applying online reinforcement post-training tailored for streaming audio-video generation. As a result, Ripple achieves ~28 FPS at 480P resolution, over faster than the teacher, while capable of coherent long-form generation. Extensive experiments on both short-video and long-video benchmarks demonstrate our superior performance over existing offline and online joint audio-video generation methods.
[54] ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures cs.CVPDF
Fahad Ahmed, Sören Auer, Jennifer D’Souza
TL;DR: 本文介绍了ICDAR 2026关于从原子层沉积/刻蚀(ALD/E)科学图表中提取信息的竞赛。该竞赛基于Sci-ImageMiner基准数据集,该数据集包含四个端到端的互补任务,并由专家标注。竞赛吸引了68名活跃参与者,提交了1263份结果。结果表明,当前最先进的多模态模型在分类和摘要任务上表现良好,但在数据提取和科学推理(尤其是视觉问答)方面存在困难。
Details
Motivation: 动机是推动科学图表理解与推理的多模态AI研究,通过整合视觉感知与领域特定推理,从研究出版物中提取文本未呈现的有意义知识。
Result: 竞赛结果显示,最先进的多模态模型在分类和摘要任务上表现良好,但在数据提取和科学推理(尤其是视觉问答)任务上表现不佳,揭示了现有模型的局限性。
Insight: 创新点在于创建了一个全面的、专家标注的Sci-ImageMiner基准数据集和社区驱动的竞赛,为科学图表理解与推理研究建立了严格的评估平台,并明确了领域感知多模态AI系统面临的挑战与机遇。
Abstract: Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
[55] Hearsay: Vision-Language Medical Diagnoses Without an Image cs.CV | cs.AI | cs.CL | cs.CYPDF
Siddharth Vohra
TL;DR: 这篇论文研究了前沿视觉语言模型在缺乏医学图像输入时,仅根据患者人口统计学描述就生成虚假诊断的‘幻觉’现象。研究发现,Claude Opus-4.7、GPT-5.4和Gemini-3.1-Pro等模型会系统性地根据患者年龄、种族和性别捏造诊断,例如年轻黑人患者常被诊断为结节病。
Details
Motivation: 动机是揭示并分析视觉语言模型在临床诊断场景中,当缺乏关键视觉信息(医学图像)时,不仅不拒绝回答,反而会根据无关的人口统计学信息产生结构化、有偏见的虚假诊断,这对临床部署的可信度构成严重威胁。
Result: 研究在胸部X光、脑部MRI和皮肤病学等多个医学领域进行了测试,发现模型输出存在系统性偏差。例如,Claude Opus-4.7对65岁白人男性‘痣’的查询几乎每次都诊断为黑色素瘤,而GPT-5.4在所有测试的人口统计单元中都产生了捏造诊断。
Insight: 创新点在于揭示了模型‘幻觉’是一种由不同故障模式构成的家族现象,而非单一问题,并指出仅审核文本描述不足以发现结构化诊断字段中的错误。核心洞见是,可靠的临床部署需要直接审计结构化输出通道,并将‘探针词敏感性’作为首要评估维度。
Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads “suspected, based on demographics and classic pattern.’’ GPT-5.4’s effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude’s dermatology effect collapses entirely when ‘skin mole’ is swapped for ‘skin lesion’ while GPT-5.4’s is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension
[56] SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation cs.CVPDF
Yunzhan Fu, Enyu Bao, Xiangyu Shen, Yihao Wu, Chunbo Jiang
TL;DR: 本文提出了SCALPEL框架,旨在解决医学视觉语言预训练中因整合大型语言模型而带来的表示坍缩、内存开销和医学幻觉问题。该方法通过临床报告对比微调将生成式LLM转化为各向同性编码器,采用非对称对齐策略进行高效训练,并引入解剖-否定感知目标来惩罚涉及方位混淆或错误否定的不匹配图像-文本对。
Details
Motivation: 现有医学VLP框架在处理冗长、术语密集的临床报告时,受限于轻量级文本编码器的有限上下文窗口和浅层表示能力。整合医学LLM虽能提供前所未有的临床推理能力,但会引入表示坍缩、巨大内存开销和医学幻觉三大瓶颈。
Result: 在MIMIC-CXR、CheXpert和IU X-Ray基准测试上的广泛实验表明,SCALPEL在跨模态检索、零样本疾病分类和医学视觉问答任务上实现了最先进的性能。
Insight: 创新点在于将生成式LLM通过领域特定的临床文本适应转化为各向同性编码器,采用离线特征缓存的非对称对齐策略以降低内存需求,并设计了能明确惩罚方位混淆和错误否定的解剖-否定感知对比损失函数,从而提升了医学多模态表示的语义对齐精度。
Abstract: Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational capacities of lightweight text encoders when processing lengthy, terminology-dense clinical reports. While integrating medical large language models (LLMs) offers unprecedented clinical reasoning capabilities, it introduces three major bottlenecks: (i) the anisotropic representational collapse of generative LLMs under standard contrastive objectives, (ii) the prohibitive memory overhead of joint end-to-end training with large batch sizes, and (iii) the medical hallucinations induced by vanilla contrastive losses that ignore fine-grained anatomical laterality and negation modifiers. To address these challenges, we propose \textbf{SCALPEL}, a \textbf{S}emantic \textbf{C}ross-modal \textbf{A}lignment framework via \textbf{L}LM-\textbf{P}owered \textbf{E}ncoder \textbf{L}earning. First, Clinical Report Contrastive fine-tuning converts a generative LLM into an isotropic encoder via domain-specific clinical text adaptation. Second, an asymmetric alignment strategy leverages offline feature caching to enable efficient training. Critically, we formulate an Anatomy-Negation Aware Objective that explicitly penalizes mismatched image-text pairs involving laterality confusion or false negations. Extensive experiments across MIMIC-CXR, CheXpert, and IU X-Ray benchmarks demonstrate that SCALPEL achieves state-of-the-art performance in cross-modal retrieval, zero-shot disease classification and medical visual question answering.
[57] CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents cs.CVPDF
Qianru Li, Xuyang Chen, Erkin Türköz, Lu Liu, Xuqin Wang
TL;DR: 本文提出了CinemaTraj框架,通过LLM智能体将自然语言描述转化为3D场景中的电影化摄像机轨迹。该方法利用结构化3D场景图,将用户指令分解为一系列原子摄像机运动(如推拉、环绕、升降等),并通过一种新颖的参数化轨迹表示进行实例化,同时生成同步的旁白和字幕。
Details
Motivation: 现有方法要么依赖2D图像先验而缺乏真正的3D空间感知,要么将轨迹生成视为脱离电影语义的几何路径规划问题。本文旨在解决从自然语言描述自动生成具有电影表现力的3D摄像机轨迹这一具有高实用价值的挑战性任务。
Result: 在真实世界的ScanNet++环境上评估表明,CinemaTraj在提示对齐度、轨迹质量和安全性指标上均优于现有方法,能生成符合提示、无碰撞且具有高电影质量的轨迹。
Insight: 核心创新在于将摄像机轨迹规划重新定义为基于语言的空间推理问题,并引入了结合电影表现力与可优化性的参数化轨迹表示。利用3D场景图作为结构化空间先验,使LLM智能体的推理能够基于环境的精确几何和语义知识。
Abstract: Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent’s reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.
[58] Prior Directions: Why GUI Grounding Gets Locked in the Past cs.CVPDF
Weile Gong, Zijian Lu, Mingcai Chen, Yiping Zuo, Xin He
TL;DR: 该论文研究了视觉语言模型在视觉锁定现象中的表现,即模型在场景变化时仍依赖过时的语言描述做出错误判断。通过控制实验,作者发现锁定强度与模型表示变化的方向性相关,并识别出可复现的’先验方向’轴,移除这些方向上的成分可恢复视觉基础能力。
Details
Motivation: 解决视觉语言模型因依赖过时语言描述而导致的视觉锁定问题,即模型在场景动态变化时无法正确更新视觉判断,从而影响基础任务的可靠性。
Result: 在受控基础设置中,多个模型表现出不同程度的视觉锁定,锁定强度与模型表示变化的方向集中度相关;移除先验方向成分的实验恢复了视觉基础能力,而移除正交成分则影响甚微。
Insight: 创新点在于提出’先验方向’概念,揭示视觉锁定源于先验诱导的表示变化沿特定可复用方向集中,而非表示变化幅度;这为理解模型错误机制和设计纠正方法提供了新视角。
Abstract: Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the verbalized prior varies. Across models, stronger lock-in accompanies smaller changes in the model representation before the final answer. This reversal suggests that lock-in depends not on how far this representation moves, but on how that movement is organized. In models that are harder to correct, prior-induced changes concentrate along a compact set of directions that repeatedly appear across examples. We call these recurrent axes the Prior Directions. They recur on held-out examples, while a descriptive four-model comparison associates greater concentration with stronger lock-in. Controlled interventions show that removing the component aligned with the Prior Directions restores visual grounding, whereas removing an equally large orthogonal component has little effect. Prior control thus arises when prior-induced changes form a coherent and reusable pattern in the representation used to produce the answer. This account explains why the same prior remains revisable in one model yet becomes dominant in another.
[59] From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models cs.CVPDF
Alexej Klushyn, Juan Rivero Sesma, Florian Seligmann, Richard Kurle, Kinh Tieu
TL;DR: 本文提出了一种轻量级后验概率扩展方法,为已训练的YOLO-Pose模型增加关键点位置的空间不确定性量化能力。该方法通过训练额外的概率头来预测每个关键点的输入相关协方差矩阵,并采用高斯或Student-t分布进行校准,从而生成校准后的二元预测分布。在COCO数据集上的实验表明,该方法能实现有效的关键点级可靠性排序,并支持基于不确定性的关键点剪枝。
Details
Motivation: YOLO-Pose模型虽能高效定位关键点,但缺乏对空间不确定性的量化能力,限制了其在需要可靠性评估的下游任务(如传感器融合)中的应用。
Result: 在COCO数据集上,该方法实现了关键点级可靠性排序,Student-t校准能最佳捕捉经验残差分布,基于不确定性的剪枝可移除不可靠关键点;在飞机着陆视觉应用中,校准后的协方差支持不确定性感知的飞机位置估计。
Insight: 创新点包括:1)为YOLO-Pose设计轻量级后验概率扩展框架;2)提出结合分布校准诊断与平均关键点精度(AKP)的评估协议;3)通过高斯/Student-t校准实现下游任务兼容性与分布保真度的平衡。
Abstract: YOLO-Pose models provide efficient keypoint localization, but do not quantify the associated spatial uncertainty. We introduce a lightweight post-hoc probabilistic extension that augments a trained YOLO-Pose model with calibrated bivariate predictive distributions over keypoint locations, centered at the model’s original predictions. Concretely, we train additional probabilistic heads with an importance-weighted negative log-likelihood to predict an input-dependent $2\times2$ dispersion matrix for each keypoint, followed by Gaussian calibration for broad downstream compatibility or Student-$t$ calibration for distributional fidelity. Complementing this, we propose an evaluation protocol that combines a suite of distributional calibration diagnostics with average keypoint precision (AKP), a keypoint-level extension of the COCO AP protocol for assessing reliability rankings. Experiments on COCO show that the learned uncertainty estimates enable effective keypoint-level reliability ranking, Student-$t$ calibration best captures the empirical residual distribution, and uncertainty-based pruning removes unreliable keypoints. A central application-level demonstration is vision-based aircraft landing, where calibrated covariances for runway keypoints support uncertainty-aware aircraft position estimation and downstream sensor fusion.
[60] Mitigating Compounding Error via Video Representation Regularization cs.CV | cs.LGPDF
Taiye Chen, Qi Zhang, Yisen Wang
TL;DR: 本文研究了基于视频扩散的世界模型在自回归长视频生成中存在的误差累积问题,发现误差累积与模型内部表示的维度坍缩紧密相关,并提出了一种轻量级的视频表示正则化方法以抑制误差漂移,从而提升长序列生成的稳定性。
Details
Motivation: 基于视频扩散的世界模型在机器人、自动驾驶和仿真等任务中可实现长序列自回归视频生成,但滑动窗口自回归推理会遭受严重的误差累积,导致帧质量随时间下降,而这一现象的内在机制和稳定生成长序列的方法尚未得到充分解决。
Result: 在VBench基准测试中,所提方法在美学质量和图像质量指标上分别从38.65提升至55.56和从44.37提升至72.08,优于Diffusion Forcing方法。
Insight: 创新点在于首次建立了自回归视频漂移与模型内部表示之间的联系,采用有效秩作为误差累积的量化指标,揭示了视频世界模型反直觉的缩放限制,并提出了一种简单有效的正则化策略来增强长视频生成的鲁棒性。
Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.
[61] Progressive Multimodal Alignment for Continual Instruction Tuning cs.CV | cs.AIPDF
Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
TL;DR: 本文提出了渐进式多模态对齐(PMA)框架,用于解决多模态持续指令调优(MCIT)中的投影器级遗忘问题。PMA通过轻量级表示描述器检测多模态分布变化,仅在需要时渐进扩展投影器专家,并利用可扩展路由器集成专家输出,同时保留预训练投影器作为稳定对齐锚点。该框架作为即插即用模块,与现有MCIT方法结合后,在多个基准测试中取得了优于先前SOTA方法的性能。
Details
Motivation: 多模态大语言模型(MLLMs)依赖投影器对齐视觉与语言表示,但在MCIT场景中,视觉分布变化和指令语义演化会导致共享投影器发生漂移,引发投影器级遗忘问题,而现有方法主要关注LLM主干而忽视了这一关键问题。
Result: 在两个最新的MCIT基准测试上进行广泛实验,结果表明,结合PMA缓解投影器级遗忘后,相比先前最先进方法取得了持续的性能提升,且PMA能够泛化到不同的MLLM主干模型,展现出鲁棒且广泛适用的MCIT性能。
Insight: 创新点在于首次明确识别并系统解决了MCIT中的投影器级遗忘问题,通过渐进式专家扩展机制平衡稳定性和可塑性,并以亚线性参数增长实现高效持续适应;从客观角度看,其轻量级分布检测和模块化设计为多模态持续学习提供了可扩展且方法无关的解决方案。
Abstract: Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Alignment (PMA), a framework that enables the projector to adapt continually while preserving previously learned alignment. PMA detects multimodal distribution shifts via a lightweight representation descriptor and progressively expands projector experts only when needed. An expandable router integrates expert outputs based on multimodal features, while the original pretrained projector is retained as a stable alignment anchor. This progressive mechanism balances stability and plasticity with sub-linear parameter growth and serves as a method-agnostic add-on to existing MCIT approaches. Extensive experiments on two recent MCIT benchmarks demonstrate that mitigating projector-level forgetting yields consistent gains over prior state-of-the-art methods when combined with PMA. Moreover, PMA scales across diverse MLLM backbones, demonstrating robust and broadly applicable MCIT performance.
[62] SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence cs.CV | cs.AIPDF
Chuanzhi Xu, Zihan Deng, Huiqi Liang, Chengkun Yue, Zhanlin Cui
TL;DR: 本文提出了SciFigAlign,一种针对科学图表质量评估的微调多模态评分模型。该方法通过结合图表、标题、引用段落及论文上下文,利用CLIP和SciBERT进行端到端微调,并引入跨模态注意力与CubeMLP融合机制,以回归和排序损失联合优化,显著提升了科学图表在清晰度、相关性、信息量和结构四个维度的评估性能。
Details
Motivation: 现有图像评估方法(如传统IQA模型、CLIP-based方法或零样本LLM/VLM)在科学图表评估中存在局限:无法判断图表是否支持论文的科学论点、缺乏对稿件上下文的理解,或导致分数过于集中且视觉与文本证据融合不足。
Result: 在包含3,857个科学图表的标注数据集上,采用论文级划分,SciFigAlign在测试集(n=396)上实现了宏观MAE为0.3524和论文内配对准确率为81.64%,相比最佳LLM-as-judge基线(MAE 0.864)相对误差降低了59%。
Insight: 创新点在于构建了面向同行评审的科学图表评估数据集,并提出了基于稿件证据的微调对齐模型,通过引用上下文去噪和排序监督,强调了科学图表评估需要学习视觉内容与稿件证据之间的对齐,而非仅依赖提示工程,即使对于先进VLM也是如此。
Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript’s claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper’s scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript context; and zero-shot LLM/VLM judges, when repurposed for figure scoring, often yield overly concentrated scores with limited fusion of visual and textual evidence. We introduce an annotated dataset of 3,857 scientific figures from peer-reviewed conference papers, each rated along four peer-review-oriented dimensions: Clarity, Relevance, Informativeness, and Structure. We propose SciFigAlign, a fine-tuned multimodal scorer that grounds figure quality assessment in manuscript evidence. Given a figure crop, caption, citing paragraphs, and light paper context, SciFigAlign fine-tunes CLIP and SciBERT end-to-end with per-modality cross-attention and CubeMLP fusion, jointly optimizing SmoothL1 regression with a within-paper ranking hinge loss. Under paper-level splits, SciFigAlign achieves a macro MAE of 0.3524 and a within-paper pairwise accuracy of 81.64% on the test set with n = 396, a 59% relative error reduction over the best LLM-as-judge baseline with MAE 0.864. Ablations confirm that manuscript-grounded inputs, citing-context denoising, and ranking supervision are all critical, showing that scientific figure assessment requires learned alignment between visual content and manuscript evidence rather than prompting alone, even with state-of-the-art VLMs.
[63] FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring cs.CVPDF
Jiatong Li, Leo Liang, Linghe Kong, Yulun Zhang
TL;DR: 本文提出了一种名为FreqForcing的无训练框架,通过谱自锚定(SSA)技术解决自回归视频扩散模型在生成长视频时出现的误差累积问题。该方法从频域角度分析误差表现为低频能量漂移,并利用锚注意力的低频分量保持视觉稳定性,结合局部注意力的高频分量保留动态运动。
Details
Motivation: 自回归视频扩散模型在实时流式视频生成中,自展开过程中引入的误差会随时间累积,导致颜色漂移、运动停滞和最终视觉崩溃。本文从频域视角将此现象表征为低频带的显著能量漂移,旨在解决长视频生成的稳定性问题。
Result: FreqForcing将基于5秒片段预训练的Self-Forcing模型扩展到两分钟生成,实现了24倍的外推。大量实验表明,该方法在定量和定性上优于现有的无训练方法,并与基于训练的代表性方法保持竞争力。
Insight: 创新点在于从频域角度分析误差累积机制,并提出谱自锚定(SSA)这一无训练框架,通过分离低频和高频注意力分量来同时维持长时程视觉稳定性和动态运动,为长视频生成提供了新的解决方案。
Abstract: Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.
[64] Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection cs.CVPDF
Hao Tan, Jun Lan, Zichang Tan, Ajian Liu, Zijian Yu
TL;DR: 本文提出Veritas++,一种感知增强的AI生成图像检测框架,通过感知导向学习和价值感知策略蒸馏来提升多模态大语言模型在细粒度异常感知上的能力,从而提高检测的泛化性和鲁棒性。
Details
Motivation: 当前基于MLLM的AIGI检测器在捕捉细粒度异常时存在感知瓶颈,主要关注视觉证据的组织与合成,而内在感知能力未充分优化,需要建立可靠的感知作为真实性推理的基础。
Result: 在标准、野外和新兴基准测试上的广泛实验表明,Veritas++实现了有前景的泛化性能,感知学习有效弥合了感知差距并带来无缝的检测增益,而VaOPD进一步实现了高效的能力演进且不牺牲现有性能。
Insight: 创新点包括将AIGI检测建立在三种基本感知能力(细粒度视觉细节、语义异常和像素级差异)上,提出感知导向学习用可验证奖励替代开放式描述监督,以及价值感知策略蒸馏通过特权自教师机制优先处理高价值蒸馏信号以内部化感知感知推理。
Abstract: The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection increasingly essential. While multi-modal large language models (MLLMs) offer a transparent alternative to black-box binary scoring, we observe that current MLLM-based detectors still exhibit notable perception bottlenecks in capturing fine-grained anomalies. They primarily focus on how visual evidence is organized and synthesized, leaving the intrinsic perception less optimized. To mitigate this gap, we present Veritas++, a perception-enhanced reasoning framework that establishes reliable perception as the foundation of authenticity reasoning. Rather than directly optimizing the model’s explanatory ability, we ground AIGI detection on three basic perception abilities, i.e., capturing fine-grained visual details, semantic anomalies and pixel-level differences. Building on this insight, we introduce Perception-oriented Learning (PoRL), which replaces open-ended description supervision with verifiable rewards to explicitly strengthen these capacities. To further integrate enhanced perception with reasoning, we introduce Value-aware On-Policy Distillation (VaOPD), an adaptive distillation mechanism that prioritizes high-value distillation signals over uniform supervision, internalizing perception-aware reasoning through a privileged self-teacher. Extensive experiments across standard, in-the-wild and emerging benchmarks demonstrate that Veritas++ achieves promising generalization. The perception learning effectively bridges the perception gap and yields seamless gains on detection, while VaOPD further enables efficient capability evolvement without sacrificing existing performance. Code and checkpoints are available at https://github.com/EricTan7/VeritasPP.
[65] SeasonStereo: Robust Dense Stereo Matching for Multi-Date Satellite Imagery via Generative AI cs.CVPDF
Álvaro Díaz-Laureano, Roger Marí, Elías Masquil, Pablo Arias, Gabriele Facciolo
TL;DR: SeasonStereo是一个用于多日期卫星图像稳健密集立体匹配的框架,通过生成式AI合成具有可控季节性外观变化的图像对进行训练,并利用基础模型的零样本几何先验。该框架无需对齐的真实多日期训练数据或LiDAR标签,就能达到最先进的LiDAR监督模型的精度,并产生更清晰的几何细节。
Details
Motivation: 解决从多日期卫星图像进行三维重建的挑战,这些图像因季节和光照条件变化而存在外观差异,且获取对齐的多日期图像和地面真实几何数据成本高昂。
Result: 在立体匹配任务中,SeasonStereo的精度与最先进的LiDAR监督模型相当,同时能生成更锐利的几何细节,且无需真实多日期训练数据或LiDAR标签。
Insight: 创新点在于利用生成式AI合成可控季节性变化的训练数据,并结合基础模型的零样本几何先验,从而减少对昂贵标注数据的依赖,实现稳健的跨季节立体匹配。
Abstract: Accurate 3D reconstruction from satellite imagery typically relies on near-simultaneous stereo pairs, limiting its applicability to diachronic settings where multi-date images exhibit varying seasonal and illumination conditions. Training dense stereo matching models robust to appearance changes is a long-standing challenge, as aligned multi-date imagery and ground-truth geometry are costly to obtain at scale. We propose SeasonStereo, a scalable framework that addresses disparity estimation from diachronic satellite images by training on synthetic image pairs with controlled seasonal appearance variation, while leveraging zero-shot geometric priors from foundation models. SeasonStereo matches the accuracy of state-of-the-art LiDAR-supervised models, while producing sharper geometric details without requiring aligned real multi-date training products or LiDAR-derived labels. As a result, SeasonStereo offers a practical path toward large-scale 3D reconstruction from heterogeneous satellite images with reduced supervision cost.
[66] Visual Credit Audit for Multimodal Spatial Reasoning cs.CV | cs.AIPDF
Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen
TL;DR: 该论文提出了视觉信用审计(VCA)方法,用于评估多模态大语言模型(MLLMs)在空间推理基准测试中的表现。VCA通过分离两个估计量来审计模型决策:一是基准图像是否比纯文本或无图像控制条件提供更多支持,二是模型是否对关系特定的视觉证据做出响应。该方法在四个开放MLLMs和两个空间基准上应用,发现12.73-26.25%的正确决策未被给予视觉信用。
Details
Motivation: 解决现有封闭式空间推理基准测试的局限性,即模型即使仅依赖文本上下文也能获得正确答案,无法准确评估图像对模型决策的实际贡献。
Result: 在四个开放MLLMs和两个空间基准测试中,12.73-26.25%的正确决策未被给予视觉信用。匹配同分割图像排列使依赖信用正确性(D-CC)降低21.25-47.80点,所有配对95%置信区间均高于零。在受控的正确但未获信用的一致决策中,对关系反转的响应率为81.57-100.00%。
Insight: 创新点在于提出了一种训练和标签无关的审计框架,将基准测试成功分解为正确性、额外图像支持和关系一致响应三个组成部分。通过控制实验(如固定像素关系对比和3x3证据源因子设计)揭示了零控制无法识别关系响应的原因,为多模态模型评估提供了更细粒度的分析工具。
Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model’s declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.
[67] Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs cs.CVPDF
Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh, Muhammad Atif Tahir
TL;DR: 该论文提出了一种通过多任务学习微调小型视觉语言模型的方法,旨在提升胃肠道内窥镜视觉问答任务的性能,并增强模型答案与图像区域之间的隐式对齐。该方法利用现有VQA数据集构建辅助的定位和描述任务,仅需少量额外标注,在Kvasir-VQA-x1数据集上对三个小型VLM骨干网络进行微调。
Details
Motivation: 胃肠道内窥镜图像分析正从单标签分类转向视觉问答,临床采纳不仅要求模型答案准确,还需要其内部表征能反映答案背后的视觉证据(即可解释性)。
Result: 在分布内和分布外数据上的评估表明,与仅使用VQA任务的微调方法相比,所提出的多任务微调方法在保持答案准确性的同时,显著改善了答案标记与相关图像区域之间的隐式对齐。
Insight: 创新点在于利用现有VQA数据集(如复用息肉掩码)和弱监督(如使用Grad-CAM定位的预训练分类器)构建辅助任务,以低成本实现模型可解释性的提升;该方法简单有效,适用于标注数据有限的医学领域。
Abstract: Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model’s internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.
[68] Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications cs.CVPDF
Piyush Jain, Kousik Dasgupta, Rajarshi Roy, Subarna Tripathi
TL;DR: 该论文提出了ByDeWay-V2框架,旨在提升多模态大语言模型(MLLMs)在决策关键应用中的空间推理能力。该框架通过整合深度线索和显式的空间关系谓词(如‘左侧’、‘内部’),以无需训练的方式增强MLLMs对细粒度空间关系的理解,并减少物体幻觉,同时保持可解释性和资源效率。
Details
Motivation: 随着MLLMs在机器人、具身AI等决策关键场景中部署,其空间判断的不透明性和在细粒度空间关系(如投影和拓扑关系)理解上的不足,限制了操作者信任和可审计性。现有方法LDP仅依赖粗略深度分层,难以解决同一几何平面内的物体间空间关系问题。
Result: 在Visual Spatial Reasoning (VSR) 和 BLINK 基准测试上评估,ByDeWay-V2显著提升了空间推理性能。例如,在BLINK空间子集上,Qwen2.5-VL的F1分数相比LDP相对提升了46%;在VSR上,BLIP-Base的性能从接近随机恢复到了具有竞争力的F1分数0.53。最轻量配置可在CPU上以40个token的上下文预算运行,适用于资源受限的实时决策场景。
Insight: 创新点在于将显式的、人类可读的空间关系谓词(通过开放词汇物体检测器计算)与深度线索结合,注入MLLM提示中,从而无需训练即可桥接3D场景深度与2D空间语义。这提供了可审计的证据,增强了模型的可解释性,同时框架设计注重资源效率,适合实时应用。
Abstract: As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination. Prior work, ByDeWay, introduced Layered-Depth-Based Prompting (LDP), a training-free framework that mitigates hallucinations by structuring prompts using monocular depth estimation. However, coarse depth layering falls short in resolving object-to-object spatial relationships within the same geometric plane, such as projective (“left of”, “above”) and topological (“inside”, “touching”) relations. We propose ByDeWay-V2, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support. Using an open-vocabulary object detector (YOLO-World-L), our framework computes pairwise geometric relations between detected objects and injects them as structured spatial predicates into the MLLM prompt, bridging 3D scene depth and 2D spatial semantics without any training. We evaluate ByDeWay-V2 on the Visual Spatial Reasoning (VSR) and BLINK benchmarks across multiple MLLMs, with hallucination grounding assessed via POPE. On the BLINK spatial subset, ByDeWay-V2 achieves a 46 percent relative F1 improvement over LDP for Qwen2.5-VL, and recovers BLIP-Base’s spatial reasoning on VSR from near-random performance to a competitive F1 of 0.53. Our lightest configuration operates under a strict 40-token context budget on CPU, showing the framework’s suitability for resource-constrained, real-time decision-support settings.
[69] Anatomy Contextualized Adaption of CT Foundation Models cs.CV | cs.AIPDF
Roshan Kenia, Stephanie L McNamara, William Lotter
TL;DR: 本文提出了一种名为解剖学上下文适应(ACA)的轻量级框架,用于调整预训练的CT视觉语言基础模型,以实现解剖级别的视觉语言对齐,同时增强全局上下文信息。该方法通过TotalSegmentator将CT体积分解为解剖级嵌入,利用Transformer捕获跨解剖关系,并与放射学报告中提取的解剖特定文本和扫描级文本对齐。在Merlin和CT-RATE基准测试中,ACA在零样本发现分类任务上优于冻结基础模型和现有细粒度方法,且训练时间短。
Details
Motivation: 现有CT视觉语言基础模型通常使用全体积表示进行训练,这会稀释细粒度解剖信号;而细粒度视觉语言预训练方法虽然对齐解剖级视觉特征与文本,但丢弃了全局上下文,且从头训练计算成本高。
Result: 在Merlin和CT-RATE基准测试中,ACA在零样本发现分类任务上一致优于冻结基础模型基线和现有细粒度方法,训练时间少于1小时(嵌入缓存后)。
Insight: 创新点在于轻量级适应框架,通过解剖分解和跨解剖关系Transformer,在保持并增强全局解剖上下文的同时实现解剖级对齐;可借鉴其结合局部细粒度与全局上下文的策略,以及利用预训练模型的效率优势。
Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA’s inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.
[70] HumanCLAW: Can Vision-Language Models Act Through a Body? cs.CV | cs.ROPDF
Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li
TL;DR: 本文提出了HumanCLAW评估框架,用于解耦视觉语言模型(VLM)的决策与低级运动控制,以评估其在物理世界中的行动智能。该框架通过将VLM生成的原子技能命令转化为具有真实物理后果的全身运动,构建了HumanCLAW-Bench基准测试集,包含1,218个长视野、以自我为中心的寻找-导航-交互任务。测试发现,当前最先进的VLM在基准上的成功率极低(最高仅16.8%),主要缺乏具身自我意识。
Details
Motivation: 现有方法难以评估VLM通过物理身体行动的能力,因为任务失败时无法区分是模型决策错误还是运动控制执行失败(如失去平衡)。本文旨在通过解耦决策与执行,专门评估VLM的瞬间动作决策智能。
Result: 在HumanCLAW-Bench(包含41个室内场景的1,218个长视野任务)上测试了9个SOTA VLM,最佳模型成功率仅为16.8%,表明当前模型均无法有效解决该基准。识别目标并非瓶颈,主要失败源于缺乏自我身体状态跟踪。
Insight: 创新点在于提出了一个解耦决策与执行的评估框架,允许VLM在真实物理环境中自由行动同时排除执行端干扰,从而专注于评估其动作智能。关键发现是当前VLM缺乏具身自我意识(如身体定位、目标到达判断和障碍感知),这为未来具身AI研究指明了重要方向。
Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM’s decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model’s action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
[71] VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion cs.CV | cs.ROPDF
Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys
TL;DR: 本文提出VidMap系统,旨在解决传统SLAM和SfM方法在无约束视频中相机标定和位姿恢复的局限性。它结合了SLAM的序列约束和离线SfM的全局优化优势,利用宽基线密集图像匹配和时间顺序信息,实现了对任意长、未标定视频的度量重建。
Details
Motivation: 现有方法如SLAM对初始化和瞬时故障敏感,且通常需要已知相机标定;而SfM缺乏对视觉对称性和极端运动的鲁棒性。本文旨在弥补这一差距,为导航和场景理解提供大规模训练数据。
Result: 在多种具有极端运动和视觉对称性的挑战性数据集上评估,VidMap比最先进的SLAM和SfM方法(无论是经典方法还是学习方法,无论相机标定已知或未知)都显著更鲁棒和准确。
Insight: 创新点在于将时间顺序作为一等公民用于可靠的闭环检测,并结合度量单目深度先验增强全局优化,从而在无约束视频中实现鲁棒的度量重建。
Abstract: Accurately recovering the camera’s calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.
[72] TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM cs.CV | cs.ROPDF
Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang
TL;DR: 本文提出了TurboVLA,一种新的视觉-语言-动作模型范式,它将传统的V→L→A路径重构为直接的V+L→A映射。该方法通过轻量级的双向视觉-语言交互独立编码视觉观察和语言指令,并使用紧凑的解码器预测连续动作块,从而在消费级RTX 4090上实现了32Hz的实时推理、低于1GB的显存占用,并在LIBERO基准上取得了与更大模型相当或更优的性能。
Details
Motivation: 传统的以LLM为中心的VLA模型(V→L→A)在每次策略调用时都会产生巨大的计算和内存开销。本文旨在设计一种更高效的VLA范式,以降低推理延迟和显存需求,使其能在消费级硬件上实时运行。
Result: 在LIBERO基准测试中,TurboVLA仅用0.2B参数就达到了97.7%的平均成功率,在RTX 4090上实现了31.2ms的推理延迟和0.9GB的推理显存占用,其性能匹配或超越了参数量大得多的VLA策略。
Insight: 核心创新在于摒弃了以LLM作为感知与动作核心接口的传统范式,转而采用轻量级的双向视觉-语言直接交互和紧凑动作解码器,直接从视觉和语言特征构建任务条件表示。这为高效连接视觉、语言和动作以实现机器人操控提供了一个新颖且有效的架构视角。
Abstract: Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional $V \to L \to A$ pathway as a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.
cs.AI [Back]
[73] Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models cs.AI | cs.CLPDF
Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani
TL;DR: 本文探究了强化学习(RL)与监督微调(SFT)微调的大语言模型在数学推理任务上性能差异的机制根源。研究发现,RL模型在隐藏层表征上具有更高的线性可分性和结构化程度,并形成了更深层更关键的分层架构,而SFT模型的层间重要性分布则更均匀。此外,研究还分析了模型在重复采样中的token数量变异性,以评估其自适应计算分配行为。
Details
Motivation: 尽管RL微调的模型在数学推理任务上普遍优于SFT微调的模型,但其性能优势的内在机制尚不明确。本文旨在探究RL与SFT模型在内部表征上的差异,以解释这种性能差距的来源。
Result: 线性探测实验表明,RL模型在预测答案正确性上比SFT模型准确率更高,表征更线性可分。平均消融研究显示,RL模型形成了深层更关键的分层架构,而SFT模型各层重要性分布均匀。在token分配变异性上,部分RL模型表现出比SFT模型更高的变异性,但并非所有RL模型都如此。
Insight: RL训练从根本上重构了模型对推理问题的表征和处理方式,使其形成了更结构化、层次化的内部表征。研究创新性地结合了线性探测、层间重要性分析和token分配变异性分析,为理解不同训练范式如何影响模型内部机制提供了多维证据。token分配变异性可作为评估模型策略稳定性和解行为确定性的潜在指标。
Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models’ superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.
[74] CG-World: A Large-Scale World-State Dataset and Protocol for World Models cs.AI | cs.CV | cs.GRPDF
Yiming Cai, Fangjie Yu, Meiqing Yu, Ziyue Shi, Pengfei Yuan
TL;DR: CG-World是一个从工业计算机图形生产流程中提取的大规模世界状态数据集和协议,旨在为世界模型提供全面的联合动态学习数据。该数据集包含约85万个1-5秒的时间对齐片段,明确记录了中间状态如多模态语义、空间结构、骨骼和控制器状态、运动曲线、相机和光照参数、物理缓存、接触事件以及多通道渲染。数据集支持干预学习和反事实推理,通过分支谱系覆盖事实轨迹、观察干预、行动干预、机制干预和严格反事实分支。
Details
Motivation: 现有视频、机器人和仿真数据集通常只捕获世界模型所需联合动态(状态、行动、事件和观察)的一部分,缺乏对中间状态的全面记录,限制了世界模型的学习能力。
Result: 在几何条件视频生成、行动预测和闭环视觉-语言-行动策略转移等任务上的评估表明,CG-World为可控生成、行动建模和具身策略转移提供了可重用的结构化监督。
Insight: 创新点在于从工业图形流程中提取并组织统一时空样本,明确分离潜在状态、观察、关系、事件和分支元数据,并定义干预分支谱系以支持反事实推理,为世界模型、物理AI和具身智能提供了共享数据基础设施的潜力。
Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.
cs.LG [Back]
[75] Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cs.LG | cs.CLPDF
Yunpeng Chu
TL;DR: 本文提出了MeRLa(元学习奖励塑形)框架,旨在改进基于人类反馈的强化学习(RLHF)。该方法通过元学习一个任务感知的奖励塑形函数,在RLHF训练前从辅助任务中学习,以生成复合奖励,从而提供更密集、任务特定的学习信号。
Details
Motivation: 当前RLHF方法依赖于静态、任务无关的奖励模型,导致学习信号稀疏和对齐效果欠佳。本文旨在解决奖励模型与任务需求不匹配的问题,以提升对齐质量和训练稳定性。
Result: 在LLaMA-3-8B模型上,于四个基准测试中(包括AlpacaEval 2.0和MT-Bench)均优于PPO、DPO、GRPO和DAPO等方法,在AlpacaEval 2.0上获得了90.8%的长度控制胜率,在MT-Bench上得分为9.14,同时训练不稳定性降低了41%。
Insight: 核心创新在于通过元学习任务感知的奖励塑形函数,结合任务判别、熵正则化和基于势能的守恒约束的元目标,在理论上保证了策略不变性并解决了熵最大化带来的激励错位问题。该方法可与基于过程和基于规则的增强奖励结合使用,提升通用性。
Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $Φ(x,y;φ)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward that preserves policy optimality while providing task-specific learning signals. Our meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence. We provide theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization. Experiments on LLaMA-3-8B across four benchmarks show consistent improvements over PPO, DPO, GRPO, and DAPO, achieving a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability. MeRLa retains its benefits when combined with process-based and rubric-based enhanced rewards.
[76] Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement cs.LG | cs.CLPDF
Haifeng Wu
TL;DR: 本文提出了IRIS框架,用于从隐式交互流中学习动态用户画像,通过从日常对话中提取行为信号,并通过预测驱动的闭环迭代优化画像表示,无需显式反馈。在合成交互流和真实世界Reddit AITA数据上的实验表明,IRIS能生成稳定的画像,并在决策预测任务上优于静态画像、仅记忆检索和无个性化基线。
Details
Motivation: 现有方法通常依赖显式偏好监督(如成对比较或人口统计属性)来个性化大语言模型,限制了其在自然交互设置中的适用性,因此需要一种直接从隐式交互流中学习动态用户画像的方法。
Result: 在合成交互流上,IRIS能生成稳定画像并区分不同用户;在100名作者的真实Reddit AITA数据上,IRIS在决策预测准确率上达到61.0%,优于所有评估方法(包括静态画像、仅记忆检索和无个性化基线)。
Insight: 创新点在于通过预测驱动的闭环迭代优化从隐式交互流中学习动态用户画像,无需显式反馈;这为个性化LLMs提供了可扩展的隐式行为建模替代方案,并为需要持续演化用户模型的自适应对话系统和具身智能体提供了实用基础。
Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. We propose IRIS, a framework that learns dynamic user personas directly from implicit interaction streams by extracting behavioral signals from everyday conversations and iteratively refining persona representations through a prediction-driven closed loop without requiring explicit feedback. We introduce an evaluation protocol based on behavior prediction, persona stability, and decision prediction. A proof-of-concept study on a synthetic interaction stream derived from public-domain autobiographical text shows that IRIS produces stable personas and distinguishes individual users while revealing limitations of memory-only approaches on recall-oriented metrics. We then validate IRIS on anonymized real-world Reddit r/AmItheAsshole (AITA) data, with personas built solely from each author’s historical interactions. Across 100 authors, IRIS achieves the highest decision prediction accuracy among all evaluated methods (61.0%), outperforming static personas, memory-only retrieval, and no-personalization baselines. These results suggest that implicit behavioral modeling provides a scalable alternative to explicit preference learning for personalized LLMs and offers a practical foundation for adaptive conversational systems and embodied agents that require continuously evolving models of their users.
cs.CY [Back]
[77] Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach cs.CY | cs.AI | cs.CLPDF
Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
TL;DR: 本文提出了一种名为认知诊断画像(CDP)的零样本框架,旨在通过提示大型语言模型(LLM)模拟具有不同认知画像的考生,以解决LLM模拟考生响应过于准确和单一的问题,从而改善心理测量校准。在分数减法数据集上的实验表明,CDP在能力分布、掌握画像和项目难度三个层面显著提升了LLM模拟考生与人类考生的对齐度。
Details
Motivation: 教育测试的心理测量校准通常依赖昂贵的人类响应数据,而LLM模拟的考生虽然为早期校准提供了可能,但其响应过于准确和均匀,无法反映真实考生的多样性。
Result: 在Tatsuoka分数减法数据集(536名考生,15个项目,五个属性)上评估了八种LLM配置。CDP显著改善了三个层面的对齐:分布重叠增加;画像级分数与人类期望的加权相关性达到0.92至0.98;项目难度恢复在排序和绝对对齐上均有提升,特别是对于具备推理能力的模型。例如,Gemini 3.0 Flash (Thinking)的单参数逻辑(1PL)难度斯皮尔曼相关性从0.24提升至0.86和0.90,均方根误差(RMSE)从6.31降至1.30和0.90。
Insight: 创新点在于将二元的属性掌握模式转化为自然语言画像进行零样本采样,并引入无信息或有信息的分布来模拟考生多样性。这为利用LLM生成更真实、多样的模拟考生数据以支持操作性测试开发提供了实用框架。
Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.
eess.IV [Back]
[78] Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance eess.IV | cs.CV | cs.LGPDF
Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker
TL;DR: 本文系统研究了胸部X光(CXR)机器学习评估中参考标准的选择如何影响模型性能评估和排名。通过收集剑桥大学医院的专家标注数据并整理MIMIC-CXR子集,论文在病理分类和图像质量评估(IQA)两个任务上证明,改变标签来源(如报告衍生标签 vs. 专家图像标注)或IQA度量标准(如SSIM、PSNR)会显著改变性能估计和模型排序,且常用指标常与临床判断不一致。
Details
Motivation: 当前CXR机器学习评估严重依赖旨在近似临床判断的参考标准,但常用的报告衍生病理标签或通用图像质量指标可能无法可靠反映真实的临床判断,这影响了模型评估的可靠性和临床有效性。
Result: 对于监督图像分类器和多种视觉语言模型,改变标签来源导致性能估计和模型排名出现显著差异;同时,IQA指标与专家对诊断可用性判断的一致性高度依赖于指标选择,SSIM和PSNR等常用指标常与专家评估不一致。
Insight: 评估参考标准的选择应被视为CXR机器学习临床有效性的核心组成部分,需要根据具体病理、成像任务和预期下游临床用途进行论证;这揭示了当前评估实践可能存在的偏差,强调了以临床相关性为中心进行模型评估的重要性。
Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.
[79] ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming eess.IV | cs.CV | cs.MMPDF
Zehao Cao, Bowei Xu, Xun Cao, Zhan Ma, Hao Chen
TL;DR: 本文提出ScalablePromptus,一种增强的基于提示的视频流传输框架。它通过语义和颜色感知的提示反演、球面线性插值以及关键的dropout训练策略,生成了具有等级顺序的提示表示,从而解决了现有Promptus框架在网络波动时因提示部分接收而导致视频质量灾难性下降的问题。
Details
Motivation: 现有基于提示的视频流传输框架(Promptus)在网络波动时,部分接收的提示会导致重建视频质量崩溃,限制了其在实际部署中的鲁棒性。
Result: 在稳定网络下,ScalablePromptus实现了适度的质量提升;在有损网络条件下,与基线相比,它将因提示截断导致的性能下降减少了82%-95%,显著提升了鲁棒性。
Insight: 核心创新在于通过dropout训练策略生成具有等级顺序的提示表示,使得接收端能够从任意截断的提示中重建有意义的视频,无需任何适配,这为基于生成的超低码率通信系统提供了关键的容错机制。
Abstract: Prompt-based video streaming transmits compact semantic prompts instead of pixel-level content for generative reconstruction, enabling ultra-low-bitrate communication. However, the state-of-the-art Promptus framework is vulnerable to network fluctuation, where partially received prompts lead to catastrophic quality collapse. We propose ScalablePromptus, which enhances Promptus with semantic and color-aware prompt inversion, spherical linear interpolation for intermediate frames, and–most critically–a dropout training strategy that produces rank-ordered prompt representations. This allows the receiver to reconstruct meaningful video from arbitrarily truncated prompts without any adaptation. Under stable networks, ScalablePromptus achieves modest quality gains. Under lossy conditions, it reduces the performance degradation caused by truncation by 82%-95% compared to the baseline, making prompt-based streaming robust enough for real-world deployment.
q-bio.NC [Back]
[80] Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition q-bio.NC | cs.AI | cs.CLPDF
Chandra Sripada, Richard Lewis
TL;DR: 这篇论文挑战了将大型语言模型(LLMs)视为与人类认知根本不同的‘外星智能’的主流观点。作者认为,尽管LLMs在物理基础、学习历史和交互环境等方面与人类存在差异,但它们在认知组织的多个核心原则上与人类认知惊人地趋同。论文从推理组织、计算架构、表征结构、预测驱动学习和类似强化学习的机制这五个维度,论证了LLMs与人类认知的深层结构相似性。
Details
Motivation: 论文旨在反驳将LLMs与人类认知的相似性简单归因于‘拟人化投射’的观点,主张这种相似性是真实且深刻的,并探讨其背后的认知科学原理。
Result: 论文未提及具体的定量实验或基准测试结果,其核心成果是理论性的,即通过系统性的维度分析,论证了LLMs与人类认知在结构上的广泛对应关系。
Insight: 论文的创新点在于提供了一个整合性框架,将LLMs的运作机制与长期在认知科学中用于解释人类智能的核心原则(如预测学习、目标导向的强化机制)联系起来,这为理解通用智能的本质提供了新的理论视角,并可能启发更类人的AI系统设计。
Abstract: LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.
cs.GR [Back]
[81] StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction cs.GR | cs.CV | cs.ROPDF
Gahye Lee, Gyoonseo Kim, Wonjong Jang, Jooeun Son, Seungyong Lee
TL;DR: 本文提出了StructureGS,一种用于铰接物体重建的结构感知高斯泼溅框架。该方法通过引入部件定向包围盒来增强空间一致性和结构连通性,从而在优化过程中解耦几何、外观和运动参数,实现高质量的重建效果。
Details
Motivation: 现有方法主要依赖光度监督,难以解耦几何、外观和运动参数,导致部件分解模糊和几何伪影。本文旨在通过结构感知约束解决铰接物体重建中的这一局限性。
Result: 大量实验表明,该方法在铰接物体重建任务上达到了最先进的性能,生成了具有清晰部件几何的高质量结果。
Insight: 创新点在于将部件定向包围盒作为结构先验,通过空间一致性和结构连通性损失注入显式结构约束,从而在3D高斯泼溅框架中实现更准确的部件分解和几何重建。
Abstract: Reconstructing articulated objects with multiple movable parts is essential for understanding object structure and enabling physical interaction. However, this reconstruction task poses significant challenges due to the entanglement of geometry, appearance, and motion parameters during optimization. Existing methods rely primarily on photometric supervision, which commonly fails to disentangle these interdependent components, resulting in poor part decomposition with blurred boundaries and geometric artifacts. To address this limitation, we introduce StructureGS, a reconstruction framework for articulated objects that integrates structure-aware guidance into 3D Gaussian Splatting. Our approach leverages oriented bounding boxes of object parts to enforce two key structural properties: spatial coherence, which constrains each part’s geometry to remain compact and spatially coherent within its designated region, and structural connectivity, which enforces physically plausible contact relationships between adjacent parts. These properties are realized through structure-aware losses that inject explicit structural constraints into the optimization process. Extensive experiments demonstrate that our method achieves state-of-the-art performance in articulated object reconstruction, producing high-quality results with well-defined part geometries.
cs.RO [Back]
[82] ContactFlow: A video action conditioning that transfers across embodiments cs.RO | cs.CVPDF
Sami Azirar, Enrico Pallotta, Jan Nogga, Jürgen Gall, Sven Behnke
TL;DR: 本文提出了一种名为Contact Flow的与具体执行器无关的动作表示方法,它通过编码执行器与目标物体之间3D接触点的轨迹来描述操作。该方法允许在人类演示和机器人执行视频上训练大规模视频生成模型,从而构建一个能预测物理合理操作结果的世界模型,并集成到一个提出-想象-验证-执行的规划流程中。
Details
Motivation: 当前基于视频的世界模型在捕捉操作(特别是接触)的物理约束方面存在困难,且其动作条件通常局限于特定的执行器(如平行夹爪)。本文旨在解决这些问题,实现跨具身(如从人类到不同机器人)的技能迁移。
Result: 在DROID数据集和真实世界桌面操作任务上的实验表明,Contact Flow能够实现从人类演示到不同机器人执行器的迁移。
Insight: 核心创新点是提出了一个抽象掉执行器特定外观和运动学的、以3D接触点轨迹为中心的动作表示,这为跨具身视频生成世界模型提供了统一的调节信号,从而促进了物理合理的预测和技能迁移。
Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.
[83] Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots cs.RO | cs.CVPDF
Hung Nguyen, Kim Nhat Minh Nguyen, Van Duc Vu, Van-Danh Le, Hoang Huy Le
TL;DR: 本文提出Speech2Grasp框架,旨在以数据高效的方式将基于文本条件的抓取检测模型迁移至语音输入,用于人形机器人。通过轻量级MLP投影器适配ALBEF模型,在保持语义判别和鲁棒性的同时,实现了优于级联ASR方案的性能并降低了推理延迟。
Details
Motivation: 解决人形机器人多模态交互中,现有视觉-语言模型通常依赖文本而非更自然的语音输入的问题,探索如何以数据高效的方式将成熟的文本条件模型迁移到语音领域。
Result: 在真实人形机器人实验中,Speech2Grasp在抓取检测任务上超越了基于ASR的级联流水线方法,同时减少了推理延迟。
Insight: 创新点在于通过轻量级MLP投影器实现文本到语音模态的高效迁移,为将现有文本条件系统扩展至语音输入提供了一种实用范式,且保持了模型的语义区分能力和鲁棒性。
Abstract: Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.
[84] From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching cs.RO | cs.CVPDF
Shiyong Meng, Bolei Chen, Ping Zhong, Yang Wan, Rongzhi Wang
TL;DR: 本文提出了一种从不确定性到确定性的粗到细视觉平面图定位框架,无需进行光线匹配。该方法通过图像条件化的位姿扩散模型参数化连续多模态位姿分布,在粗粒度阶段引导随机初始化的位姿粒子向不同的候选模式汇聚;在细化阶段,利用局部优化器从候选位姿为中心的平面图裁剪中预测有界的亚米级位姿残差。该方法在S3D(完整)和ZInD基准测试中实现了最先进的精度和鲁棒性。
Details
Motivation: 视觉平面图定位由于跨模态信息不对称和室内布局重复性,本质上受到多模态位姿分布的挑战,即视觉上相同的观测可能对应空间上分离的不同位置。现有基于光线匹配的方法需要预测稀疏的几何或语义光线,导致信息损失,并且需要资源密集的预处理和推理时的穷举匹配。
Result: 在S3D(完整)和ZInD基准测试上的综合结果表明,该方法在精度和鲁棒性方面达到了最先进水平。
Insight: 创新点在于绕过了中间的光线匹配范式,采用粗到细的两阶段框架:粗粒度阶段利用扩散模型处理多模态不确定性,细化阶段在局部消除结构歧义进行精确优化。该方法无需任何离线地图预处理或测试时查找表,有效平衡了全局多假设跟踪和局部亚米级细化。
Abstract: Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.
[85] DLAM: Distributional Latent Actions with Temporal Constraints cs.RO | cs.AI | cs.CVPDF
Zuojin Tang, Feifan Luo, Haoyun Liu, Botai Yuan, Dekang Qi
TL;DR: 本文提出了DLAM(Distributional Latent Actions with Temporal Constraints),一种分布式的潜在动作模型,用于解决视觉-语言-动作(VLA)模型中机器人动作标注数据稀缺的问题。该方法利用无动作标签的视频数据,通过将每个状态转移建模为对角高斯分布,并结合时间约束(如归一化组合和反转),学习到更具时间一致性的潜在动态。该模型在保持编码器冻结的情况下,通过流匹配策略联合生成转移序列和机器人动作,并在多个基准测试中提升了策略性能。
Details
Motivation: 动机在于解决VLA模型因机器人动作标注数据稀缺而受限的问题,同时利用丰富的无动作视频观察数据。现有潜在动作模型虽能提取先验,但其重建训练的编码可能缺乏与机器人动作联合生成所需的结构,且确定性转移点可能导致局部推断误差在递归组合中传播和累积。
Result: 在保留的转移序列上,DLAM比现有潜在动作基线学习了更具时间一致性的潜在动态,并在保留视频上实现了更强的直接和累积重建能力。在相同的受控策略转移协议下,DLAM在MetaWorld MT50、LIBERO和真实世界操作任务上提升了策略性能。消融实验表明,归一化均值约束贡献了大部分重建增益,而学习的方差和相关性感知组合为下游控制提供了补充改进。
Insight: 创新点在于将状态转移建模为对角高斯分布,通过参考帧条件重建将均值锚定在观察到的视觉变化中,并利用等间隔三元组的归一化组合和反转来约束均值和逐维度方差。方差组合使用轻量级共享相关系数来考虑共享中间帧的相邻转移间的依赖性,而反转则取反均值并保持方差。这为从无动作视频中学习结构化潜在表示以支持机器人策略学习提供了新思路。
Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $π_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.